Claude Opus 5.5 Has a Measurable Writing Fingerprint

Graphite has catalogued 13,000 phrases that occur at least twice as often in AI-generated prose as in human writing, using that twofold gap as the working definition of a tell TechCrunch. The total shows scale. This is not a handful of giveaway words. It is a large and measurable tilt in vocabulary.
Claude Opus 5.5 has its own fingerprint inside that set. It uses "dependable" 23 times more often than the human samples, which Graphite identifies as its biggest tell. Two explanatory framings score higher on raw ratio. "This matters" appears 116 times more often than in human writing. The variant "why X matters" appears 92 times more often.
The comparison rests on 10,000 articles published before the release of ChatGPT, treated as the human baseline. Graphite also compared model generations. In its samples, Opus 5.5 used em dashes 99% less often than Opus 5.
The broader context here is where these habits come from. They are not deliberate choices. They are residue from instruction tuning and preference data, the training stages where a model is rewarded for being helpful, explicit and reassuring. Instruction tuning means fine-tuning on examples of following directions. Preference data means human rankings of better answers. Small shifts in sampling, ranking and style rewards can move thousands of small word choices at once.
Looking at detection, the lesson cuts both ways. A 116-fold ratio sounds decisive. In practice phrase-level signals are brittle. Writers adapt, prompts suppress them, and the next fine tune redistributes them. The em dash swing shows the problem. A marker tied to one generation all but vanished in the next, without the underlying challenge changing. For teams maintaining classifiers, that volatility complicates precision, recall and calibration over time.
In my view the durable value is not a checklist for spotting machine text. It is feedback for training and evaluation. If "dependable" is overrepresented by 23 to 1, that points to a narrowed patch in the lexical distribution, a place where sampling has collapsed toward a safe default. For teams tuning prompts, filters or style controls, that skew is actionable. It can be measured, penalized and widened during iteration.
Worth flagging for readers who review a lot of synthetic copy is where these tells cluster. "This matters" and "why X matters" do structural work, telling a reader where to look and why to care. That habit fits an assistant optimized to be useful. It also flattens prose. Human technical writing more often lets the implication sit unstated, trusting a knowledgeable reader to connect it. I have watched my own children move from asking software for answers to asking for summaries of why the answer counts, and the models have obliged by becoming more didactic.
Looking ahead, the practical opening is better style control. System instructions, decoding constraints or post-processing, different ways of steering word choice, can recover much of the lost variance without sacrificing clarity. The long arc favors writers and builders, not detectors. Once a skew can be counted at this resolution, it can be managed.


