AI RACE— The AI Race
Language Models

Frontier Models Still Have Signature Tells Despite Efforts to Eliminate Them, Study Finds

A study by marketing firm Graphite identified over 13,000 phrases disproportionately favored by modern LLMs, showing that while models have phased out old habits like em-dashes, new stylistic tics continue to emerge.

10/02/2026, 00:50
Language Models

A New Map of Generative AI's Linguistic Fingerprints

Even as early artificial intelligence writing hallmarks such as the em-dash and the word "delve" fade from view, state-of-the-art language models continue to rely on distinct, repetitive linguistic habits. A study conducted by marketing firm Graphite analyzed the prose of leading frontier models to identify the phrases and constructions that separate synthetic text from human output.

To isolate these patterns, Graphite built a control dataset of 10,000 human-authored articles published prior to the launch of ChatGPT. Researchers then instructed several cutting-edge AI systems to rewrite these pieces based solely on summaries, minimizing bias inherited from the original source text. In analyzing the paired outputs, Graphite pinpointed roughly 13,000 phrases that appeared at least twice as frequently in machine-generated text as in the human baseline—a frequency the firm used as the baseline definition of a "tell."

The findings highlight a divergence in how different model families mirror human prose over time. "It turns out that Claude models are actually getting closer to the human word distribution over time," Graphite Chief AI Officer Greg Druck told TechCrunch. "And for the GPT models, it’s getting further away."

How Opus 5.5 and Astra Reveal Themselves in Prose

The study found that individual models exhibit sharply defined stylistic signatures. Anthropic’s Claude Opus 5.5, for instance, frequently leans on the word "dependable," utilizing it 23 times more often than human writers. While the model has largely abandoned the classic "it's not X, it's Y" antithesis, it has substituted another contrast formula, frequently asserting that a subject "is more than an X, it's a Y." Opus 5.5 also shows an overwhelming tendency to emphasize significance: the exact phrase "this matters" appeared 116 times more often than in the human control group, while "why X matters" surfaced 92 times more frequently.

By contrast, OpenAI’s Astra model presents an entirely different vocabulary. Astra heavily favors the phrase "another dimension" and frequently softens claims with speculative hedges such as "may provide" or "can provide." Its most pronounced tell is what Graphite terms "corrective framing"—constructions that frame topics as "not simply X" or propose alternatives "rather than relying on X." Graphite recorded these corrective patterns more than 100 times more often in Astra’s text than in human writing.

Punctuation habits have shifted drastically across the board. In response to widespread criticism regarding em-dash overuse, frontier labs appear to have tuned their models away from the mark. Opus 5.5 used em-dashes 99% less often than its predecessor, Opus 5. Astra deployed the punctuation mark 88% less frequently than the human benchmark, while Google’s Gemini 3.1 Pro has nearly excised the em-dash from its prose entirely.

The Difficulty of Scrubbing Tells from Massive Models

Despite targeted updates designed to eliminate well-known stylistic giveaways, Graphite found that the sheer volume of AI tells remains broadly consistent. Rather than disappearing, the tells simply rotate as new architectures and fine-tuning datasets are deployed. "It’s not like the tells are decreasing," Druck said. "They are managing to remove the most well-known tells, but other ones pop up. And every model version has its own."

The persistence of these habits runs counter to the positioning of major AI labs, which have consistently marketed newer releases as more fluent and natural. Anthropic promoted Opus 5.5 by highlighting that early testers found its prose clearer and easier to follow, while OpenAI made similar commitments with GPT-6 versions Sol and Luna, promising reduced jargon and fewer strange phrasing choices.

Druck suggested that completely ironing out stylistic quirks may be beyond the reach of existing training techniques. "A general hypothesis I have is that the labs are less able to control some of these things than you might expect," Druck noted. "These are giant models with billions of parameters. They have some finite number of tests they can run, and things slip through."

◗ Sources

TechCrunch10/02

Related stories