Skip to content
AI Watermark Remove

How the score is built, and how we know it works

Most tools in this space give you a number and ask you to trust it. This page is the opposite: what the score measures, what it is tested against, what it scores today, and what we deliberately refuse to count.

What the score is, and what it is not

The style score is a count of surface habits that unedited assistant text tends to have. It runs from 0 to 100, and two numbers on that scale are load-bearing: 25 is where a text starts reading as possibly machine-shaped, 55 is where it reads strongly like unedited model output. Those two boundaries are chosen against a corpus, not picked by eye — more on that below.

It is not a probability that AI wrote the text, and it is not a watermark detector. It reads how the words are arranged and which phrases recur. It can be high on a careful human writer whose prose happens to be formal and even, and it can be low on model output that has been edited. Both facts are stated in the report rather than smoothed over, because a score that pretends to certainty is worse than no score.

Two kinds of evidence

The score adds up two families of measurements. The first is lexical: stock phrases, connective openers, the closing offer to revise. Precise, and narrow — when one fires it is nearly always right, and a writer told to "write plainly, no filler" evades all of them at once. A scorer built only on these is a phrase list, and a phrase list is an arms race the phrases always eventually lose.

The second family measures arrangement rather than vocabulary. Four measurements, none of which can be evaded by avoiding a phrase list:

Punctuation range
Model prose runs almost everything through the comma and the full stop. Human writing of the same register spends a visible share on the colon, the bracket, the question mark, the semicolon. We measure how much of the punctuation falls outside the two default marks — the strongest single measurement on the distributional side.
Sentence-length shape
The variance of sentence length alone is easy to satisfy: a model alternates a 14-word sentence with a 22-word one and reads as varied. So the range is read alongside the variance — what a human draft has and model output does not is a five-word sentence sitting next to a thirty-five-word one.
Lexical evenness
Model prose reaches for a synonym where a person repeats the word they just used. That habit reads as a rich vocabulary and measures as a flattened frequency distribution: the most-used words are less dominant than they would be in human prose.
Recurrence dispersion
The same habit seen along the text. Someone writing about moorland says "moor" four times in one paragraph and then not again; a model spaces its repeats almost evenly. We measure how evenly repeats are scattered, windowed so text length does not masquerade as style.

On top of these sit the visible tells the report names directly: dash density, curly quotes, bolded bullet lead-ins, templated structure, chat-assistant address, hedging. The full list the scorer can currently fire on:

  • Stock AI phrasing ("plays a vital role", "it is important to note")
  • The closing offer to revise
  • Unnaturally even sentence rhythm
  • Heavy dash use
  • Typographic quotes and apostrophes
  • Repeated "X, Y, and Z" triples
  • Bulleted lists with bolded lead-ins
  • Templated article structure
  • Chat-assistant address ("As an AI language model…")
  • Agentless hedging

What it is measured against

A scorer without a labelled corpus is an opinion. This one is tested on every change against 49 samples in two languages:

Human (31 samples): prose published before assistants could have written it — Wikipedia and Wikivoyage revisions as of mid-2018, and essays published between 2001 and 2005. Chosen to be hard rather than easy: formal, well organised, about the subjects people ask assistants about. That is exactly the confounder that makes a naive scorer call competence an AI tell.

Machine (18 samples): output from the major assistants in both languages, plus ten samples written specifically to evade phrase lists — plain register, no em-dashes, no stock phrases, casual where casual fits. Those ten are the ones that used to score as human on the phrase-list-only version.

The numbers, as of 1 September 2026

The suite reports separation as area under the curve — the probability that a random machine sample outscores a random human one. 0.5 is a coin; 1.0 is perfect separation.

Both languages
0.963
18 machine samples against 31 human ones.
English
0.953
10 against 18.
German
0.986
8 against 13.

Before the distributional measurements existed, the same corpus scored 0.74 — close enough to a coin that lightly-edited model output landed in the "reads human" band. The band outcomes matter more than the average: no human sample scores in the high band, and three of the eighteen machine samples score below 25. Two of those three are the deliberate residue of honest thresholds; the suite documents them rather than tuning them away. One is long, varied model prose that reads like careful writing on every measurement because on every measurement it is.

The suite asserts floors, not decimals: it requires separation above 0.9 and fails on a regression below it, and it does not chase the last sample, because forty-nine samples do not support a claim about the third decimal place — and a threshold tuned until the corpus is perfectly split is a threshold that calls ordinary human writing a machine.

What we deliberately refuse to score

Every candidate measurement is audited by class and by source before it earns a place in the score, because a signal that separates the classes by separating the genres inside the human corpus is worse than no signal. Three candidates were computed, measured on the corpus, and excluded:

Function-word share
How much of the text is made of "the", "of", "and". It measured like the strongest signal in the file — and it is a genre detector. Every human essay lands high and every human encyclopedia entry lands low, which means on a mixed corpus it is a coin weighted by genre, not evidence of authorship.
Vocabulary richness
Type–token ratios fall as texts get longer, so the windowed version is honest — and still tracks register harder than authorship. Computed, reported internally, and excluded from the score.
Average word length
Measured at chance on the corpus in both languages. It is in the report because removing it silently would be hiding a null result; it contributes nothing.

The same audit runs over the phrase lists: every pattern in both language packs is run across the human corpus, and a pattern that fires only on human samples is removed. A pattern matching human writing is not evidence of anything except that we wrote a bad pattern.

Where the thresholds come from

Every threshold is per-language, because German inflects and compounds in ways that move all of these measurements for reasons that have nothing to do with who wrote the text. And each threshold sits at roughly the lower quartile of the human corpus rather than at the midpoint between the classes. A threshold placed to split the two classes evenly maximises corpus accuracy and calls every second honest essay machine-shaped — which is the failure this product cannot afford, and the reason the one hard machine sample that still slips through is left in on purpose.

The raw severity points are not the score. The sum goes through a fixed monotone curve so that the two band boundaries mean the same thing on a 200-word paragraph and on a 2,000-character document. Adding the points up and clamping at 100 put a third of ordinary human writing above the first boundary; the curve exists so that 25 and 55 are worth what the interface says they are.

What would change this page

These numbers describe this scorer on this corpus today. They are re-measured on every change to the scorer, and this page carries the date of the last measurement. What would make them stronger: a larger and more varied human corpus, especially informal prose in German; and, on the watermark side rather than the style side, a public verifier from a vendor — the moment one exists, every statistical claim in this space becomes testable, including ours.

What this page is not: a claim that the score identifies authorship. It identifies habits. A reader who wants to know what is definitely in their file — the characters, the metadata — gets definite answers from the checker itself; the score is the part where the honest answer is "here is what a careful reader would notice".

Try it on your own text

The check is free and complete: every finding, the score, and the exact words that produced it.

Check my text