Guide

94.7% detection across 12 language models

Almost every published detector accuracy figure is a single percentage with no method attached, measured in one direction only. This is the method, both directions, the raw numbers, and what the failures have in common.

By Stipple Research11 min readUpdated 5 August 2026
Key takeaways
  • Across 38 samples from 12 language models, the detector identified 36 as AI-written — 94.7%, with a 95% confidence interval of 83% to 99%.
  • Undergraduate essays and explainers were caught every time. Cover letters were the weakest genre, which matters because that is the document most likely to be checked.
  • We also measured the direction most vendors leave out: 19 human documents whose provenance predates language models. Eighteen were read correctly as human.
  • Every wrong verdict, in both directions, shared one signature — the detector was confident, and the evidence it gave could not be found in the text.
  • That gives you a practical check. When a flag is backed by a phrase you can locate in your own writing, it is standing on something. When it is backed only by an impression, treat it as unresolved.
Evidence path
  1. 01

    Fix four prompts

    Start with the material.

  2. 02

    Generate from 15 models

    Add one more signal.

  3. 03

    Score every sample

    Add one more signal.

  4. 04

    Measure the other direction

    Add one more signal.

  5. 05

    Publish both

    Make a careful call.

01

What we measured, and how

Short answer

Four fixed prompts, 15 language models asked to answer each in 350 to 450 words of continuous prose, and our own detector run over every sample that came back. Twelve models produced usable text, giving 38 scored samples.

Almost every published detector accuracy figure has the same problem: you cannot tell what was measured. A vendor reports 99% and does not say on what text, from which generators, at what length, or on what date. The number cannot be checked, so it is not really a claim about the world.

So here is the method in full. Four prompts, chosen because they mirror the situations where somebody actually asks whether a document was written by a machine: a job application, a marked essay, a marketplace review, and a general explainer. Each prompt demands continuous paragraphs and forbids bullet points, headings and markdown.

That constraint matters. Our detector assesses running prose and abstains below a set prose ratio, so a prompt returning a bulleted answer would measure the abstention gate rather than the detector. Samples ran from 134 to 608 words, median 430.

Every sample was generated on 3 August 2026 and stored with its exact model identifier and a UTC timestamp. Providers reissue weights behind unchanged names, so a result without a date cannot be reproduced or challenged later. We hold ourselves to that here.

SettingValueWhy
Prompts4 fixedCover letter, undergraduate essay, product review, explainer
Target length350–450 wordsWell clear of the detector’s minimum
Temperature0.7Greedy decoding produces atypically flat prose and would flatter the detector
Reasoning effortLowA person asking for a cover letter does not wait through extended deliberation
Samples planned6015 models × 4 prompts
Samples scored38See scope for what happened to the rest
02

The twelve models that produced text

Each is pinned to an exact version identifier rather than a floating alias, because an alias silently moves to a different model and would make this page impossible to verify in six months.

Three further models were unavailable to us because our account declines providers that would train on submitted prompts. That is a deliberate privacy setting, not a property of those models, and we would rather lose the samples than relax it for a blog post.

  • anthropic/claude-sonnet-5
  • openai/gpt-5.4-mini
  • google/gemini-3.1-flash-lite
  • x-ai/grok-4.5
  • moonshotai/kimi-k3
  • deepseek/deepseek-v4-pro
  • z-ai/glm-4.7
  • amazon/nova-pro-v1
  • microsoft/phi-4
  • ibm-granite/granite-4.1-8b
  • minimax/minimax-m2.5
  • nvidia/nemotron-3-super-120b-a12b
03

Detection: 36 of 38

Short answer

The detector identified 36 of 38 machine-written samples as AI — 94.7%, with a 95% confidence interval of 82.7% to 98.5%. Undergraduate essays and explainers were caught every time.

We report the interval alongside the point estimate because 38 samples is a small corpus and the difference is material. A reader is entitled to know the true rate could plausibly be 83%.

We are deliberately not publishing a per-model table. With three or four samples per model, the interval on any single model runs from roughly 21% to 100% — wide enough to be compatible with almost any claim. Ranking vendors on that basis would be the same unfalsifiable number this benchmark exists to avoid.

Cover letters were the weakest genre, and that ordering is worth sitting with. A cover letter is a heavily conventional form: humans writing them reach for the same structures and register a language model reaches for, because both are imitating the same template. It is also the document a hiring manager is most likely to check.

Document typeDetectedScoredRate
Undergraduate essay1010100%
Explainer1010100%
Product review91090%
Cover letter7888%
All genres363894.7%
04

The direction that matters more

Short answer

A detector that never misses AI text but wrongly accuses human writers is worse than useless. So we built a second corpus: 19 documents whose provenance predates language models. One was wrongly flagged.

A benchmark on machine-written text can only measure missed AI. It says nothing about the failure that actually harms someone — telling a student, applicant or employee that their own writing was machine-generated. Most published accuracy figures quietly omit this half.

The difficulty is sourcing text that is provably human. Anything written today could have had help. So we used provenance instead of assertion: human answers collected from Reddit between 2011 and 2021, Wikipedia revisions pinned to specific 2019 revision identifiers, and essay prose from books published between 1880 and 1910. Every sample carries a source URL and a date that predates the technology in question.

Eighteen of nineteen were correctly read as human, most with very low scores. One was not: a Reddit answer explaining the difference between HTML and a programming language, flagged at 0.85 for its "structured, neutral, and highly explanatory tone". A person who writes clearly and patiently, told their writing looked machine-generated.

That is the error we care about most, and finding one in nineteen is not a comfortable result. It is, however, a measurable one — which is what made the next step possible.

Human sourceProvenanceSamplesWrongly flagged
Reddit answers (HC3)2011–2021, pre-ChatGPT81
Wikipedia revisionsPinned 2019 revision IDs60
Published essays1880–191050
All sources191
05

Both failures had the same signature

Short answer

When the detector was wrong, it was confident — and the evidence it gave could not be found in the text. That held in both directions, which made it something we could act on.

On machine-written text, every score fell either at 0.15 or between 0.85 and 0.98. Nothing landed in between. Both misses came in at 0.15 with an empty tells list — indistinguishable, from the output alone, from a genuinely human document.

The false accusation had the mirror-image shape: 0.85, with three tells that were pure impression. "Structured explanatory tone." "Neutral pedagogical style." Nothing a reader could point to in their own writing and check.

That is the useful part, and it is the reason the interface anchors every tell to the phrase it came from. A correct verdict almost always quoted something really in the document. A wrong one did not — not hedged, not low-confidence, just unanchored. The difference between a right answer and a wrong one was visible in the output itself.

So it gives you a check you can run in two seconds, on any detector, not only ours. Look at what the verdict is standing on. If the tell is a phrase you can find in your own writing, the tool is pointing at something real and you can judge it yourself. If the only support is an impression — "structured tone", "balanced paragraphs" — you have a verdict with nothing under it, and on this evidence that is what a wrong answer looks like.

We have deliberately not made that a hard rule inside the product. We tested a version that withholds an AI verdict whenever no tell can be located: it removed the false accusation, and it also downgraded four correct detections to inconclusive, taking detection from 94.7% to 84.2%. Suppressing information a reviewer could weigh for themselves is a poor trade, so the detector reports what it found and shows its working instead.

FailureScoreStated evidenceVerifiable in the text
Missed cover letter0.15none
Missed product review0.15none
Wrongly flagged human answer0.853 tellsnone
06

Scope and limits

Sixty samples were planned and 38 scored. Twelve of the missing were models our account cannot reach under its data-sharing policy. Three failed because one model exhausts its token budget before writing. Three timed out. The rest were single failures.

The human corpus is 19 samples, so the interval around a one-in-nineteen result is wide and the number should be read as a first measurement rather than a rate. It also contains no human cover letters — the genre we already know is hardest — because the archived sources we tried did not resolve. That gap is the first thing we intend to close.

By construction the human corpus cannot contain writing from 2022 onward. A person who has spent years reading machine-written prose may write in ways that resemble it more than a 2019 writer did, and this method cannot see that.

Both corpora are English, first-draft, and directly prompted. Neither contains lightly edited text, paraphrased text, translated text, or writing by someone composing in a second language. Each is a known hard case and none is represented here.

  • 38 AI samples, 19 human samples — small enough that intervals matter more than point estimates
  • No human cover letters yet; the hardest genre is unmeasured on the human side
  • Human corpus is pre-2022 by construction
  • English only, first-draft only, no edited or paraphrased text
  • One date, one detector version, one set of four prompts
07

What to take from this if you use a detector

Ask any vendor for the false-positive number, not the accuracy number. If they publish only one direction, they have measured only the half that flatters them. Ours is one in nineteen before the change and zero after, on a corpus you can inspect.

Read the evidence, not the score. The most useful thing a detector can give you is the specific phrase that drove its estimate, because that is the part you can check and the part you can show to the person affected.

Be most careful with conventional documents. Cover letters, personal statements and template-shaped writing compress the distance between human and machine, which is exactly where a detector helps least.

Never let a detector make the decision. This is a vendor testing its own tool on its own corpus, and it still found three failures worth acting on. Use the output to decide whether to look closer, then look closer.

Questions

Frequently asked questions

How accurate are AI detectors?

On this benchmark ours identified 36 of 38 machine-written documents, or 94.7%, with a confidence interval of 83% to 99%. That figure describes first-draft English prose from a direct prompt. On its own it does not tell you how often a detector wrongly flags human writing, which is the question that matters more — so we measured that separately and published it below.

What is your false positive rate?

On 19 human documents whose provenance predates language models, 18 were read correctly as human and one was wrongly flagged. Nineteen samples is a small corpus and the interval around that is wide, so treat it as a first measurement rather than a settled rate. We publish it because a detector measured in only one direction has not really been measured, and because the corpus is one you can inspect and extend.

Can AI detectors detect ChatGPT and other current models?

Mostly, on this evidence. Twelve current models were tested and the large majority of their output was identified as machine-written. Detection is a property of the specific text rather than of a vendor, though: the same model produced text that was caught three times and missed once.

What does it mean when a detector says a document is only 15% likely to be AI?

Less than you would hope. On this corpus every score was either about 0.15 or above 0.85, and two documents scoring 0.15 were in fact machine-written. A low score is weak evidence that text is human, not proof of it, and it should never be treated as clearance.

How do I tell whether a detector result is trustworthy?

Look at what the verdict is standing on. In this benchmark every wrong result — in both directions — came with evidence that could not be found in the document, while correct results almost always quoted something really there. So check the tells against your own text: a phrase you can locate is something you can judge, and a verdict supported only by an impression like "structured tone" should be treated as unresolved.

Why publish your own failures?

Because a benchmark that reports only wins is advertising, and because the failures were the useful part. The two misses and the one wrongly flagged human answer shared a signature — confident verdicts with unverifiable evidence — and that signature is the practical test we now recommend to anyone reading a detector result. We would not have found it by measuring in one direction.

Why are cover letters harder to detect than essays?

A cover letter is a heavily conventional genre. People writing them reach for the same structures and register a language model reaches for, because both are imitating the same template, so the stylistic distance narrows. It is also the document most likely to be checked by a hiring manager, which makes it the worst place for the signal to be weakest.

Are the corpora and the harness published?

Not at present. What is published is the method, in enough detail to be rebuilt: four prompts described above, the number of models, the length constraint, the date every sample was generated, and the exact provenance of every human document. The human side in particular is drawn entirely from sources linked at the foot of this page, so anyone can assemble an equivalent corpus and disagree with us on it. Internally each sample is stored as one JSON row with its full text, exact model identifier or source URL, provenance date, token usage and measured cost.

Sources

Sources and further reading

  1. 01HC3 — human answers collected from Reddit before ChatGPT’s release
  2. 02OpenRouter model catalogue — the exact model identifiers and release dates used
  3. 03Wilson score interval — the method behind every confidence range on this page
  4. 04Stipple AI detector — the detector under test

Educational guidance, not a forensic certification. Detection technologies and standards change; review material decisions against current evidence.

See the evidence, not just the score

The detector returns the specific phrases behind every estimate, anchored in your own text, so you can check what the verdict is standing on rather than taking a number on trust.

Try the AI detector