What we measured, and how
Short answer
We asked fifteen language models for two short reports, each required to cite at least five sources with full URLs. Nineteen usable reports came back with 101 citations between them. Every citation was resolved by machine, and every failure was then verified by hand before it was allowed to count as fabricated.
Hallucination statistics are usually quoted with no method attached: “AI makes things up X% of the time”, source unclear, task unstated, date missing. We wanted a number we could defend, for the narrow question that matters most in documents: when a model cites a source, does that source exist?
Two prompts, chosen to represent the situations where people genuinely ask AI for sourced research — and to vary one thing deliberately: how well-covered the topic is. The first asked for a research brief on remote work and productivity, one of the most-published workplace questions of the decade. The second asked for a policy brief on AI-writing detection in university integrity processes — a real but far thinner literature. Both demanded at least five cited sources with full URLs, and both forbade placeholder links.
Each report then went through the same reference checker that powers our public fact-check tool: every URL resolved, paper identifiers matched against the papers they claim to be, claims tested against the sources that survived, internal arithmetic recomputed.
The run cost eleven cents of generation and produced every number in this article.
| Setting | Value | Why |
|---|---|---|
| Models asked | 15, version-pinned | The same set as our detection benchmark; five produced nothing usable (see the model list) |
| Reports checked | 19 | Ten models produced reports — nine completed both genres, one completed one |
| Citations required | ≥5 per report, with full URLs | URLs make verification mechanical — and raise the pressure to fabricate when the model has no real source |
| Citations collected | 101 | Every one resolved, none sampled |
| Verdicts hand-verified | All 50 failures | A machine “not found” is a candidate, never a conclusion |
The ten models that produced reports
Each is pinned to an exact version identifier rather than a floating alias, because an alias silently moves to a different model and would make this page impossible to verify later. Nine models completed both briefs; Nemotron completed one of its two.
Five produced nothing usable and are absent, not innocent: three returned our account’s data-policy refusal (we decline providers that train on submitted prompts, and would rather lose the samples than relax that for an article), and two failed to return a usable report. The rate describes the reports we could check.
- anthropic/claude-sonnet-5
- openai/gpt-5.4-mini
- google/gemini-3.1-flash-lite
- x-ai/grok-4.5
- amazon/nova-pro-v1
- moonshotai/kimi-k3
- z-ai/glm-4.7
- microsoft/phi-4
- ibm-granite/granite-4.1-8b
- nvidia/nemotron-3-super-120b-a12b (one of two briefs)
What counts as a fabricated citation
Short answer
A citation counts as fabricated only if the reference as written does not exist: its domain does not resolve, or the page returns a definitive not-found — confirmed by a second, browser-identical fetch. Sources that merely blocked our checker were excluded, not counted.
Definitions decide benchmarks, so ours was locked before the run. “Fabricated as cited” means the specific reference the model wrote — that URL, that identifier — does not exist. Sometimes a similar real article exists somewhere else; that does not rescue the citation, because a reference whose details are invented sends every reader who checks it to a dead end.
The conservative side of the definition matters just as much. Five citations pointed at real publishers whose servers refuse automated requests — the paywalled and bot-walled academic press. We could not confirm those references, and so they were excluded from the fabrication count entirely. Unverifiable is not the same as invented, and a benchmark that conflates them inflates its own headline.
One more exclusion by design: a report that cited no URLs at all would have been classed “unverifiable sourcing”, not fabrication. None of the nineteen took that exit — every model confidently produced linked references.
The results: 38.6% overall, doubling on the thin topic
Short answer
Of 101 citations, 39 did not exist as cited — 38.6%, with a 95% interval of 29.7% to 48.4%. On the heavily published topic the rate was 24.5%; on the thinner topic it was 54.2% — a majority of everything cited.
The split between the two genres is the finding. The remote-work brief draws on one of the deepest evidence bases in workplace research, and three-quarters of its citations were real. The policy brief, on a subject with a real but much thinner literature, crossed into majority-fabricated: models kept producing five confident references for a topic that does not have five canonical sources to give.
That is exactly the mechanism our hallucination guide describes: a language model knows the shape of a citation far better than the substance of one. Where the literature is dense, shape and substance coincide. Where it thins, the model keeps the shape — a plausible newsroom URL, a valid-format journal identifier, a believable university page — and invents the substance.
What the fabrications looked like matters too. These were not garbled strings. They were a news article on a real newspaper’s domain that was never written; a guidance page on a real university’s site that does not exist; journal identifiers in perfect publisher format resolving to nothing. Every surface signal of legitimacy, attached to nothing.
One citation failed differently: a real, resolvable paper attributed a claim it does not make — the mis-attribution failure, rarer here but harder to catch by resolution alone. And in the other direction, none of the claims tested against sources that did resolve were contradicted by them: when the source was real, it tended to genuinely support the sentence citing it.
| Corpus | Fabricated | Citations | Rate | 95% interval |
|---|---|---|---|---|
| All reports | 39 | 101 | 38.6% | 29.7% – 48.4% |
| Research brief — dense literature | 13 | 53 | 24.5% | 14.9% – 37.6% |
| Policy brief — thin literature | 26 | 48 | 54.2% | 40.3% – 67.4% |
| Excluded as inconclusive (bot-walled) | — | 5 | not counted | — |
| Real paper, wrong claim | 1 | — | mis-attribution | — |
What a fabricated citation actually looks like
Short answer
Not garbled text — perfect form. A fully-styled academic reference naming a real journal, with invented authors, an invented article, and an invented domain. A news URL on a real newspaper’s site for an article that was never written. Every surface signal of legitimacy, attached to nothing.
Four fabrications from this run, exactly as the models wrote them (URLs abridged). Each was confirmed nonexistent in the hand-review.
The Phi-4 example repays a close look. The IZA Journal of Labor Policy is a real journal — but the authors, the article, and even the domain the citation points at are all invented. The model reproduced the entire costume of scholarship around a source that has never existed. That is the mechanism in miniature: the shape of a citation, learned perfectly; the substance, absent.
And it is why “does this reference look credible?” is the wrong test. Every row below looks credible. The only test that works is resolution — and that test is mechanical.
| Model | The reference, as written | What is actually there |
|---|---|---|
| Phi-4 | Ryberg & Pichler (2021), “Remote Work and Productivity…”, IZA Journal of Labor Policy — izajolp.org/article/view/1344 | The domain does not exist. The journal is real; the authors, article and website are invented |
| Gemini Flash | Vanderbilt University (2023), “AI Detection Tools”, Center for Teaching — vanderbilt.edu/cet/ai-detection-tools/ | No such page. The centre is real; the cited guidance page is not |
| Grok 4.5 | nytimes.com/2023/05/18/technology/ai-chatbot-cheating.html | No such article. A plausible NYT URL for a story that was never written |
| Nova Pro | Gallup (2020), “State of the Global Workplace: 2020 Report” — gallup.com/workplace/326371/… | The report series is real; the cited URL resolves to nothing |
The review pass caught our own tool — twice
Short answer
Fifty citations initially failed to resolve. Hand-verification cleared eleven of them: four were real sources behind bot-walls, five were excluded as unverifiable, and two were correct citations that our own URL parser had truncated. We fixed the parser, re-ran the checks, and recounted before publishing.
The machine pass said 50 of 101 citations were dead — 49.5%. We did not publish that number, because a benchmark that accuses without verifying is doing the same thing the models are.
Re-fetching all fifty with a browser-identical client told a finer story. Four resolved immediately: real articles whose publishers block automated tools. Five more sat behind hard bot-walls at real academic publishers — plausible, unverifiable, excluded.
And two failures were ours. Certain scholarly URLs legally contain parentheses; our extractor treated the first parenthesis as the end of the URL, checked the truncated fragment, and reported it dead. Two models had cited real journal articles correctly and were about to be accused of fabricating them. We fixed the parser, re-resolved the citations, confirmed both live, and corrected the count — and the fix now ships in the public fact-check tool, which was silently mis-flagging the same URLs for real users.
That is what the review pass is for. It moved the headline from 49.5% to 38.6%, all eleven corrections in the models’ favour — and it found a real bug in the measuring instrument. A benchmark honest about its tool’s failures is the only kind entitled to publish the tool’s successes.
What this benchmark does not tell you
Short answer
It measures one narrow thing: whether references in short, prompted reports exist as cited. It is not a ranking of models, not a general hallucination rate, and not evidence about retrieval-backed tools that browse while they write.
The limits, plainly. Nineteen reports is a small corpus, and the intervals are wide — the true overall rate could plausibly be 30% or 48%. Two reports per model is far too few to rank vendors, so we deliberately publish no per-model table; the aggregate and the genre split are what the data supports.
The models were asked to cite from memory — no browsing, no retrieval. Tools that search while writing will fabricate less and mis-attribute more; that is a different benchmark, worth running. The two topics are two points on a spectrum of literature density, not a map of it. And “resolves” is not “true”: a real URL can host a weak source. Existence is the floor of citation quality, not the ceiling.
Finally, the five models that never produced usable reports — three refused under our no-training privacy setting, two erroring — are absent, not innocent, and one more completed only half its brief. The rate describes the reports we could check, on the day we checked them, with the versions we pinned.
What to do with an AI-written report
Short answer
Treat every reference as a claim to verify, not a source to trust. Check that the references resolve before you read a word of the argument — on a thin topic, more of them may be invented than real.
The practical readings of these numbers:
- Check citations first, argument second. A report whose evidence does not exist has answered your quality question early and cheaply.
- Be most suspicious on niche topics. Fabrication climbed past half exactly where you are least able to spot a fake source on sight.
- Plausibility is not a filter. Every fabricated reference looked real — right domain, right format, right tone. Resolution is the filter.
- Existence is only the first check. A resolving source still has to be the claimed source, and still has to support the sentence citing it.
- Automate the mechanical part. One reference is a minute; a report is an afternoon. This entire benchmark’s checking step was one tool call per report.
Frequently asked questions
How often does AI make up citations?
In this benchmark, 39 of 101 citations across 19 AI-written referenced reports did not exist as cited — 38.6%, with a 95% interval of 29.7% to 48.4%. The rate was 24.5% on a heavily published topic and 54.2% on a thinner one. Those numbers describe short, prompted, non-browsing reports on two topics in August 2026 — not a universal constant.
What does “fabricated as cited” mean?
The reference as the model wrote it — that URL, that identifier — does not exist: the domain does not resolve or the page returns a definitive not-found, confirmed by a second browser-identical fetch. A similar real article existing under a different address does not rescue an invented reference, because the citation as given sends anyone who checks it nowhere.
Why do models fabricate more on some topics?
Because a language model reproduces the shape of a citation far more reliably than the substance of one. On a dense literature the two coincide — there really are five canonical sources to cite. On a thin literature the model keeps the shape and invents the substance, which is why the fabrication rate doubled when the topic thinned out.
Which model fabricated the most?
We deliberately do not publish a ranking. With two reports per model, per-model intervals span most of the possible range — a table would be an unfalsifiable league standing dressed as measurement. What the data does support: every model that produced reports fabricated at least one citation somewhere, and none was immune on the thin topic.
Were any of the “fake” citations actually real?
Eleven of the fifty machine-flagged failures did not survive review: four were real articles behind bot-walls, five were unverifiable at hard-blocking publishers and excluded, and two were correct citations our own URL parser had truncated — a bug we fixed and now ship corrected in the public tool. Every published number is post-review, with all corrections in the models’ favour.
Does this mean AI-written reports can’t be trusted?
It means their references are claims to verify, not sources to trust — and that verification is mechanical. Notably, where citations were real, the claims they supported generally held up: none of the claims tested against resolving sources were contradicted. The failure concentrates in the sourcing, which is precisely the part you can check.
Would retrieval or web-browsing fix this?
It changes the failure rather than removing it. A model that browses while writing fabricates fewer nonexistent sources but mis-attributes more real ones — the source exists and does not say what the report claims. This benchmark measured memory-cited reports; a retrieval benchmark is a follow-up worth running.
How can I check the citations in a report I received?
Resolve each reference, confirm it is the source the document claims, and check that it supports the sentence citing it. Doing that by hand works for one reference; our fact-check tool automates all three steps for a whole report in one pass, and it is the same engine this benchmark used — including the parser fix the benchmark itself prompted.
Sources and further reading
- 01Stipple reference checker — the engine used for every check in this benchmark
- 02AI hallucinations guide — why fabricated citations are the checkable hallucination
- 03Wilson score interval — the method behind every confidence range on this page
Educational guidance, not a forensic certification. Detection technologies and standards change; review material decisions against current evidence.