What we scanned, and how
Short answer
All 771 AFCA determinations published between 1 June and 23 August 2026, fetched in full text and searched for more than 50 generative-AI phrases — product and model names, “AI-generated”, “AI-drafted”, “robo-advice”, “chatbot”, “language model” and similar — with three rounds of false-positive filtering. One determination matched.
The Australian Financial Complaints Authority publishes its determinations, full-text and free. That makes it one of the few places where the arrival of AI-generated evidence in financial disputes can actually be counted rather than anecdoted. So we counted.
Method, in full: every determination published in the window (771 documents) was fetched whole. We searched for an expansive phrase list — over 50 terms covering AI product names, model vocabulary, and the phrasings a decision-maker would use (“AI-generated”, “AI-drafted”, “machine-generated”, “chatbot”, “robo-advice”). Matching was exact-string; three filtering passes removed regex artifacts where ordinary legal language (“produced”, “generated by the bank’s system”) collided with the net.
The result: one determination — 0.13% of the window — engages with AI-generated documents as evidence. We are not aware of a prior published count of this kind for AFCA; if one exists, we would genuinely like to compare notes.
The one case: AI transcripts, spotted and refused
Short answer
Case 12-25-282452 (Westpac Banking Corporation, deposit taking, determined 2 June 2026): a complainant submitted what he described as written notes and memory-based summaries of phone calls with the bank. The ombudsman found they appeared to be AI-generated transcripts, could not be satisfied the underlying recordings were lawfully made, and declined to rely on them. The determination went in Westpac’s favour.
The complainant — detained in an immigration detention centre at the time — submitted documents he described as “written notes and memory-based summaries” of phone calls with Westpac. The ombudsman looked closer, and the determination records exactly what gave the documents away. It is worth quoting, because it is a working checklist of how machine-generated documents betray themselves:
“I have reviewed these documents, and they appear to be AI-generated. This is because the documents: include the words ‘transcribed’ by ‘AI Minutes’ with a specific date. For example, one of the documents stated ‘Notes created on October 22 2025 at 4.15 by AI Minutes’ … often include random percentage figures after each sentence, which I understand is often done by AI transcription services as an indication of the probability of accuracy in its transcription. For example, ‘speaker 2 • 0:09 - 0:11 • 94%’ … include the lyrics of songs, which appears to be music which was played when the call was on hold”.
Three tells, all artifacts of the tool rather than the content: the product left its name and timestamp in the output; it left its per-utterance confidence markers in; and it faithfully transcribed the hold music. In the ombudsman’s assessment, the “memory-based summaries” were likely machine transcripts of recorded calls — recordings AFCA could not be satisfied were lawfully made.
The refusal then rested on more than authenticity. In the ombudsman’s words: “My view is these documents are likely transcripts of phone calls recorded by the complainant with the bank. I cannot be satisfied these documents have been made in accordance with relevant Commonwealth and State laws which require consent prior to recording calls.” And: “Consequently, it would not be appropriate for AFCA to rely on such transcripts as we expect parties to act in good faith”.
What the case actually demonstrates
Three things, each more durable than the headline.
First: the direction of the problem. The AI-generated documents came from the consumer, not the firm. Most institutional worry about AI evidence runs the other way — firms worrying about their own disclosure obligations. The one case surfaced in our window is a reminder that every party to a dispute now has a transcript generator in their pocket.
Second: the detection was artifact-dependent. The ombudsman caught these documents because the tool signed its work — name, timestamp, confidence percentages, hold-music lyrics. Strip those artifacts (a copy-paste into a clean document does it) and the same submission reads as handwritten notes. Human review caught documents that announced their own origin; it has no comparable method for a submission whose artifacts have been cleaned before it arrives.
Third: the ruling gives future parties a template. The refusal chained authenticity (are these what they claim to be?), legality (could they have been lawfully made?), and good faith. That three-link chain — not a technical authenticity finding alone — is what sank the evidence. Practitioners advising on AI-generated material in disputes will find the structure reusable.
What one-in-771 does and does not mean
Short answer
It is a floor, measured honestly: explicit engagements with AI-generated evidence in published determinations over three months. It is not a measurement of how much AI-generated material entered disputes undetected or unremarked — which is precisely the quantity nobody can currently count.
The temptation with a number like this is to inflate it into a trend. We decline. One case in 771 determinations says AI-generated evidence has arrived at AFCA and is, so far, rarely surfaced in written reasons. It does not say how often it arrives undetected — by construction, a successful fake generates no mention.
The limitations, stated plainly: the window is three months, not two years; the corpus is published determinations, which follow the underlying complaints by months; exact-string matching finds explicit mentions, not paraphrased discussion; and a determination only records AI involvement when a party or the decision-maker raised it. Every one of those biases pushes the count down. 0.13% is the visible tip; the size of what is below it is the open question.
That is also why this page will be re-run. The measurement is repeatable by design — same corpus source, same phrase net, same filtering — and the interesting number is the delta. When the count moves from 1, that movement will be the story.
The gap the case exposes — and our stake in it
Disclosure: we build document-verification tooling, so we have an obvious interest in this gap. Judge the reasoning accordingly.
The ombudsman’s three tells are metadata and artifact analysis, done by an attentive human. That works when the tool labels its output. The harder cases — the ones this count cannot see — are documents where the artifacts were cleaned: a transcript pasted into a fresh file, a generated statement re-typed, a fabricated exhibit exported through a converter that strips the fingerprints. Catching those requires reading the document the way software can and people cannot: internal arithmetic that fails to reconcile, formatting inconsistent with the claimed source, structural traces of generation and editing.
The honest version of our claim is modest: forensic tooling automates and extends the checklist this determination performed by eye, and it produces its evidence in a form a decision-maker can quote — which, as this case shows, is what a refusal ultimately has to stand on.
Frequently asked questions
How many AFCA determinations mention AI-generated documents?
In the window we measured — all 771 determinations published between 1 June and 23 August 2026, full text — exactly one engages with AI-generated documents submitted as evidence: case 12-25-282452, involving Westpac. That is 0.13% of published determinations in the period.
How did the ombudsman detect the AI-generated documents?
Three artifacts, quoted in the determination: the transcription tool’s own name and timestamp embedded in the text (“Notes created … by AI Minutes”), per-utterance confidence percentages (“speaker 2 • 0:09 - 0:11 • 94%”), and hold-music song lyrics faithfully transcribed into the record.
Did AFCA accept the AI-generated evidence?
No. The ombudsman refused to rely on the documents — both because they appeared to be transcripts of calls recorded without the consent required by Commonwealth and State law, and because relying on them would sit poorly with AFCA’s expectation that parties act in good faith. The determination went in the financial firm’s favour.
Does this mean AI-generated evidence is rare in financial disputes?
It means explicit, surfaced engagement with it is rare in published AFCA determinations so far. Undetected material generates no mention by definition, published determinations lag the underlying complaints, and our method counts exact phrases. The honest reading is a floor: it has arrived, and the visible count is small.
Will this count be updated?
Yes — the scan is repeatable by design (same source, same phrase net, same filtering), and the movement of the count is the durable story. We intend to re-run it as new determinations publish.
Can Stipple detect documents like the ones in this case?
The artifacts the ombudsman spotted — tool watermarks, confidence markers, structural oddities — are the class of signal our document verifier checks mechanically, alongside things human review cannot do at speed, like recomputing a document’s internal arithmetic. Disclosure: that is our product, and this page is our research. The determination itself is linked below so you can read the primary source.
Sources and further reading
- 01AFCA determination 12-25-282452 (Westpac Banking Corporation, 2 June 2026) — the primary source
- 02AFCA — published decisions search (the corpus)
- 03AFCA — Datacube complaint statistics
Educational guidance, not a forensic certification. Detection technologies and standards change; review material decisions against current evidence.