Why looking at the document does not work
Short answer
A modern forgery is not a bad photocopy. It is a real template with edited numbers, and it looks correct because it is correct everywhere except the values someone changed.
Editing a payslip used to require skill. It now requires a text editor and a few minutes, and the output has consistent fonts, plausible spacing and clean metadata at a glance. Manual review still catches the careless fakes — a misaligned logo, a font that changes mid-number — and the careful ones go straight through, because there is nothing to see.
The reason is structural rather than a failure of attention. A reviewer checks whether a document LOOKS like the documents they have seen. A forger optimises for exactly that. The two are playing the same game, and the forger gets to move second.
A forensic check plays a different game. It does not ask whether the document looks right; it asks whether the document is internally consistent with itself and with the world outside it. Those are questions with answers, and the answers do not depend on how convincing the layout is.
The three things a check inspects
Short answer
Arithmetic, identifiers, and the file structure underneath the picture. None of them are visible to a reader, and all three are expensive for a forger to get right at once.
ARITHMETIC. Gross minus tax minus deductions should equal net pay. Year-to-date figures should be consistent with the period figures. A superannuation contribution should be the statutory percentage of ordinary time earnings. Each of those is a calculation, and a calculation either balances or it does not — there is no judgement to argue with, and no amount of visual polish changes the result.
IDENTIFIERS. Australian company and business numbers carry check digits: an ABN verifies against a weighted modulus-89 sum, an ACN against its own scheme. A number invented to look plausible fails that arithmetic. This is the cheapest check in the set and one of the most decisive, because a forger has to either use a real entity’s identifier — which is traceable — or produce one that verifies, which requires knowing the algorithm.
STRUCTURE. A PDF is not a picture; it is a program describing how to draw a page. That program records things the rendered image does not show: whether text was added in a later incremental update, whether a value is drawn in a different font from its neighbours, whether an annotation is painted over the top of the original figure, whether the file was produced by a word processor or by a payroll system. A screenshot of a payslip destroys all of it, which is itself worth knowing.
The 43 signals
Short answer
Thirteen categories, from PDF structure to visual tamper analysis to domain arithmetic. Each one states what it tested and what it found, so a result can be argued with rather than trusted.
The count matters less than the spread. A check that only reads metadata is defeated by stripping metadata; one that only recomputes arithmetic is defeated by a forger with a calculator. The families below fail independently, which is what makes defeating all of them at once expensive.
Every signal reports its own evidence. "Superannuation contribution deviates from the statutory minimum" arrives with the claimed figure, the expected figure, and the rate it was measured against — not as a score to be taken on trust.
| Family | Checks | What it inspects |
|---|---|---|
| PDF structure | 7 | Incremental updates, annotation overlays, embedded files, signature integrity, active content, hidden text, text/render mismatch |
| Visual forensics | 7 | Compression inconsistency, error-level analysis, spectral smoothness, JPEG grid analysis, render and paint order, copy-move cloning |
| Domain arithmetic | 6 | Pay reconciliation, superannuation, year-to-date totals, disbursement matching, field consistency, identifier checksums |
| Transform history | 5 | Screenshot or scan detection, mobile capture, messaging-app capture, screen recording |
| Content consistency | 5 | Amount words against figures, table arithmetic, date anomalies, label integrity, font discontinuity inside a value |
| File integrity | 3 | Header consistency, hash availability, size plausibility |
| Provenance | 3 | Modification history, generation stack, HTML-to-PDF render origin |
| Metadata | 2 | AI-generation provenance in documents and in images |
| Reading quality | 5 | Document-type classification, extraction completeness, grid anomalies, low-confidence OCR regions, whole-document review |
Seventeen checks that never raise a fraud flag
Short answer
Twenty-six of the 43 checks can contribute to a risk finding. The other seventeen run, report exactly what they found, and are structurally unable to flag fraud — because on real documents they fire on honest artifacts.
Every check in the set was run against genuine documents as well as forged ones. Some of them turned out to fire on the genuine ones: hidden text spans, which are a normal by-product of rendering a web page to PDF and of OCR text layers. Screenshots, which mean somebody photographed their payslip rather than that they edited it. Missing text layers, which are what a scan legitimately looks like.
A check like that is not a detection capability. It is a false-accusation generator with good intentions, and the usual response — leave it in, weight it down a little, hope the score absorbs it — just spreads the error thinner. So those checks were moved onto a second axis instead. They still run. They still report what they saw. They cannot raise a fraud flag, at any weight, ever.
Ten of the seventeen carry no severity value at all. The remaining seven kept theirs, which looks odd until you see what the axis does: severity describes how notable a finding is, and the axis decides whether it is allowed anywhere near an authenticity verdict. Hidden text is worth telling you about. It is not evidence against anybody.
The practical consequence is that a low-risk result is not automatically a good result. A document can be low risk because nothing suspicious was found, or low risk because very little could be tested — and those are indistinguishable in a single score. Which is why coverage travels with every result rather than sitting in a footnote.
What a demoted check looks like
Short answer
One check catches a real forgery family and was still stopped from raising fraud flags — because on genuine documents it was more often reporting a bad read than a forgery.
Year-to-date reconciliation compares period figures against the running totals. It is a legitimate check and it catches a real class of forgery. In the benchmark it caught that family zero times, and the reason is instructive: on genuine payslips it fired often, because YTD figures legitimately diverge — a mid-year start, a corrected prior period, a payroll migration, a document that simply prints the numbers differently.
A check that fires on honest documents more often than dishonest ones is not a detection capability. It is a source of false accusations wearing one. So it was demoted: it still runs, it still reports what it found, and it no longer contributes a fraud flag.
That decision costs measured detection rate, and it is published rather than hidden, because the alternative is a headline number bought with other people’s false accusations. A checker that never demotes anything is a checker that has never looked closely at its own false positives.
What it cannot do
Short answer
It cannot prove a document is genuine, it cannot catch a forgery that is internally consistent, and it cannot see anything outside the file.
THE CONSISTENT FORGERY PASSES. A forger who adjusts every figure so the arithmetic still balances, uses a real employer’s valid identifier, and produces the file through a genuine tool has defeated the arithmetic, the checksums and the structure at once. The benchmark includes eighteen documents built exactly that way, and none of them were flagged — which is reported as a control rather than buried, because it is the honest boundary of the method.
ABSENCE IS NOT PROOF. A clean result means these checks found nothing in the portion of the document they could read. It is not a statement that the document is genuine, and any workflow that treats it as one has built its confidence on the wrong side of the evidence.
THE FILE IS THE WHOLE WORLD. The check cannot ring the employer, cannot see the bank account the money went into, and cannot know the person exists. That is precisely why a consequential decision should confirm one figure from outside the document — a bank credit, an employer identifier, a direct call. One external confirmation defeats the entire class of internally-consistent forgeries that no file-level check can reach.
STATUTORY RATES CHANGE, AND OLD DOCUMENTS ARE NOT WRONG. The superannuation check holds the legislated schedule back to 2002 and selects the rate in force for the pay period, so a correct 2024–25 payslip showing 11.5% is not flagged for failing to be 12%. A checker that compares every document against today’s rate generates a false positive for every historical document it sees.
How to use a result
Short answer
As the start of a review, with the evidence attached — not as a decision. The result is built to be argued with.
Read the signals, not the band. The band is a summary; the signals are the finding. "Net pay does not reconcile: claimed $4,210.00, computed $3,980.00" is something a reviewer can check in thirty seconds and a customer can explain. A risk score is neither.
Read the coverage before the conclusion. If the document was a photograph, the quality axis will say so, and the right response is to ask for the original PDF rather than to act on a thin read.
Confirm one figure externally when the decision matters. Not every figure — one. It is the cheapest defence against the one forgery family the method cannot reach, and it takes a phone call.
And treat a flag as a question, never an accusation. Someone wrongly suspected of forging a payslip has a real grievance, and the asymmetry is deliberate: the cost of a missed forgery is money, and the cost of a false accusation is a person. The result is designed around that difference.
Frequently asked questions
What does document fraud detection actually check?
Three things a reader cannot see: whether the arithmetic reconciles (gross minus tax minus deductions against net pay, period figures against year-to-date), whether identifiers pass their own checksums, and what the file structure records about how the document was produced and edited. Forty-three checks in total, each reporting the evidence it stood on.
Can it tell me a payslip is fake?
No. It can tell you that specific checks failed, and show you what each one tested. "Net pay does not reconcile" is a fact about the document; "this payslip is fake" is a conclusion about a person, and the second is not ours to draw. Treat a flag as a reason to look closer.
Does a clean result mean the document is genuine?
No, and the result says so. A clean result means these checks found nothing in the portion of the document they could read. Coverage travels with every result for exactly this reason — a document that could barely be read produces a clean result for the wrong reason.
What kind of forgery gets through?
One that is internally consistent. If every figure is adjusted so the arithmetic still balances, the employer identifier is real and valid, and the file was produced by genuine software, there is nothing left in the file to catch. That is why a consequential decision should confirm one figure from outside the document.
Does a photo or scan of a payslip still work?
It is read, and the result reports lower coverage. Rasterising a document destroys the structural evidence a born-digital PDF carries, so the checks that inspect the file’s internals have nothing to work with. The quality signals report that honestly rather than returning a confident verdict on a thin read.
Why do seventeen checks not affect the risk score?
Because they answer a different question. "We could not read this well" and "this has been altered" are separate findings, and a single score collapses them into something that cannot be acted on. Seventeen checks report reading quality; eleven of those carry no severity at all and can never produce a fraud flag.
Which superannuation rate is a payslip checked against?
The rate in force for its own pay period. The engine holds the legislated schedule back to 2002 and selects the last step on or before the pay date, so a 2024–25 payslip is measured at 11.5% and a current one at 12%. Checking every document against today’s rate would flag every historical payslip ever submitted.
Sources and further reading
- 01Payslip forgery benchmark — 210 documents, six families, with confidence intervals
- 02Developer playbook — the checks, the schemas, and a worked response
- 03ATO — Super guarantee percentage, the legislated schedule
- 04AI-generated evidence has reached AFCA — 771 determinations reviewed
Educational guidance, not a forensic certification. Detection technologies and standards change; review material decisions against current evidence.