Guide

Four payslip forgery types caught every time, one we gave up

Most document-fraud claims are a single accuracy number with no corpus behind them. This is the corpus, the six ways we forged it, the checks that fired, and the one whole family of forgeries we chose to let through.

By Stipple Research10 min readUpdated 5 August 2026
Key takeaways
  • Across 174 forgeries that a forensic check should catch, 138 were caught — 79.3%, with a 95% confidence interval of 72.7% to 84.7%.
  • Four of the six forgery families were caught without a single miss: broken net-pay arithmetic, wrong superannuation, reversed pay periods, and invalid employer identifiers.
  • One family — inconsistent year-to-date totals, 36 documents — was caught zero times. That is deliberate. The check still runs, but it no longer raises a fraud flag, because on real payslips it was far more often reporting a bad read than a forgery.
  • An honesty control matters as much as the detection rate. Eighteen documents were forged so that the arithmetic still balanced, and none of them were flagged — proof the checks are testing something specific rather than reacting to any edit.
  • The corpus is synthetic. A clean-document false-positive rate measured on generated payslips is the easy case, and we say so rather than quoting it as a headline.
Evidence path
  1. 01

    Build 210 payslips

    Start with the material.

  2. 02

    Forge them six ways

    Add one more signal.

  3. 03

    Run 17 forensic checks

    Add one more signal.

  4. 04

    Measure both directions

    Add one more signal.

  5. 05

    Publish the trade-off

    Make a careful call.

01

What we measured, and how

Short answer

We built 210 Australian payslips from three templates, left 18 untouched, and forged the other 192 in six specific, catalogued ways. Every document then went through the same 17 forensic checks, and we recorded which checks fired on which document.

Document-fraud tools tend to publish one number. "99% accurate." On what documents, forged how, with what left uncaught? The number cannot be checked, so it is not really a claim about anything.

So here is the corpus. Two hundred and ten Australian payslips, generated from three layout templates, all born-digital PDFs, all containing entirely fictional people and employers. Eighteen were left untampered as a control. The remaining 192 were forged, one field at a time, in six families chosen because they are the ways payslips are actually altered when somebody wants to overstate income.

Each forged document carries a manifest row naming the field that was changed, the original value, the altered value, and the specific check that should catch it. That last column is what makes this a test rather than a demonstration: a forgery only counts as detected when the check we said would catch it is the check that fires. A different signal firing for an unrelated reason is not a hit.

One deliberate exclusion. The engine separates risk signals from quality signals — the second kind report how much of a document could be read, not whether it was altered. Only risk signals count here. If a poorly-rendered forgery were caught because the engine could not read it properly, that would be a coverage artefact masquerading as forensics, and it would flatter the results.

02

The six ways we forged a payslip

Short answer

Broken net-pay arithmetic, an altered superannuation figure, reversed pay-period dates, an invalid employer identifier, an inconsistent year-to-date total, and — as a control — a forgery edited consistently enough that the arithmetic still balances.

Each family isolates one kind of lie. Together they cover the difference between a careless forgery and a careful one.

The last row is the important one. In that family we changed gross, net and year-to-date together, so every figure on the document still reconciles. Nothing arithmetic can catch it, and a detector that flags it anyway is not doing arithmetic — it is reacting to the fact that the file was edited at all. That control is how you tell the two apart.

FamilyDocumentsField alteredShould be caught by
Arithmetic break54Net payGross-to-net reconciliation
Superannuation rate54Super amountStatutory super check
Year-to-date mismatch36YTD netYTD reconciliation
Date order18Pay period start / endField consistency
Consistent tamper18Gross + net + YTD togetherNothing — must not flag
Identifier checksum12Employer ABNIdentifier checksum
03

The results, family by family

Short answer

Four families were caught without a single miss across 138 documents. One family was caught zero times out of 36. The honesty control was correctly left alone in all 18 cases, and none of the 18 untampered payslips were flagged.

Point estimates are misleading on corpora this size, so every rate below carries a 95% Wilson interval. For detection we report the lower bound, because you should only claim the recall you can demonstrate. For false alarms we report the upper bound, because you should be pessimistic about harm you might be causing.

Read the identifier-checksum row carefully. Twelve out of twelve is a perfect score, but with only twelve documents the honest lower bound is 75.8%. The measurement is thin, not the capability — but it is thin, and rounding that away would be the exact dishonesty this article exists to avoid.

FamilynDetected95% interval
Arithmetic break54100%lower bound 93.4%
Superannuation rate54100%lower bound 93.4%
Date order18100%lower bound 82.4%
Identifier checksum12100%lower bound 75.8%
Year-to-date mismatch360%upper bound 9.6%
Consistent tamper (must not flag)180 flaggedcorrect in ≥82.4%
Untampered control180 false alarmsupper bound 17.6%
All families that should be caught17479.3%72.7% – 84.7%
04

The family we gave up, and why

Short answer

Year-to-date reconciliation was moved off the fraud axis on purpose. On real payslips with a two-column "this pay / year to date" layout, a YTD mismatch was far more often a misread column than a forgery, so leaving it as a fraud signal meant accusing genuine documents.

Thirty-six documents in this corpus have a deliberately inconsistent year-to-date figure. Not one of them is flagged. On a scoreboard that is a 36-document hole, and it is the single biggest drag on the headline number — without it the aggregate would be 100% instead of 79.3%.

It is there because of what happened when the same check met real documents. Australian payslips very often print two columns side by side: the current pay period, and the year-to-date running total. When a document is read automatically, those columns are easy to confuse, and a confused read produces a year-to-date total that does not reconcile. The check fires. The document is genuine.

Measured against a corpus of real payslips, that one behaviour was responsible for a large share of all false accusations. Demoting it, together with a guard that detects the column-mixing at the source, cut the false-positive rate on that corpus from 78% to 30%.

So the trade was 36 synthetic catches for a very large reduction in real-world false accusations. We think that is obviously the right way round — a missed forgery costs you one bad application, while a false accusation costs a real person a job, a rental or a loan, and costs you the credibility to be believed the next time. But it is a trade, it is visible in the number, and we would rather explain it than quietly drop the family from the corpus.

The check still runs. It now reports as a coverage note — "the year-to-date column could not be fully verified" — which is what it was always actually measuring.

05

Why the control group matters more than the score

Short answer

Eighteen documents were forged so carefully that every figure still reconciles. None were flagged. Without that result, a high detection rate would be indistinguishable from a system that flags any edited file.

It is trivial to build a document checker that catches every forgery in this corpus. Flag everything. You would score 100% on all six families and the number would be worthless.

The consistent-tamper family is the guard against that. Those 18 payslips have had gross, net and year-to-date all altered in step, so the document is a lie but the arithmetic is sound. Every arithmetic check should pass, and every one did — zero of 18 were flagged.

That is a real limitation stated plainly: this engine does not catch a competent forgery of this kind by arithmetic alone. Catching it needs something outside the document — the disbursement actually credited to a bank account, or the employer confirming the figure. The engine has a disbursement check for exactly this, and it is why we describe results as coverage rather than clearance.

It is also why we publish the control at all. A detection rate without one tells you nothing about whether the system is discriminating or just alarmed.

06

How long the checks take

Short answer

All 210 documents were processed in 13.6 seconds — a median of 54.7 ms per document, with 95% completing inside 65.8 ms.

Forensic checks that take a minute per document cannot sit inside an application form. These are arithmetic, structural and file-level checks over an already-extracted document, so they are fast enough to run on every submission rather than on a sample.

The distribution is tight: a median of 54.7 ms, a 95th percentile of 65.8 ms, and a 99th of 87.8 ms. One outlier took 1.83 seconds — first-run initialisation, not a property of the document.

The extraction step that reads the document before these checks run is a separate and much slower stage. These figures cover the forensic pass only, and it would be misleading to quote them as end-to-end latency.

07

What this benchmark does not tell you

Short answer

The corpus is synthetic, Australian, born-digital, and built from three templates. Each of those is a real limit on how far the numbers generalise.

Taking the limitations one at a time, because they matter more than the headline.

The zero-false-alarm figure deserves particular scepticism. Eighteen untampered synthetic payslips are the easy case — they were generated by the same code that generated the forgeries, so they are internally perfect in a way real documents never are. Real payslips are scanned, re-saved, printed and re-photographed, carry inconsistent fonts from payroll software, and reconcile imperfectly for legitimate reasons like salary sacrifice and mid-period rate changes. The genuinely hard false-positive work happened on a separate corpus of real documents, and it is the reason the year-to-date family was demoted at all.

LimitWhat it means
Synthetic corpusGenerated, not collected. No scanner noise, no re-saves, no payroll-software quirks.
Australian payslips onlyThe super and identifier checks are jurisdiction-specific and do not transfer.
Born-digital onlyNo scanned or photographed documents, where extraction is the dominant error source.
Three templatesLayout diversity is far below what a real intake queue receives.
One field per forgeryReal forgeries are often multi-field, which is why the consistent-tamper control exists.
Coverage, not clearanceA clean result means these checks found nothing, not that the document is genuine.
08

What to do with a payslip you are unsure about

Short answer

Check the arithmetic first, because it is cheap and it is where careless forgeries fail. Then get one figure confirmed from outside the document, because that is the only thing a careful forgery cannot survive.

The practical lesson from this corpus is that forgeries fall into two groups, and they need different responses.

Careless forgeries change one number and leave the rest of the document to contradict it. Every one of those was caught here, without exception, in under a tenth of a second. Gross minus deductions should equal net. Superannuation should be the statutory percentage of ordinary earnings. The pay period should start before it ends. The employer identifier should pass its checksum. Any of those failing is not a judgement call.

Careful forgeries adjust every figure in step, and no amount of arithmetic will find them — our own control group proves it. For those, you need one fact from outside the paper: the amount actually credited to the bank account, a super fund contribution record, or direct employer confirmation. One external anchor beats any number of internal checks.

And treat a clean forensic result as what it is. It means these specific checks found nothing, on this document, in this format. That is coverage. It is not a certificate.

Questions

Frequently asked questions

How do you detect a fake payslip?

Start with the internal arithmetic, because careless forgeries fail there. Gross minus deductions must equal net, superannuation must match the statutory percentage of ordinary earnings, the pay period must start before it ends, and the employer identifier must pass its checksum. On this benchmark every forgery of those four kinds was caught, 138 out of 138. A forgery that adjusts every figure in step will pass all of them, which is why anything consequential needs one figure confirmed from outside the document.

What is the accuracy of payslip fraud detection?

On this corpus, 138 of 174 forgeries that a forensic check should catch were caught — 79.3%, with a 95% confidence interval of 72.7% to 84.7%. That single number hides the shape of the result, though: four of the six forgery families were caught without a single miss, and the entire shortfall is one family we deliberately stopped flagging.

Why does the benchmark miss an entire category of forgery?

Because catching it was costing more than it was worth. Year-to-date reconciliation looks decisive on synthetic documents, but on real Australian payslips, which usually print current-period and year-to-date figures in adjacent columns, a mismatch was far more often a misread column than a forgery. Leaving it as a fraud signal meant flagging genuine documents at a high rate. Demoting it cut false positives on a real-document corpus from 78% to 30%, at the cost of the 36 synthetic catches you can see in the results table.

Can a payslip forgery get past these checks?

Yes, and we built 18 of them to prove it. If gross, net and year-to-date are all altered consistently, every arithmetic check passes because the document genuinely reconciles — it is just false. None of those 18 were flagged, correctly. Detecting that kind of forgery needs evidence from outside the document, such as the amount actually credited to a bank account or confirmation from the employer.

How many false positives does it produce on genuine payslips?

None of the 18 untampered documents in this corpus were flagged, which puts the upper bound of the interval at 17.6%. Treat that figure with caution: the corpus is synthetic, so the clean documents are internally perfect in a way real payslips never are. The meaningful false-positive work was done separately against real documents, and it is what drove the year-to-date demotion described above.

Does this work on payslips from outside Australia?

The structural checks do — arithmetic reconciliation, date ordering, file and layout consistency are not jurisdiction-specific. The superannuation check and the employer identifier checksum are specifically Australian and do not transfer. This benchmark measured Australian payslips only, so nothing here should be read as a claim about other countries.

Is this fast enough to run on every application?

The forensic pass took a median of 54.7 milliseconds per document, with 95% finishing inside 65.8 milliseconds — 210 documents in 13.6 seconds. That is fast enough to run on every submission rather than a sample. Note that reading the document beforehand is a separate and slower stage, so this is not an end-to-end figure.

Is the corpus published?

No, and one part of it never will be. A public library of 192 realistic forged Australian payslips would be more useful to the people making them than to the people checking them, so those files stay internal. What we have published instead is the method in full: how many documents, how each was altered, which check was expected to catch it, and what happened — including the family that was caught zero times. That is enough to argue with the result, which is the point of publishing it.

Sources

Sources and further reading

  1. 01ATO — the superannuation guarantee rate the super check is measured against
  2. 02Australian Business Register — the ABN checksum algorithm used by the identifier check
  3. 03Wilson score interval — the method behind every confidence range on this page
  4. 04Stipple document verification — the engine under test

Educational guidance, not a forensic certification. Detection technologies and standards change; review material decisions against current evidence.

Run your own document through these checks

Upload a payslip and see exactly which forensic signals fired, which abstained, and what each one is standing on — the same engine, the same signals, the same coverage-not-clearance limits stated in every result.

Verify a document