Research

How accurate is bank statement categorisation — and how we measure it

A statement reader is judged on two numbers, not one. How we measure ours, what it reads today, and what holds it back.

· 5 min read

General information, not legal or financial advice.

Two numbers, not one

Every figure a lender takes from a bank statement — income, rent, other lenders, gambling — starts as a category on a transaction. A reader can fail in two ways: it can put a line in the wrong category, or leave it unread. So we publish two numbers.

  • Precision: of the lines it put in a category, the share that were right.
  • Coverage: the share of lines it put in any category. It declines the rest, and says so.

On the 963 transactions its rules never saw, measured on 4 October 2026, our statement reader is right 99.1% of the time it answers, and answers 76.6% of lines (published figures).

One “accuracy” figure hides the choice between the two. A reader that answers only the easy lines can be right every time and read half a statement. One that answers everything reads it all, and more of it wrongly.

Why we decline rather than guess

A declined line is visible: it is marked unread, for a person to check. A wrong one is not. A repayment to another lender read as shopping hides a commitment; a pay credit read as a transfer understates income. In a credit file, an honest gap costs less than a confident error.

Every category at 95%, not the average

Overall precision can look excellent while one small category is poor, because a small category barely moves the average. A lender adding up one category — rent, other lenders, buy now pay later — relies on that category, not on the mean.

So each category is held to 95% precision on its own. A category is judged once it has given 20 answers on a test set. Below that, one wrong answer moves it by more than the margin, so we show a count, such as “9 of 9”, not a percentage. Today all 8 categories with that many answers on the held-out set meet the floor; the weakest is shopping, at 96.3% (by category).

Strict, balanced, broad

How far to trust the reader is the lender’s choice, in its settings or on each call, and every answer says which setting gave it (each setting, measured).

  • Strict: the rules alone. Every answer is a name or a kind of business a rule read.
  • Balanced, the default: the rules, and a model’s reading where it keeps its category at 95% precision or better.
  • Broad: the model may answer where its category stays at 90% or better.
Precision and coverage by setting, by line and by household
SettingPrecisionCoverage
Strict99.4%97.7% by household73.6%81.5% by household
Balanced default99.1%97.7% by household76.6%83.1% by household
Broad99.0%97.7% by household80.1%84.9% by household
On 963 transactions the rules never saw, measured 4 October 2026. By household, each of the 14 source datasets counts as at most 20 lines. Upper bar by line, lower by household; the mark is 90% coverage. Every category and release →

Broad reads 80.1% of lines for a small loss of precision. The model’s own answers are where the risk sits: 26 of 29 were right under the default, and 58 of 62 under broad (by setting).

Held out, not written from

A figure measured on the lines the rules were written from says nothing about the next statement. So the labelled public datasets are split in two by a hash of each description, which keeps the same wording on one side only. Rules are written, and the model tuned, without ever reading the held-out half. It is measured once, at the end, and nothing is changed after.

The split does its job. On 4 October 2026, changes that lifted the development half left the held-out half where it was, and we counted them as no gain.

A second independent set checks names rather than statement lines: 1,143 Australian businesses from public registers and map data, each with a known answer. The default reads 60.6% of them at 100.0% precision (test sets). Sets the rules were written from, such as the banks’ own wording, are published too, but only to show nothing was lost. They do not measure accuracy.

The limits

The test set is narrow. 779 of the 963 held-out lines come from one bank’s synthetic test account: one invented household, at the same shops again and again. Spending dominates it. The lines a lender cares about most are thin: 6 salary lines, 4 rent and 3 buy now pay later, which is why those categories show counts (test sets).

So we also count by household. Beside every figure counted by line, we publish the same figure with each household weighing at most 20 lines; in the public datasets, each of the 14 sources counts as one. The synthetic account is then 16% of the figure, not 81%. Counted that way, the default is right 97.7% of the time and answers 83.1% of lines (both views). Precision is lower by household and coverage higher. Neither view replaces the other.

Coverage is short of 90%. To read 90% of the held-out set, the default would need to read 129 more lines, and broad 96 more. Coverage has risen from 53.4% at the first release on 19 September 2026 to 76.6% today, with precision held (every release). What is left is harder:

  • the synthetic account’s own shops, which we will not add by name: the score would rise and no real statement would be read better;
  • the model’s thresholds, set on business names from maps rather than on statement lines, so at the default it gives no answer at all in some spending categories;
  • shops known only by a name, with no word such as café or pharmacy to say what they are;
  • placeholder lines in synthetic data, which should not be read at all.

What would lift it: labelled lines from many real households; a few thousand statement lines, kept apart from every test set, to set the model’s thresholds on; and a list of the payees that appear most often on real statements, each checked against a public register. The way to the first is built: with a lender’s agreement, its statements can be kept de-identified and labelled, split by household so no account is both written from and measured on, and published as a test set once 30 households are held out (how it is measured). None is published yet.

What to ask any provider

  • Precision and coverage, both, for each category, not one average.
  • Measured on what: lines the rules were written from, or lines held out?
  • Who labelled the test set, and how many households are in it?
  • What happens to a line it cannot read?

Sources

  1. How well it reads: precision, coverage and F1 for every category, setting, test set and release, measured 4 October 2026. The same figures for a program: /v1/statement/accuracy.
  2. Creditcrest’s coverage measurement of 4 October 2026, on what stands between today’s coverage and 90%; and the measurement script’s rules for a lenders’ test set, split by household and published from 30 held-out households.