[ Benchmarks ]

Bank Statement Analysis Accuracy: What 200,001 Transactions Revealed

A head-to-head categorisation test across 200,001 transactions, same raw statements, two tools. 92.8% versus 80.3%, and a breakdown of exactly where the errors fall.

Zeus Dhanbhoora

8 min read · 20 August 2026

Bank Statement Analysis Accuracy: What 200,001 Transactions Revealed
Contents
  1. The problem with "accuracy" in this market
  2. Test design
  3. Results
  4. Reading the result correctly
  5. Where the errors actually fall
  6. What both tools got wrong
  7. What the difference costs
  8. Run this test yourself
  9. Limitations
  10. Frequently asked questions

Every bank statement analysis vendor in India claims high accuracy. None of them publish what they measured, on what data, against what baseline, or scored by whom. The number is an assertion, and assertions are not comparable to each other.

We ran the test instead. Same raw statements, two tools, 200,001 transactions, every disagreement adjudicated. This post gives the methodology, the full result matrix, the transaction-level examples, and the residual errors we did not get right — so you can reproduce it on your own files rather than take our word for any of it.

The problem with "accuracy" in this market

Three different things get sold under one word.

Extraction accuracy asks whether the tool read the page correctly — did ₹88,201 come through as ₹88,201. On native PDFs this is close to solved. Any competent vendor is above 99%, which is why everyone quotes it.

Categorisation accuracy asks whether the transaction was assigned to the right class. This is where tools actually diverge, and almost nobody publishes it.

Credit-signal accuracy asks whether the categorisation preserved the information an underwriter needs. A transaction labelled "Transfer from ABC Enterprises" is not wrong. It is simply void. It passes a naive correctness check and tells the credit team nothing.

The gap between the second and third is the entire subject of this post. Categorisation errors are not randomly distributed across a statement. They concentrate in precisely the transaction types that carry credit signal — borrowings, failed payments, and income of ambiguous quality. A tool can be 80% accurate overall and wrong on most of the transactions that determine the decision.

Test design

Dataset. 200,001 transactions drawn from live credit cases across salaried, self-employed and SME borrower profiles, spanning multiple banks and both native and scanned statement formats.

Method. Identical raw statements were processed independently by Fiscus and by an established bank statement analyser already in production use at partner institutions. Neither tool saw the other's output.

Scoring. Every transaction was assigned one of four outcomes: both tools correct, Fiscus correct alone, the other tool correct alone, or both incorrect. Disagreements were adjudicated against the underlying narration and account context by credit review teams at the partner banks and lenders whose cases these were — the same analysts who would have reclassified the output manually in the ordinary course. Their conclusions matched ours.

What this is and is not. This is a vendor-run study, jointly adjudicated with the institutions whose data it used. It is not a third-party audit. We have set out the full protocol in the last section so you can run it against your own current provider, with your own analysts scoring, on your own files.

Results

OutcomeTransactionsShare
Both correct153,79976.9%
Fiscus correct, other tool wrong31,84015.9%
Other tool correct, Fiscus wrong6,7393.4%
Both incorrect7,6233.8%
Total200,001100.0%

Aggregate accuracy:

  • Fiscus: 92.8% (185,639 of 200,001)
  • Other tool: 80.3% (160,538 of 200,001)

Reading the result correctly

The headline gap is 12.5 percentage points, which sounds incremental. It is the wrong way to read it.

Invert to error rates. The other tool misclassified 39,463 transactions. Fiscus misclassified 14,362. That is 64% fewer errors — the same result stated in the unit that determines your analyst's workload, because analysts spend their time on the errors, not on the correct classifications.

Per 1,000 transactions: 197 errors versus 72. On a 24-month multi-account SME case running around 900 transactions, that is roughly 177 items needing review against 65.

The second number worth extracting is the disagreement set. The two tools disagreed on 38,579 transactions, 19.3% of the sample. Within that set, Fiscus was correct 82.5% of the time. Where the tools diverge, the divergence is not symmetrical noise. It runs overwhelmingly one way.

Where the errors actually fall

The 31,840 transactions Fiscus classified correctly and the other tool did not are not spread evenly. They cluster into four families, and each one distorts underwriting in a specific, predictable direction.

1. Borrowing recorded as revenue

A large NEFT credit from a corporate counterparty was labelled by the other tool as a transfer from that party. It was a loan disbursal.

This is the single most expensive error class in business lending, because it is wrong twice. The credit inflates apparent revenue, so the borrower looks larger than it is. And the obligation behind it never enters the picture, so leverage looks lower than it is. DSCR moves in the wrong direction on both the numerator and the denominator.

The same pattern appears with bill discounting. A narration prefixed 177ILBD — BILLS DISCOUNTED is a drawdown against receivables. Read as a generic inflow, it becomes revenue. It is borrowing.

2. Failed payments read as successful ones

ACHRETCHG is an ACH return charge. It is the bank debiting the borrower a penalty because a mandate bounced.

The other tool classified it as a loan repayment.

An EMI failure was recorded as an EMI success. This does not degrade the conduct assessment — it inverts it. A borrower with a bounce history reads as a borrower with a clean one, and the technical-versus-non-technical distinction that credit policy actually turns on never gets made.

3. Income quality collapsed into a single bucket

Three credits, all correctly identified as inflows by both tools, all treated as equivalent by one of them:

  • A credit under a state government DBT scheme, coded APBCR — a transfer benefit, not earned income, and not repeatable.
  • A payment gateway settlement, NEFT — [PG operator] settlement — genuine merchant revenue, the most creditworthy inflow in the account.
  • An IMPS credit from a fantasy gaming platform — winnings.

A subsidy, a business receipt and a gambling payout are not the same input to an income assessment. Collapsing them to "transfer" or "other credit" removes the distinction the underwriter is paid to make.

4. The named-counterparty problem

Across a large share of the disagreement set, the other tool's output was of the form "Transfer to [string from narration]" — the narration echoed back with a category prefix.

Nothing there is false. Nothing there is usable either. You cannot compute buyer concentration, you cannot identify related-party circularity, and you cannot separate a supplier payment from a promoter withdrawal from a tax remittance. Examples from the sample: a cheque to a road transport company classified as a cheque withdrawal rather than logistics spend; a CGST charge classified as income tax; processing, CIBIL and CERSAI fees classified as generic bank charges rather than borrowing costs; a quick-commerce purchase classified as a utility rather than groceries.

Each is individually trivial. In aggregate they are the difference between a categorised statement and a credit picture.

What both tools got wrong

7,623 transactions — 3.8% — defeated both systems. A representative failure: a ₹5,000 debit narrated only as DOCUMENTATION C. One tool called it "others", the other called it a transfer. Neither was right, because the narration does not contain enough information to be right.

There is an irreducible floor here. Some Indian bank narrations are truncated to a length that carries no recoverable signal. Any vendor claiming to be near it should be asked how they handle a narration with nothing in it. The honest answer is that the transaction gets flagged for review rather than confidently mislabelled — a low-confidence flag is a better output than a wrong category asserted at full confidence.

What the difference costs

Two things, and the second is worth more than the first.

Analyst time. At 30 to 45 minutes of reclassification on a typical multi-account business case, the error volume is the workload. Cutting misclassifications by 64% removes most of the manual pass, which is where the reduction in per-case analyst time comes from.

Decision quality. This is the one that does not show up in an efficiency calculation. A reclassification pass catches errors an analyst notices. It does not catch a loan disbursal sitting quietly in the revenue line, because there is no visual anomaly to notice — the number is correct and the label is plausible. Those errors survive review and go into the credit appraisal memo intact.

Run this test yourself

The methodology is not proprietary. It takes about a day.

  1. Pull 20 to 30 closed cases your team has already analysed manually, weighted toward the borrower segments where your current tool underperforms. Include scanned statements, multi-account cases and at least a few cooperative or small finance bank formats.
  2. Run both tools on the identical raw files. No pre-processing, no format hints, no cherry-picking. Whatever your team receives in production is what goes in.
  3. Score only the disagreements. Where both tools agree, adjudication adds nothing. Extract the disagreement set — it will be roughly a fifth of transactions — and have your own credit analysts, not either vendor, assign the correct category.
  4. Report the four-cell matrix, not a single percentage. Both correct, A only, B only, neither. A single accuracy figure hides which tool is right when it matters.
  5. Then segment the errors by category. This is the step everyone skips and the one that decides the evaluation. Count misclassifications separately for loan disbursals, EMI outflows, bounce and return charges, and income credits. Overall accuracy is a vanity metric. Accuracy on the transactions your credit policy depends on is the real one.

If your current provider will not run step 2 against a competitor on live cases, that is itself a finding.

Limitations

Stated plainly, because a benchmark that does not disclose its weaknesses is marketing.

The comparison covers one competing tool, not the full market. The sample is drawn from Indian lending cases across specific borrower segments and will not generalise to portfolios shaped differently. The study was run by us and adjudicated jointly with the partner institutions whose cases were used, which is stronger than an internal claim and weaker than an independent audit. And categorisation accuracy, though it is the metric this market most needs and least publishes, is not the only thing that matters in a BSA tool — format coverage, tamper detection and integration effort all belong in an evaluation.

None of which changes the underlying point. The numbers are on the page and the protocol is above. Test it.

Frequently asked questions

What is a good accuracy rate for a bank statement analyser? Ask which accuracy. Extraction accuracy above 99% is common and near-meaningless as a differentiator. Categorisation accuracy is the number that varies between tools, and published, adjudicated figures are rare. In this study the two tools scored 92.8% and 80.3% on identical data.

Why do bank statement analysers misclassify transactions? Most were built as extraction engines with a classification layer added on top, and that layer typically maps narration strings to categories by pattern. Indian bank narrations are heavily abbreviated, bank-specific and often truncated, so pattern matching produces plausible-looking labels — usually "transfer to" or "from" the entity named in the narration — that are technically defensible and analytically empty.

Which transaction categories matter most for credit decisions? Loan disbursals and repayments, bounce and return charges, related-party and self-transfers, and income credits distinguished by source and repeatability. Errors in these four classes change DSCR, obligation coverage and conduct assessment directly. Errors in discretionary spend categories mostly do not.

How do I compare two bank statement analysis tools? Run both on identical raw files from closed cases, score only the transactions where they disagree, have your own analysts adjudicate, and report a four-outcome matrix segmented by transaction category. The full protocol is set out above.


Fiscus reads bank statements, GST filings and bureau data for banks and NBFCs across salaried, self-employed and SME borrowers. To run this benchmark against your current provider on your own files, book a parallel evaluation.

Written by

Zeus Dhanbhoora

Zeus Dhanbhoora is the CEO of BridgeUp Tech, the company behind Fiscus. He previously co-founded Bacferim Technologies and was an associate at the law firm Bharucha & Partners. He writes the Fiscus credit desk blog on benchmarks, fraud detection and credit underwriting methods.

See it on your own cases.

Run a parallel evaluation. Bring statements your team has already analysed and compare output, depth and time.

Book a demo