[ Buyer's Guide ]

How to Evaluate a Bank Statement Analysis Tool Without Being Sold To

Bank statement analysis software comparison, done properly. Every tool looks excellent on a clean salary statement chosen by the vendor. How to build the test set, score categorisation rather than extraction, and the ten questions that separate tools.

Zeus Dhanbhoora

10 min read · 17 September 2026

How to Evaluate a Bank Statement Analysis Tool Without Being Sold To
Contents
  1. Build the test set before you talk to anyone
  2. Score categorisation, not extraction
  3. Ten questions that separate tools
  4. Measure analyst time correctly
  5. Run parallel, do not pilot
  6. Be realistic about switching cost
  7. The regulatory floor underneath all of this
  8. Questions about the vendor, not the product
  9. The short version
  10. Frequently asked questions

When did you last see a tool fail in its own demo?

Every bank statement analyser performs beautifully on a clean, native-PDF salary statement from a large private bank. That is the easy case, it is the case vendors bring, and it tells you nothing about how the tool will behave on a twelve-month multi-account SME file with two scanned cooperative bank statements in it.

Bank statement analysis software comparison has a structural problem, and it is not that vendors are dishonest. It is that the selection of test data sits with the party being tested. This is a guide to taking it back.

Build the test set before you talk to anyone

Assemble it first, independently, and give every vendor the identical set. Twenty to thirty closed cases is enough if the composition is right.

Weight it to your actual book, not to convenience. If 40% of your volume is self-employed, 40% of the test set is self-employed. Most evaluation sets skew salaried because those files are tidier to assemble, which systematically tests the segment where tools differ least.

Include the formats that break things. Scanned statements. Cooperative banks, small finance banks and regional rural banks. Multi-account cases with four or more accounts. At least a few twelve-month business files.

Include cases that went bad. Accounts you approved that later defaulted are the most valuable test data your institution owns, and almost nobody uses them for this. The question is not whether a tool categorises them correctly — it is whether the signals that would have flagged them are visible in the output. If a file that defaulted in month seven shows nothing unusual in the tool's report, that is your finding.

Include the cases your analysts found hard. Every credit team has files that took three passes. Those are where the difference between tools is largest.

Score categorisation, not extraction

Extraction — did the tool read ₹88,201 as ₹88,201 — is close to solved on native PDFs. It is also the easiest number to score well on, which is exactly why it is the number quoted.

Categorisation is where tools diverge and where credit decisions are made. Score it properly:

  1. Run both tools, or the candidate against your incumbent, on the identical raw files. No pre-processing, no format hints.
  2. Extract only the transactions where the outputs disagree. Where both agree, adjudication adds nothing and consumes your analysts' time.
  3. Have your own credit analysts adjudicate — not either vendor.
  4. Report a four-cell matrix: both correct, A only, B only, neither.

A single accuracy percentage hides which tool is right when it matters. We published our own results on this basis across 200,001 transactions, including the full protocol and the cases where both tools failed.

Then segment the errors by category. This is the step that decides the evaluation and the one almost everyone skips. Count misclassifications separately for loan disbursals, EMI outflows, bounce and return charges, and income credits. A tool at 88% overall that is wrong on a third of loan disbursals is worse for you than a tool at 84% that gets them all right, because a disbursal read as revenue inflates turnover and hides leverage in the same transaction.

Overall accuracy is a vanity metric. Accuracy on the categories your credit policy turns on is the real one.

Ten questions that separate tools

Each of these is difficult to answer convincingly without the underlying capability actually existing. Ask them in a working session, not in an RFP response.

1. Are the two kinds of cheque return separated? One is the borrower's own payment failing, which Indian banks record as an inward return because the cheque came in through their clearing. The other is a cheque the borrower deposited being returned, which comes back as an outward return and means their customer defaulted. Treating them identically penalises a borrower for their debtor's failure. Ask which convention the tool uses as well, because some label from the borrower's point of view and reverse both. This one question reliably distinguishes tools built for business lending from tools adapted to it.

2. Are bounce reason codes retained and classified? Not just recorded — classified into technical failures, genuine shortfalls and deliberate refusals such as stopped payments and cancelled mandates. Our reason code reference sets out which codes fall where.

3. Is the counterparty named on every transaction, or inferred from the category? "Transfer to [narration string]" is not false and not usable. Without resolved counterparties you cannot compute concentration, detect circularity, or notice that a top buyer has disappeared.

4. Are self-transfers matched on both legs and netted? Flagging them individually is not the same as eliminating them. Ask to see the netting on a four-account case.

5. Are fraud rules tuned by segment? Ask specifically what the round-figure-credit rule does differently for a sole proprietor and a corporate. If the answer is nothing, the false positive rate is being paid by your analysts, who will learn to ignore the flags.

6. Does the balance trail reconcile across page boundaries? Many tools reconcile within a page and not across. Page breaks are exactly where a tamper survives.

7. Does a flag carry a page and line reference? If your analyst has to search the statement to find what an alert refers to, the alert costs more than it saves.

8. Does the taxonomy include business-specific classes? Bill discounting, facility interest, inter-company settlement, drawings. Either these exist as categories or business transactions are being mapped into retail buckets.

9. What happens to a narration with no recoverable information? Some Indian bank narrations are truncated past the point of meaning. The honest answer is a low-confidence flag for review. A confident wrong category is worse than an admitted gap.

10. Will you run against our incumbent on our live files? The answer to this is itself a finding.

Measure analyst time correctly

The efficiency case is usually made on licence cost against headcount. The number that actually matters is reclassification minutes per case.

Measure it directly: take ten cases, time how long an analyst spends correcting and supplementing the tool's output before it is usable, and do it for both tools on the same files. In our own parallel runs, per-case reclassification on multi-account business cases lands between 30 and 45 minutes, and that figure is driven almost entirely by the error volume. It is why the 64% reduction in misclassifications we measured in our 200,001-transaction benchmark matters more to a cost base than a difference in licence fee. Both figures are ours, measured by us, which is exactly the kind of claim this post is telling you to test rather than accept.

Measure it separately by segment. A tool may be efficient on salaried files and expensive on SME ones, and if SME is where you are growing, the blended average conceals the thing you need to know.

Run parallel, do not pilot

A pilot switches a slice of volume to a new tool and compares outcomes over time. It is slow, it disrupts the cases in the slice, and it produces a comparison confounded by everything that changed in between.

Parallel running sends the same live cases through both systems for four to six weeks while the incumbent continues to be the system of record. Nothing is disrupted, the comparison is on identical inputs, and your analysts see both outputs side by side on files they already understand.

Two things to insist on. The parallel period should cover a full month-end cycle, because that is when volume patterns and pressure are real. And the outputs should be reviewed by the analysts who will use the tool, not by the project team — the people doing reclassification know within a week which output costs them less.

Be realistic about switching cost

Integration. REST API into your LOS or LMS is the clean path but requires engineering time on your side. A browser-based dashboard requires none, which matters if your technology queue is twelve months long. Ask which is available and what each actually takes — onboarding within 48 hours and onboarding within two quarters are both claims you should test against a reference.

Configurability. Whether report fields, categorisation rules and key metrics can be aligned to your existing credit policy, or whether your policy has to bend to the tool's output format. The second is a hidden cost paid by every analyst, every case, forever.

Account Aggregator handling. Whether AA-delivered JSON is ingested natively or requires conversion, and whether the same analysis logic applies to both AA and document paths. If the two channels produce different outputs, you are maintaining two credit processes. The boundaries of what AA does and does not solve are worth understanding before this conversation.

Contract structure. Bundled commitments, minimum volumes and multi-year lock-ins are common and are the mechanism by which a switching decision is deferred past the point where it matters.

The regulatory floor underneath all of this

Since 28 November 2025 the outsourcing rules have been consolidated and rewritten. The Reserve Bank of India (Commercial Banks — Managing Risks in Outsourcing) Directions, 2025 and the matching directions for NBFCs, small finance banks, urban and rural cooperative banks and all India financial institutions replaced the earlier outsourcing and IT outsourcing instructions outright. Existing IT outsourcing agreements have to be brought into line by 10 April 2026 or their renewal date, whichever falls first.

Scope is layered for NBFCs. Base Layer entities are covered only for outsourcing of financial services; Middle Layer and above are covered for IT services as well. A bank statement analysis vendor sits in the IT services chapter.

Five of the requirements convert directly into evaluation questions.

  • Due diligence is mandatory, not discretionary. The directions require an NBFC to perform appropriate due diligence when considering or renewing an outsourcing arrangement, to assess the service provider's capability to meet its obligations on an ongoing basis. A selection made on the strength of a demo does not meet that standard, and the parallel run described above is the cheapest way to meet it.
  • Audit rights belong in the contract. The regulated entity must retain the right to audit the service provider, through its own or external auditors, and to obtain copies of audit and review reports.
  • Data stays in India. Storage of data only in India, as applicable under extant regulatory requirements. Ask where processing happens, not only where the contract was signed.
  • Concentration risk is yours to assess. A bank is required to assess the impact of concentration risk from multiple outsourcing arrangements with the same service provider. If your statement analyser, your bureau connector and your fraud engine all come from one group, that is a finding you are expected to have made and documented.
  • Grievances stay with you. Responsibility for redressal of customer grievances relating to outsourced services rests with the regulated entity. A misclassification that produces a wrong decline is yours to answer for, not the vendor's.

There is also an exit requirement: the IT outsourcing policy must contain a clear exit strategy ensuring business continuity during and after exit. That reframes the lock-in question. A multi-year commitment with no tested exit path is not only a commercial problem, it is a gap in a policy your regulator expects you to hold.

Questions about the vendor, not the product

Three that tend to be more informative than any feature discussion.

Will they publish their methodology? Any vendor can assert an accuracy figure. Ask what was measured, on what data, against what baseline, adjudicated by whom. A vendor unwilling to describe the test has told you the test was not designed to be described.

RBI has proposed to put this on a formal footing. Its August 2024 draft circular on regulatory principles for the management of model risks in credit, since replaced by a fresh draft issued for comments in June 2026 as Guidance on Regulatory Principles for Model Risk Management, would require that agreements with third party model providers give the regulated entity access to minimum technical documentation covering the design, configuration and operation of the model, and states that the regulated entity remains ultimately responsible and accountable for the integrity and outcomes of outsourced models. It would also require validation independent of model selection, at least annually. Whether or not it is finalised in its current form, it is a reasonable description of what to write into a contract now, and a categorisation engine driving credit decisions is a model in every sense that matters.

What is their policy on the residual? Every tool has transactions it cannot resolve. A vendor claiming otherwise is either not measuring or not telling you. The useful answer describes what happens to those transactions.

Who trained the categorisation logic? Categorisation quality is largely a function of whether the people building it have read credit files. A taxonomy designed by engineers without credit input produces categories that are technically coherent and analytically unhelpful.

The short version

Bring your own files, weighted to your own mix, including the ones that went bad. Score categorisation on disagreements, adjudicated by your people. Segment the errors by the categories your policy depends on. Ask the ten questions. Run parallel through a month-end. Measure reclassification minutes, not licence cost.

None of this requires a consultant and it takes about a week of an analyst's time. The alternative is choosing a tool based on how it performed on a file someone else selected.

Frequently asked questions

How do you compare bank statement analysis tools objectively? Run both on identical raw files drawn from your own closed cases, score only the transactions where the outputs disagree, have your own analysts adjudicate rather than either vendor, and report a four-outcome matrix segmented by transaction category.

What accuracy figure should a BSA vendor be able to provide? Categorisation accuracy, with the methodology disclosed — what was measured, on what data, against what comparison, adjudicated by whom. Extraction accuracy on native PDFs is the easiest figure to score well on and the least informative one to compare.

How long should a bank statement analysis evaluation take? Roughly a week of analyst time to assemble and score a test set, plus four to six weeks of parallel running covering at least one full month-end cycle.

What is the difference between a pilot and a parallel run? A pilot diverts a portion of live volume to the new tool, disrupting those cases and producing a comparison confounded by time. A parallel run sends the same cases through both systems while the incumbent remains the system of record, giving a direct comparison on identical inputs with no disruption.

Does choosing a bank statement analysis vendor fall under RBI outsourcing rules? For IT services it does, for banks and for NBFCs in the Middle Layer and above, under the Managing Risks in Outsourcing Directions issued on 28 November 2025. They require documented due diligence before engagement or renewal, contractual audit rights, data storage in India, an assessment of concentration risk where one provider supplies several services, and a clear exit strategy. Responsibility for customer grievances stays with the lender.

What is the most commonly missed question in a BSA evaluation? Whether errors are segmented by transaction category. A tool with strong overall accuracy that misclassifies loan disbursals is more dangerous than a weaker tool that identifies them correctly, because a disbursal read as revenue inflates turnover and conceals leverage simultaneously.


Fiscus runs parallel evaluations on live cases against whatever you use today, with your analysts scoring the output. Book one here.

Written by

Zeus Dhanbhoora

Zeus Dhanbhoora is the CEO of BridgeUp Tech, the company behind Fiscus. He previously co-founded Bacferim Technologies and was an associate at the law firm Bharucha & Partners. He writes the Fiscus credit desk blog on benchmarks, fraud detection and credit underwriting methods.

See it on your own cases.

Run a parallel evaluation. Bring statements your team has already analysed and compare output, depth and time.

Book a demo