Confidence Scores in Practice: What They Tell You and What They Don't
Last reviewed 11 July 2026

In short
A confidence score is the model's own estimate of how sure it is about a draft — useful for ordering review work and gating automation, but never proof of correctness. We surface the most doubtful drafts first, set thresholds per account alongside materiality limits, and record every score in the audit trail.
Every entry the AI drafts in our platform carries a confidence score. It is one of the most useful signals in the system and one of the easiest to oversell, so this article tries to do something slightly unusual for a software company: explain what the number genuinely means, where it breaks down, and how we use it anyway.
What a confidence score actually is
When the AI drafts a double-entry transaction from a receipt or a bank line, it also reports how sure it is about what it produced. That self-estimate is the confidence score.
It is not a measurement taken from outside the model. It is the model's own opinion of its work — closer to a colleague saying "I'm fairly sure about this one" than to a laboratory instrument reading. Colleagues who say that are usually, but not always, right, and their "fairly sure" is not calibrated the same way across every kind of task. Model confidence behaves similarly.
What it does tell you
Used honestly, the score is valuable for one thing above all: ordering attention.
Across a day's drafts, higher-confidence entries are, on the whole, more likely to be routine and correct than lower-confidence ones. That relative signal — this pile is more doubtful than that pile — holds up well in practice even when the absolute numbers should not be read as probabilities. It lets us do two useful things:
- Sort the review queue doubtful-first. Human attention is a budget, and it is freshest at the start of a session. Doubtful-first ordering spends that fresh attention on the drafts the model itself flagged as shaky — the ambiguous categorisations, the odd document layouts, the counterparties it has not seen before. The routine tail of the queue gets reviewed too; it just is not allowed to exhaust the reviewer before the hard cases arrive. This is also a defence against rubber-stamping: a queue that opens with easy approvals trains reflexes, while a queue that opens with genuine puzzles trains scrutiny.
- Gate automation. Where you have agreed an automation policy for an account, the confidence score is one of the two thresholds a draft must clear before it may post without queueing. It is never the only one — more on that below.
What it does not tell you
Three limits matter, and we design around all of them rather than hoping they average out.
High confidence is not correctness. Models are sometimes confidently wrong — most often when a document resembles a familiar pattern but differs in a detail that matters. A supplier who usually bills for materials sending a one-off invoice for equipment hire can produce a high-confidence, wrong categorisation. This is precisely why confidence alone never authorises posting to a material account.
Low confidence is not wrongness. A blurry photo of a perfectly ordinary receipt yields a hesitant, often correct draft. Low scores mean "look at this", not "discard this". Low-confidence drafts are held for review or queried — never silently dropped, for the same reason suspected duplicates are held for a person rather than auto-discarded.
Scores are not uniformly calibrated. "90%" on a task the model has seen ten thousand times and "90%" on a rare document type are not the same strength of evidence. We treat scores as ordering and gating signals within a context, not as universal probabilities, and we check them against outcomes: every reviewer correction is recorded, so systematic over-confidence in a particular account or document type shows up in the record and informs the thresholds.
Why thresholds are set per account
A single global threshold would be convenient and wrong. The risk carried by a draft depends on where it is posting, not just how sure the model is. A £12 stationery receipt from a supplier you have used for years and a £4,000 subcontractor invoice are different decisions, even at identical confidence.
So thresholds live in per-account automation policies, agreed with you, and they work in tandem with materiality limits: a draft posts automatically only if its confidence clears the account's threshold and its value sits below the account's materiality limit. Sensitive accounts can simply run fully manual, where every draft queues regardless of score. The default for every account is manual; automation is something you opt into deliberately, account by account, not something the software assumes.
Scores in the audit trail
Every draft's confidence score is recorded permanently, alongside the validation result, any edits, and the approval decision. That matters for a reason that goes beyond tidiness: if an entry posted automatically, the trail shows exactly which score, which threshold, and which policy version permitted it. "Why is this entry in my books?" always has a specific, reconstructable answer — a person approved it, or a policy you agreed to accepted it, with the numbers on the record. And when thresholds change, that is a logged human decision too; the system does not quietly re-tune itself.
The honest summary
A confidence score is a useful opinion from a capable but fallible drafter. We use it to point human attention where doubt is highest, to gate narrow and agreed automation, and to keep an evidential record of every decision — and we decline to use it as a substitute for judgement. AI drafts. AgentLedger validates. People approve. The score helps decide which person looks first and how hard; it never decides that no one looks at all, unless you have explicitly told it to for that account, within limits you set.
Want this handled for you — with a person accountable for it?
Check eligibility / Ask for DetailsNo payment on the first step. No free trial.