Skip to content
AI Digital Hub

By use case

Document Processing & Extraction

Turn unstructured paperwork into validated records

Usually owned by

  • Finance
  • Operations
  • Compliance
  • Claims
85-95%
Documents processed without human touch

Structured and semi-structured formats after tuning

> 99%
Field-level accuracy on posted records

With confidence thresholds and validation rules

4 hrs
Invoice receipt to posted, from 3 days

The change

What actually differs afterwards

Not a maturity model. The concrete difference in how the work happens.

Today

  • Two people spend most of the week typing from PDFs
  • Error rates discovered downstream in reconciliation
  • Backlog grows whenever volume spikes or someone is on leave
  • No consistent audit trail for what was read from what document
  • Suppliers chase payments delayed by data entry

Afterwards

  • Documents parsed, validated and posted automatically
  • Confidence scores route only genuine exceptions to a person
  • Volume spikes absorbed without a backlog
  • Every extracted field traceable to its source page
  • Cycle time in hours instead of days

The most reliable automation there is

Document processing is where AI automation most consistently pays back, for three reasons: the input is high volume, the correct answer is verifiable, and the current process is expensive and universally disliked.

It is also where accuracy claims deserve scrutiny. A system that is 95% accurate on invoice totals is not a system you post to a ledger unattended — it is a system that needs the other 5% routed to a person, which is exactly how we build it.

The pipeline

Classify. What kind of document is this, and does it belong in this process at all.

Extract. Fields as typed values with a confidence score and a reference back to the page region they came from. That reference is what makes review fast and audit possible.

Validate. Business rules, not just format checks. Does the total match the line items. Does the supplier exist. Does the PO have remaining budget. Is the tax calculation right for that jurisdiction.

Route. Clean records post automatically. Anything below threshold or failing validation goes to a review queue with the document and the flagged field side by side.

Learn. Corrections made in review feed the evaluation set, so accuracy is tracked over time rather than assumed to hold.

How to read an accuracy claim

Any number quoted without these four qualifiers is close to meaningless, and asking for them is the fastest way to evaluate a proposal — including ours.

Measured on whose documents. A vendor benchmark on clean public invoices tells you nothing about your suppliers' scans. Insist on a figure from your own sample.

Field-level or document-level. A document with twelve fields at 98% per-field accuracy is correct end-to-end about 78% of the time. Both numbers are honest; only one answers the question you are actually asking.

Before or after the confidence threshold. Accuracy on records that were posted automatically should be very high, because the uncertain ones were routed away. That figure and the straight-through rate have to be quoted together — either alone can be made to look good by sacrificing the other.

Which fields. Accuracy on a supplier name is not accuracy on a tax amount. We report per field, because the fields differ enormously in both difficulty and the cost of getting them wrong.

Where the model can be confidently wrong

Extraction failures are usually obvious — a garbled value fails validation and routes to review. The dangerous case is a plausible wrong answer: a total read from the wrong column that happens to be a valid number, or a date correctly extracted from the wrong one of three dates on the page.

This is hallucination in its most mundane and most expensive form, and it is why validation rules are business rules rather than format checks. Does the total match the line items. Does the supplier exist in the master. Does the tax compute for that jurisdiction. Those catch the plausible-wrong cases that a confidence score alone will not, because the model can be confident and mistaken at the same time.

Where this shows up by sector

The same pipeline, with different documents and different evidential burdens:

  • Financial services — KYC packs, statements and settlement files, with an audit trail that has to satisfy an inspector.
  • Healthcare — claims, pre-authorisation and coding, with a clinician or coder accountable for every output.
  • Logistics — bills of lading, customs paperwork and freight invoices, with more partner format variation than any other sector.
  • Manufacturing — supplier quotations and quality records, where the extracted data has to reconcile against a part master.

FAQ

Questions we get asked

Our documents are scans, and some are bad ones. Does that work?

Mostly yes, and the modern models are far better at degraded scans than previous OCR generations. Quality still degrades on genuinely poor input, which is why every field carries a confidence score and low-confidence records go to review rather than into your ledger.

Do we need to tell it about every layout?

No. Unlike template-based OCR, this generalises across layouts, which is what makes it viable when you receive invoices from four hundred suppliers in four hundred formats. Unusual layouts may need examples added to the evaluation set.

How do you stop a wrong number reaching the accounts?

Validation rules plus thresholds. Totals must reconcile, dates must be plausible, supplier and PO must exist, tax must compute correctly. Anything failing validation or below the confidence threshold goes to a human queue with the source page shown alongside.

What is the realistic saving?

For a team spending two full-time equivalents on data entry, typically 70-85% of that time, with the remainder going to genuine exception handling. We model it against your actual volumes in the readiness audit rather than quoting a benchmark.

What does a first engagement cost and how long does it take?

Four weeks to put one document type into production, integrated with the system that receives the records and monitored. That is the Automation Sprint. If you have several document types, we sequence them rather than building them together — the first one carries the integration and evaluation scaffolding, so the second and third are materially cheaper and faster.

Does this work on handwritten documents?

Partially, and the honest answer depends on what kind. Printed forms with handwritten field entries work reasonably well. Free-form handwritten prose is unreliable enough that we would not post it to a system of record unreviewed. We test against your actual documents during the audit and give you a measured accuracy figure per document type rather than a general claim, because the variance between document types is larger than the variance between vendors.

What about documents in Hindi or regional languages?

Supported, with accuracy that varies by language and script quality. Devanagari handles well. Some regional scripts degrade more than English on poor scans, which pushes more records into the review queue rather than producing wrong answers — the confidence threshold is doing its job. Mixed-language documents, which are common in Indian commercial paperwork, are usually fine. We measure per language on your own sample before quoting a throughput number.

How does this compare to the OCR vendor we already use?

Template-based OCR is faster and cheaper where the layout is fixed and it already works — keep it for those. The difference shows up on the documents it cannot handle, which is usually where your remaining manual effort sits. The two coexist well, and a reasonable architecture routes fixed layouts to the existing tool and everything else to the model. We are happy to build that rather than replace something that works.

How much does the review queue cost to run?

More than most proposals admit, and it is a real line in the business case. At a 90% straight-through rate, one in ten documents still needs a person, and that person needs a review interface good enough that the check takes seconds rather than minutes. Building that interface properly is a meaningful share of the engineering, and skipping it means your reviewers are slower than the manual process they replaced.

Let's find out what is actually automatable

Bring a process that annoys you. In 30 minutes we will tell you whether AI helps, what it would cost, and where it would fail — even if the answer is don't bother.

Or email [email protected] · we reply within 1 business day