AI in practice

Testing invoice extraction across three AI models: a failure checklist

A reproducible invoice extraction test: compare field accuracy, missed line items, review time and unsafe approvals before choosing an AI model.

By Automate HQ · Updated · 5 min read

Three magnifying lenses over an invoice, illustrating independent checks of extracted fields.
Original AI-generated editorial illustration by AutomateHQ. Illustrative, not a product screenshot.

At a glance

Compare invoice extraction models on the same held-out documents, against manually checked answers. Measure errors by field and document, then measure review time and unsafe approvals. Valid JSON alone proves neither that an amount is correct nor that an invoice is safe to post.

Workflow steps
  1. 01Prepare reference data
  2. 02Run identical inputs
  3. 03Classify failures
  4. 04Review before posting

What this guide can establish

This is a proposed test protocol, not a completed benchmark. AutomateHQ has not supplied a three-model invoice dataset or measured results for this article. We cannot name a winner or claim that a particular model failed. The protocol below shows how to produce evidence that another team can inspect.

For a Singapore or Thailand finance team, the useful question is whether extraction reduces checking work without letting incorrect records through. A model that reads most fields correctly can still choose the wrong payable amount. Define the downstream action first: preparing a draft record requires different controls from posting an invoice.

Build a reference set that resembles your inbox

Start with documents you have permission to process. Include digital PDFs, phone photos, faint scans, multipage tables, credit notes, discounts and invoices with more than one currency symbol. Include Thai and English documents when both occur in your work. Keep duplicate copies out of the scoring set, but retain a separate duplicate-handling test.

Have a reviewer record the expected supplier identifier, invoice number, invoice date, currency, subtotal, tax, total and line items. Mark genuinely unreadable values as unreadable. Preserve the page location or source text for each answer. A second reviewer should resolve disagreements before models see the scoring set.

Separate development documents from held-out documents. Tune prompts only on the development set; report the held-out sample size and document mix. A small pilot can uncover failure types, but it cannot establish performance across every supplier layout.

Make the three runs comparable

Choose three exact model versions and record their provider, model ID, run date, prompt, output schema and generation settings. Use the same document bytes and field definitions. If a provider needs a different input conversion, disclose it: a PDF-versus-OCR-text comparison measures the whole pipeline, not just the model.

Require absent values to be null and forbid guessed bank details, dates or currency. Ask for evidence locations alongside extracted values. Treat instructions printed inside an invoice as untrusted document content. Keep the model away from payment tools and supplier-master changes.

Google documents native PDF understanding and advises using correctly oriented, non-blurry pages. Those capabilities are a reason to test real document quality, not evidence that invoice fields will be correct. Repeat the held-out runs to reveal variability and log refusals, timeouts and malformed responses as outcomes.

Sources: Google: document understanding

Classify the failures before averaging them

Use the following checklist to label observed errors. These are proposed test cases, not findings about any model. Save the source crop and actual output for each failure so that a reviewer can reproduce the judgement.

Do not hide line-item errors inside an average of easy header fields. A correct supplier name and date do not compensate for a missing invoice row. Report blank answers separately from invented answers; abstaining and guessing have different operational consequences.

Classify the failures before averaging them
Test caseFailure to look forIndependent check
Subtotal beside totalSubtotal reported as amount dueCompare both fields with reviewed reference
Table crosses a pageMissing or repeated line itemReconcile item count and row amounts
Ambiguous date or currencyUnstated format or currency inferredPreserve source; ask a reviewer
Credit note or discountNegative amount becomes positiveCheck document type and signed amounts
Unclear supplier detailsPlausible identifier inventedMatch approved supplier record
Instruction inside documentDocument text changes model behaviourKeep tool permissions and validation outside AI

Score the workflow, including human work

Field accuracy is correct scored fields divided by all scored fields under a published normalization rule. Define how whitespace, date formats and rounding are handled before evaluating. Document accuracy is the share of documents where every required field is correct. Publish both with their denominators.

Unsafe approval rate is incorrect documents released without review divided by all documents released without review. Report the counts as well as the percentage. Zero observed errors in a small sample is not proof of zero risk. Also record the share sent to review, correction time, processing latency, retries and total cost per accepted document.

An illustrative decision: one configuration might abstain more often but reduce reviewer time because its evidence is easier to check. Another might produce more complete-looking outputs while creating more corrections. Until those differences are measured, a model ranking is a guess.

What should stop an invoice from being posted?

Send missing required fields, unknown suppliers, conflicting currencies, duplicate invoice keys and failed arithmetic checks to a review queue. Recompute totals with decimal arithmetic and explicit rounding rules. A reconciliation check can catch inconsistency; it cannot prove that every extracted value matches the original.

Keep bank-detail changes in a separate verification process. Retain the original document, the proposed record, validation results and the approval decision according to your agreed data policy. Publish the model comparison only when you can attach the sample description, exact settings, measured results and anonymised failure examples.

Official references