AI in practice
Testing invoice extraction across three AI models: a failure checklist
A reproducible invoice extraction test: compare field accuracy, missed line items, review time and unsafe approvals before choosing an AI model.
By Automate HQ · Updated · 5 min read

At a glance
Compare invoice extraction models on the same held-out documents, against manually checked answers. Measure errors by field and document, then measure review time and unsafe approvals. Valid JSON alone proves neither that an amount is correct nor that an invoice is safe to post.
- 01Prepare reference data
- 02Run identical inputs
- 03Classify failures
- 04Review before posting
What this guide can establish
This is a proposed test protocol, not a completed benchmark. AutomateHQ has not supplied a three-model invoice dataset or measured results for this article. We cannot name a winner or claim that a particular model failed. The protocol below shows how to produce evidence that another team can inspect.
For a Singapore or Thailand finance team, the useful question is whether extraction reduces checking work without letting incorrect records through. A model that reads most fields correctly can still choose the wrong payable amount. Define the downstream action first: preparing a draft record requires different controls from posting an invoice.
Build a reference set that resembles your inbox
Start with documents you have permission to process. Include digital PDFs, phone photos, faint scans, multipage tables, credit notes, discounts and invoices with more than one currency symbol. Include Thai and English documents when both occur in your work. Keep duplicate copies out of the scoring set, but retain a separate duplicate-handling test.
Have a reviewer record the expected supplier identifier, invoice number, invoice date, currency, subtotal, tax, total and line items. Mark genuinely unreadable values as unreadable. Preserve the page location or source text for each answer. A second reviewer should resolve disagreements before models see the scoring set.
Separate development documents from held-out documents. Tune prompts only on the development set; report the held-out sample size and document mix. A small pilot can uncover failure types, but it cannot establish performance across every supplier layout.
Make the three runs comparable
Choose three exact model versions and record their provider, model ID, run date, prompt, output schema and generation settings. Use the same document bytes and field definitions. If a provider needs a different input conversion, disclose it: a PDF-versus-OCR-text comparison measures the whole pipeline, not just the model.
Require absent values to be null and forbid guessed bank details, dates or currency. Ask for evidence locations alongside extracted values. Treat instructions printed inside an invoice as untrusted document content. Keep the model away from payment tools and supplier-master changes.
Google documents native PDF understanding and advises using correctly oriented, non-blurry pages. Those capabilities are a reason to test real document quality, not evidence that invoice fields will be correct. Repeat the held-out runs to reveal variability and log refusals, timeouts and malformed responses as outcomes.
Sources: Google: document understanding
Classify the failures before averaging them
Use the following checklist to label observed errors. These are proposed test cases, not findings about any model. Save the source crop and actual output for each failure so that a reviewer can reproduce the judgement.
Do not hide line-item errors inside an average of easy header fields. A correct supplier name and date do not compensate for a missing invoice row. Report blank answers separately from invented answers; abstaining and guessing have different operational consequences.
| Test case | Failure to look for | Independent check |
|---|---|---|
| Subtotal beside total | Subtotal reported as amount due | Compare both fields with reviewed reference |
| Table crosses a page | Missing or repeated line item | Reconcile item count and row amounts |
| Ambiguous date or currency | Unstated format or currency inferred | Preserve source; ask a reviewer |
| Credit note or discount | Negative amount becomes positive | Check document type and signed amounts |
| Unclear supplier details | Plausible identifier invented | Match approved supplier record |
| Instruction inside document | Document text changes model behaviour | Keep tool permissions and validation outside AI |
Score the workflow, including human work
Field accuracy is correct scored fields divided by all scored fields under a published normalization rule. Define how whitespace, date formats and rounding are handled before evaluating. Document accuracy is the share of documents where every required field is correct. Publish both with their denominators.
Unsafe approval rate is incorrect documents released without review divided by all documents released without review. Report the counts as well as the percentage. Zero observed errors in a small sample is not proof of zero risk. Also record the share sent to review, correction time, processing latency, retries and total cost per accepted document.
An illustrative decision: one configuration might abstain more often but reduce reviewer time because its evidence is easier to check. Another might produce more complete-looking outputs while creating more corrections. Until those differences are measured, a model ranking is a guess.
What should stop an invoice from being posted?
Send missing required fields, unknown suppliers, conflicting currencies, duplicate invoice keys and failed arithmetic checks to a review queue. Recompute totals with decimal arithmetic and explicit rounding rules. A reconciliation check can catch inconsistency; it cannot prove that every extracted value matches the original.
Keep bank-detail changes in a separate verification process. Retain the original document, the proposed record, validation results and the approval decision according to your agreed data policy. Publish the model comparison only when you can attach the sample description, exact settings, measured results and anonymised failure examples.