# Invoice Data Extraction

Extract fields from a text-based PDF invoice into structured JSON and a review spreadsheet. Nothing is paid or posted automatically.

By AutomateHQ, an AI and workflow automation business serving companies in Singapore and Thailand.

Guide: https://automatehq.org/resources/n8n-invoice-data-extraction/

## Status and permission to use

Free to use and adapt in personal or commercial workflows. Provided as a starter without warranty. No signup. No secrets or pinned customer data. Local fixture tests cover Code node validation and export structure. Not yet imported into a live n8n instance or verified with provider credentials. Do not treat these checks as end-to-end certification.

## Requirements

- An n8n instance with Code, HTTP Request and Webhook nodes. The downloads use built-in nodes only; import compatibility still needs checking on your installed version.
- Your own OpenAI API account with billing enabled. Store the key in an n8n Header Auth credential named OpenAI: header Authorization, value Bearer followed by your key.
- A Google Sheets OAuth2 credential in n8n and a dedicated test spreadsheet. Enable the Sheets API and grant the connected account access to that spreadsheet.
- A readable, unencrypted PDF with selectable text. This workflow does not include OCR.
- An n8n Webhook Header Auth credential with header X-AutomateHQ-Key and a long random value. Use a trusted upload client.

## Setup

1. Import the JSON and create a sheet tab named InvoiceReview with the headers below.
2. Connect OpenAI, Sheets and webhook credentials, then set Configure workflow.
3. Select Listen for test event in Receive invoice. Send a multipart/form-data request with data containing the PDF and invoice_id containing your unique test ID. Let your HTTP client set the multipart boundary.
4. Inspect the extracted text, validated JSON and review row against the source PDF. Use a synthetic invoice first.
5. Try a scanned PDF and confirm the run fails before writing. Try an invoice with missing currency and confirm it creates a flagged review row, not invented data.
6. Keep a person responsible for checking the PDF before copying any amount into an accounting system. Activate the production webhook only after testing.

## Sheet columns, in order

```text
invoice_id,vendor,invoice_number,invoice_date,currency,subtotal,tax,total,issues,needs_review
```

Sheets append uses RAW input, so text is not executed as spreadsheet formulas. No write node retries automatically.

## Configuration

### Configure workflow

Set spreadsheetId, sheetName and model. Start with a dedicated InvoiceReview spreadsheet tab.

### Receive invoice

Select your Webhook Header Auth credential. The upload field must be named data; invoice_id is a separate form field.

### Extract PDF text

Keep operation PDF and input binary field data.

### Classify with OpenAI

Select the OpenAI Header Auth credential. Unknown fields must remain null; do not ask the model to guess.

### Append review row

Select your Google Sheets OAuth2 credential.

## Example input

```json
{
  "invoice_id": "invoice-demo-001",
  "data": "invoice.pdf (multipart binary upload)",
  "sample_pdf_text": "Example Supplies | Invoice INV-1042 | 2026-10-01 | SGD | Subtotal 100.00 | Tax 9.00 | Total 109.00"
}
```

## Illustrative output (model output varies)

```json
{
  "invoice_id": "invoice-demo-001",
  "vendor": "Example Supplies",
  "invoice_number": "INV-1042",
  "invoice_date": "2026-10-01",
  "currency": "SGD",
  "subtotal": 100,
  "tax": 9,
  "total": 109,
  "issues": [],
  "needs_review": true
}
```

## Troubleshooting

### Missing file or oversized upload

Use multipart field data and invoice_id. The workflow checks the file signature and a 5 MB limit before PDF extraction; also enforce upload limits at your reverse proxy.

### Empty or unreadable PDF text

Scans, encrypted PDFs and malformed documents stop the run. Add a separate OCR path if needed; do not interpret missing text as a valid empty invoice.

### Missing fields or arithmetic mismatch

The row remains for review with issues. Totals use a 0.02 tolerance; discounts, rounding and credit notes require additional business rules.

Failed runs remain visible in n8n Executions; configure a separate error workflow for alerts. Webhook failures return an error rather than a success result. Check your caller's timeout; a timeout does not prove that no external write happened.

## Limitations

- Text-based PDFs only. No OCR, email attachment trigger, line-item extraction, tax compliance validation or accounting posting.
- A valid JSON shape does not prove the extracted values match the document. Every row requires review against the PDF.
- The text limit is 20,000 characters. Longer documents fail explicitly instead of silently truncating totals.

## From demo to production

- Use invoice_id plus a file hash and supplier invoice number to detect repeat uploads. An append-only spreadsheet is not a payment ledger.
- Retain source evidence with controlled access and connect each extracted field to a page or text span. Record corrections and reviewer identity in an audit trail.
- Add OCR, malware scanning and document isolation before accepting arbitrary uploads. Queue large documents with backpressure and per-file cost limits.
- Validate currency, rounding, credits and supplier-specific layouts with deterministic rules. Require approval before creating ERP entries; keep payment authorization separate.

## Security and retention

PDF text is extracted in your n8n environment and sent to OpenAI. Extracted fields are saved to Google Sheets; binary files and text may also remain in n8n execution storage. Use synthetic invoices first, restrict access and agree deletion rules with the data owner. Successful production execution payloads are not saved by this export; manual and failed executions are saved for debugging. Set pruning and binary-data retention explicitly, and delete synthetic test runs when finished. Header credentials may appear in webhook execution inputs; restrict access and rotate test keys.

## Costs

One model request and one Sheets append per readable invoice, plus n8n hosting or execution charges. PDF extraction runs in n8n; no OCR service is included. Retries may create additional model charges. Estimate per 1,000 runs as (input tokens × input price per million + output tokens × output price per million) / 1,000, then add hosting and retries. Use current pricing at https://openai.com/api/pricing/ and https://n8n.io/pricing/.

## Customization

- Add an OCR branch with file-size and page-count limits.
- Create a review screen that shows fields next to the original PDF.
- Write approved records to accounting software with duplicate checks.

## Live acceptance checklist

- Record n8n version and import the JSON; confirm every node and credential type resolves.
- Configure only sandbox accounts and synthetic data.
- Run a valid case; compare each external write to the validated result.
- Run invalid input and simulated model refusal; confirm no downstream writes.
- Simulate provider failure and inspect partial completion.
- Check duplicate/retry behavior and delete test data.
- Record results before submitting to n8n or claiming live compatibility.

Need this connected to your actual systems? https://automatehq.org/#request-demo
