Free n8n template · Intermediate

Invoice Data Extraction

Extract fields from a text-based PDF invoice into structured JSON and a review spreadsheet. Nothing is paid or posted automatically.

How the starter works
  1. 01PDF upload
  2. 02Extract text
  3. 03AI + validation
  4. 04Review in Sheets

What it does

Who it is for
Finance and operations teams testing document intake with a small set of text-based invoices.
Input
One text-based PDF, up to 5 MB, uploaded as the multipart field data, plus an invoice_id field for tracking.
Process
Check the upload, extract PDF text locally in n8n, ask AI for invoice fields, validate types and basic arithmetic, then append a review row.
Output
Vendor, invoice number, date, currency, subtotal, tax and total, plus review issues and the source ID.

Validation status: Export structure and Code node logic are checked with local fixtures. Live n8n import and authenticated provider runs have not been verified. Follow the setup guide with synthetic data before using it with real records.

Set up the workflow

Tools and prerequisites

  • An n8n instance with Code, HTTP Request and Webhook nodes. The downloads use built-in nodes only; import compatibility still needs checking on your installed version.
  • Your own OpenAI API account with billing enabled. Store the key in an n8n Header Auth credential named OpenAI: header Authorization, value Bearer followed by your key.
  • A Google Sheets OAuth2 credential in n8n and a dedicated test spreadsheet. Enable the Sheets API and grant the connected account access to that spreadsheet.
  • A readable, unencrypted PDF with selectable text. This workflow does not include OCR.
  • An n8n Webhook Header Auth credential with header X-AutomateHQ-Key and a long random value. Use a trusted upload client.

Installation steps

  1. Import the JSON and create a sheet tab named InvoiceReview with the headers below.
  2. Connect OpenAI, Sheets and webhook credentials, then set Configure workflow.
  3. Select Listen for test event in Receive invoice. Send a multipart/form-data request with data containing the PDF and invoice_id containing your unique test ID. Let your HTTP client set the multipart boundary.
  4. Inspect the extracted text, validated JSON and review row against the source PDF. Use a synthetic invoice first.
  5. Try a scanned PDF and confirm the run fails before writing. Try an invoice with missing currency and confirm it creates a flagged review row, not invented data.
  6. Keep a person responsible for checking the PDF before copying any amount into an accounting system. Activate the production webhook only after testing.

Spreadsheet column order

Create these headers in row 1. The workflow appends values in this order using RAW input, which keeps text from becoming a spreadsheet formula.

invoice_id, vendor, invoice_number, invoice_date, currency, subtotal, tax, total, issues, needs_review

Configuration points

Configure workflow
Set spreadsheetId, sheetName and model. Start with a dedicated InvoiceReview spreadsheet tab.
Receive invoice
Select your Webhook Header Auth credential. The upload field must be named data; invoice_id is a separate form field.
Extract PDF text
Keep operation PDF and input binary field data.
Classify with OpenAI
Select the OpenAI Header Auth credential. Unknown fields must remain null; do not ask the model to guess.
Append review row
Select your Google Sheets OAuth2 credential.

Example input & output

Synthetic examples, not customer data. AI output may differ on each run. A well-formed response still needs review.

Input

{
  "invoice_id": "invoice-demo-001",
  "data": "invoice.pdf (multipart binary upload)",
  "sample_pdf_text": "Example Supplies | Invoice INV-1042 | 2026-10-01 | SGD | Subtotal 100.00 | Tax 9.00 | Total 109.00"
}

Illustrative output

{
  "invoice_id": "invoice-demo-001",
  "vendor": "Example Supplies",
  "invoice_number": "INV-1042",
  "invoice_date": "2026-10-01",
  "currency": "SGD",
  "subtotal": 100,
  "tax": 9,
  "total": 109,
  "issues": [],
  "needs_review": true
}

What can go wrong

Missing file or oversized upload

Use multipart field data and invoice_id. The workflow checks the file signature and a 5 MB limit before PDF extraction; also enforce upload limits at your reverse proxy.

Empty or unreadable PDF text

Scans, encrypted PDFs and malformed documents stop the run. Add a separate OCR path if needed; do not interpret missing text as a valid empty invoice.

Missing fields or arithmetic mismatch

The row remains for review with issues. Totals use a 0.02 tolerance; discounts, rounding and credit notes require additional business rules.

Failed nodes stop execution and remain visible in n8n’s Executions view. No write node retries automatically. A caller timeout does not prove that no write happened: inspect downstream systems before replaying. Add an error workflow if you need alerts.

Limitations

  • Text-based PDFs only. No OCR, email attachment trigger, line-item extraction, tax compliance validation or accounting posting.
  • A valid JSON shape does not prove the extracted values match the document. Every row requires review against the PDF.
  • The text limit is 20,000 characters. Longer documents fail explicitly instead of silently truncating totals.

Before your team depends on it

From demo to production

A successful test run proves one path works. Production engineering also accounts for repeated requests, unavailable services and people correcting the result.

  • Use invoice_id plus a file hash and supplier invoice number to detect repeat uploads. An append-only spreadsheet is not a payment ledger.
  • Retain source evidence with controlled access and connect each extracted field to a page or text span. Record corrections and reviewer identity in an audit trail.
  • Add OCR, malware scanning and document isolation before accepting arbitrary uploads. Queue large documents with backpressure and per-file cost limits.
  • Validate currency, rounding, credits and supplier-specific layouts with deterministic rules. Require approval before creating ERP entries; keep payment authorization separate.

Security & privacy

PDF text is extracted in your n8n environment and sent to OpenAI. Extracted fields are saved to Google Sheets; binary files and text may also remain in n8n execution storage. Use synthetic invoices first, restrict access and agree deletion rules with the data owner.

Secrets belong in n8n’s credential store. Successful production payloads are not saved by this export; manual and failed executions are saved for debugging. Configure pruning and binary-data retention, restrict execution access and delete test runs when finished. Webhook headers can appear in execution inputs, so rotate test keys.

Estimated running costs

One model request and one Sheets append per readable invoice, plus n8n hosting or execution charges. PDF extraction runs in n8n; no OCR service is included. Retries may create additional model charges.

For 1,000 runs, estimate the model cost as (average input tokens × input price per million + average output tokens × output price per million) ÷ 1,000. For example, at hypothetical rates of $1 input and $4 output per million tokens, 1,500 input and 250 output tokens per run would cost $2.50 per 1,000 runs, before retries and hosting. These are illustrative rates, not a provider quote.

Use your execution usage and the current OpenAI pricing and n8n plan to budget. The template download is free; connected services may charge.

Make it your own

  • Add an OCR branch with file-size and page-count limits.
  • Create a review screen that shows fields next to the original PDF.
  • Write approved records to accounting software with duplicate checks.

Official setup references

Importing n8n workflowsGoogle OAuth credentials in n8nOpenAI structured outputs