On this page9 sections
AI can extract data from PDFs and invoices reliably when it is part of a pipeline: get the text out of the document, ask a large language model to fill a fixed set of fields, validate every field with ordinary code (totals add up, dates are real, the supplier exists), and send anything that fails or looks uncertain to a person. It works well on varied layouts that defeat fixed templates. It fails on poor scans, handwritten notes, complex tables spread across pages, and any case where nobody checks the output before it is used.
We build these pipelines for invoices, purchase orders, delivery notes and forms. This post explains how each stage works and where the errors come from.
Why is extracting data from PDFs hard?
A PDF describes how a page looks. It does not describe what the content means. The text "Total 1,250.00" is just characters placed at a position on a page; nothing in the file says it is the invoice total.
Older approaches handled this with templates: for supplier A, the total is in this box. Templates are accurate for a fixed set of suppliers but break every time a layout changes or a new supplier appears. Large language models read the page more like a person does, using labels and context, which is why they handle variety much better.
Step 1: How do you get text out of the document?
There are two kinds of PDF, and they need different handling.
- Digital PDFs, generated by accounting software, contain real text. A PDF library can extract it, along with the position of each word, accurately and cheaply.
- Scanned PDFs and photos contain only an image. They need optical character recognition (OCR) to turn the image into text. OCR quality depends heavily on the scan: resolution, skew, shadows and stamps all reduce accuracy.
Check which kind each document is before processing. Running OCR on a digital PDF throws away accurate text and replaces it with guesses.
Some models can read page images directly. That helps with layout and tables, but the same rule applies: a blurry image produces unreliable output whichever method reads it.
Step 2: How does the model extract the fields?
Give the model the extracted text (or the page image) and a fixed schema: the exact fields you want and their types. For an invoice, typically:
supplier_name, supplier_tax_id, invoice_number, invoice_date,
due_date, currency, subtotal, tax_amount, total,
line_items[description, quantity, unit_price, amount]
Ask for structured output such as JSON that matches the schema, and instruct the model to return an empty value when a field is not present instead of guessing. Ask it to quote the text it used for each field where possible, which makes checking much easier.
Keep the instructions specific to your documents: which date format your suppliers use, how to treat credit notes, what to do when several tax rates appear.
Step 3: How do you validate the output?
This is the step that makes the system trustworthy, and it uses plain code with no AI involved:
- Arithmetic. Line amounts equal quantity times unit price. Line amounts add up to the subtotal. Subtotal plus tax equals the total.
- Formats. Dates are valid and plausible. Tax IDs match the expected pattern for the country. Currency codes are real.
- Reference data. The supplier matches one in your supplier list. The purchase order number exists and is open.
- Duplicates. The same supplier and invoice number has not been processed before.
- Source check. Each extracted value can be found in the document text.
Any document that fails a check goes to a person with the failed check highlighted.
Step 4: How do you handle the uncertain cases?
Route documents to human review when:
- Any validation check fails.
- Required fields are missing.
- The document type is unclear, such as a statement or a reminder instead of an invoice.
- The supplier is new.
- The amount is above a threshold you set.
The review screen should show the document beside the extracted fields so a person can confirm or correct in seconds. Store every correction, because those are your test cases for improving the instructions later.
What works well?
- Varied layouts from many suppliers, without building a template for each.
- Header fields such as supplier, dates, invoice number and totals on clean digital documents.
- Normalising values, for example converting different date formats to one standard.
- Classifying documents: invoice, credit note, receipt, statement, purchase order.
- Multiple languages, within reason, without separate setups.
Where does it fail?
- Poor scans and phone photos, where the text itself is wrong before the model sees it.
- Handwriting, especially numbers.
- Long tables that run across pages, with merged cells or subtotals in the middle.
- Ambiguous documents, such as an invoice that also lists previous unpaid balances.
- Confident errors. A model may fill a field with a plausible value that does not appear in the document. Validation and source checks exist to catch exactly this.
- Silent drift. Supplier layouts and your own needs change. Without regular sampling of results, accuracy can drop without anyone noticing.
How do you measure accuracy before going live?
Collect a set of real documents, say a hundred or more, covering your common suppliers and your awkward ones. Record the correct values by hand. Run the pipeline and measure field-level accuracy, plus how many documents were routed to review. Rerun this test set every time you change the instructions or the model.
Decide in advance what accuracy is acceptable for automatic posting and what goes to review. Many teams start with every document reviewed, then relax review for suppliers and fields with a strong record.
If this is your first automation project, our guide on which process to automate first explains why a human approval step belongs in anything that touches money.
Working with Syntora Ai
Syntora Ai builds AI extraction pipelines for PDFs, invoices and emails with validation and human review built in, and every system can stop and ask a person when its confidence is low. If documents are piling up in your inbox, write to hello@syntorahq.ai or see our AI systems practice.