AI · October 8, 2026 · 5 min read

Automating invoice and document processing with AI

Typing numbers from invoices, receipts, purchase orders and forms into another system is one of the most common jobs AI can take off a team's plate. Modern language models can read a messy PDF or a phone photo and pull out the fields you care about far better than older template-based tools.

But "the model reads the invoice" is only one step. A system you can trust with your books needs extraction, validation, human review and a way to measure accuracy. Here's how those pieces fit together.

Step 1: Define exactly what you need

Start with the output, not the model. Write down every field you need from each document type, its format and whether it's required.

{
  "supplier_name": "string",
  "invoice_number": "string",
  "invoice_date": "YYYY-MM-DD",
  "currency": "ISO 4217 code",
  "subtotal": "number",
  "tax": "number",
  "total": "number",
  "line_items": [{ "description": "string", "quantity": "number", "unit_price": "number" }]
}

A fixed schema does two things. It tells the model precisely what to return, and it gives your code something to check against. Most major model APIs now support structured outputs or JSON schemas, which make the response far more reliable than asking for JSON in plain text.

Step 2: Get the document into a readable form

Documents arrive as digital PDFs, scanned PDFs, photos and email attachments. Each needs slightly different handling.

  • Digital PDFs usually contain real text you can extract directly, which is cheaper and more accurate than reading an image.
  • Scans and photos need a vision-capable model or a separate OCR step first.
  • Multi-page documents should keep their page order, and long ones may need splitting so line items aren't cut off.
  • Keep the original file, linked to every extracted record, so anyone can check the source later.

Step 3: Validate everything the model returns

Never write extracted data straight into your accounting system. Run it through rules first. These catch a surprising share of errors, and they cost nothing to run.

  1. Schema checks: every required field is present and in the right format.
  2. Arithmetic checks: line items add up to the subtotal, and subtotal plus tax equals the total.
  3. Lookup checks: the supplier exists in your records, and the purchase order number matches an open order.
  4. Duplicate checks: the same supplier and invoice number haven't been processed before.
  5. Sanity checks: dates aren't in the future, amounts fall in a normal range for that supplier, and the currency matches what the supplier usually bills in.

A document that fails any rule goes to a person. A document that passes all of them is a candidate for automatic processing.

Step 4: Design human review on purpose

Human review isn't a sign the automation failed. It's what makes it safe. The goal is for people to spend their time on the documents that need judgment, not on retyping.

  • Show the original document side by side with the extracted fields, with problem fields highlighted.
  • Let reviewers correct a field in one click rather than re-entering the whole record.
  • Route by risk: high amounts, new suppliers and changed bank details always get a human look, however confident the system is.
  • Record every correction. Those corrections are your best test data.

Be careful with model "confidence scores". A model saying it's confident is not the same as being right. Rule checks and comparison with your own records are much better signals for what can skip review.

Step 5: Measure accuracy honestly

Before going live, build a test set: a few dozen real documents from different suppliers and formats, with the correct values filled in by hand. Run the pipeline against it and measure accuracy per field, not per document.

  • Field accuracy tells you which fields are weak. Totals may be near perfect while line-item descriptions are not.
  • The share of documents that pass every check tells you how much work is actually automated.
  • Errors that passed all checks are the dangerous ones. Look at each one and add a rule if you can.

Re-run the test set whenever you change the prompt, the model or the validation rules. A change that fixes one supplier's invoices can quietly break another's.

Start small, then widen

Begin with one document type and a handful of high-volume suppliers. Run the system alongside your current process for a few weeks, compare results, then let documents that pass every check flow through automatically. Expand to new suppliers and document types one at a time.

Two practical cautions. Invoices often contain personal and financial data, so check how your model provider handles that data before sending it. And watch for changed bank details on an invoice, a common fraud pattern that no amount of extraction accuracy will catch on its own.

Deeraf builds document processing pipelines like this as a Build Sprint, with validation, a review screen and a test set included so you can see the accuracy before you rely on it.

Keep reading

Want a second pair of eyes on your app?

Book a Tech Check