Skip to content

Reference architecture

Automated document intelligence and OCR extraction pipeline

Turn large volumes of PDFs, scans and invoices into validated relational data with hybrid OCR and LLM extraction, sending low-confidence fields to people.

Design constraints

What the design has to hold to

Targets for the scenario this reference is sized for. A real engagement starts by replacing them with your own numbers.

  • 01Field-level accuracy target set per document type and measured on a labeled set (for example 99% on totals)
  • 02100-page financial filings processed in under 15 seconds
  • 03Confidence scoring that routes low-confidence fields to human reviewers
  • 04PII and sensitive financial data redacted before any cloud model sees it

Component topology

System components and technologies

Subsystems with separate responsibilities, clear contracts between them and storage that scales on its own. The stack named for each is typical, not mandatory.

Stack topology

Automated document intelligence and OCR extraction pipeline

Illustrative reference architecture

  1. 01

    Document ingestion gateway

    PDF upload, rasterization and virus scanning

    FastAPI + Amazon S3

  2. 02

    OCR & layout engine

    Text coordinates and bounding boxes from scanned pages

    Amazon Textract / Google Document AI

  3. 03

    LLM extraction & schema normalizer

    Maps extracted text into strict Pydantic JSON schemas

    Claude API with schema-validated output

  4. 04

    Human verification queue

    Side-by-side review UI with bounding-box highlights

    Next.js + PDF.js

Subsystem 01

Document ingestion gateway

PDF upload, rasterization and virus scanning

Typical stack

FastAPI + Amazon S3

Subsystem 02

OCR & layout engine

Text coordinates and bounding boxes from scanned pages

Typical stack

Amazon Textract / Google Document AI

Subsystem 03

LLM extraction & schema normalizer

Maps extracted text into strict Pydantic JSON schemas

Typical stack

Claude API with schema-validated output

Subsystem 04

Human verification queue

Side-by-side review UI with bounding-box highlights

Typical stack

Next.js + PDF.js

Data lifecycle

End-to-end data flow

  1. A user or an email bot uploads a multi-page invoice PDF to a secure S3 bucket.

  2. A Lambda dispatcher runs Amazon Textract to produce text coordinate blocks and table arrays.

  3. A layout-aware prompt builder passes the OCR text to the LLM with a strict Pydantic JSON schema.

  4. The LLM returns validated JSON (vendor, tax ID, line items, totals) with a confidence score per field.

  5. Fields above a per-field threshold (such as 95% confidence) sync to the ERP (NetSuite); the rest go to the human verification queue.

Reliability and resilience

Failure modes and how each is contained

Failure mode 01

Handwritten or blurry scans

Mitigation

Preprocess images (deskew, contrast enhancement, binarization) before OCR.

Failure mode 02

Invented line-item numbers

Mitigation

Deterministic arithmetic checks (line items must sum exactly to the total) flag any discrepancy.

Failure mode 03

Token limits on 500-page documents

Mitigation

Chunk by logical section, extract each section in parallel, then merge.

Questions

What teams ask about this design

Template OCR breaks whenever a vendor changes its invoice layout. An LLM reads the document's meaning, so it finds the right fields across layout changes, and the validation rules catch the cases where it doesn't.

The review interface shows the PDF beside the extracted fields and highlights each field's bounding box on the original scan, so an operator can confirm or correct it at a glance.

Planning a system like this?

Send us your requirements, expected load and budget. We'll reply within one business day with an honest read on the design, and on whether we're the right team to build it.