Subsystem 01
Document ingestion gateway
PDF upload, rasterization and virus scanning
Typical stack
FastAPI + Amazon S3
Reference architecture
Turn large volumes of PDFs, scans and invoices into validated relational data with hybrid OCR and LLM extraction, sending low-confidence fields to people.
Design constraints
Targets for the scenario this reference is sized for. A real engagement starts by replacing them with your own numbers.
Component topology
Subsystems with separate responsibilities, clear contracts between them and storage that scales on its own. The stack named for each is typical, not mandatory.
Stack topology
Automated document intelligence and OCR extraction pipeline
Illustrative reference architecture
Document ingestion gateway
PDF upload, rasterization and virus scanning
FastAPI + Amazon S3
OCR & layout engine
Text coordinates and bounding boxes from scanned pages
Amazon Textract / Google Document AI
LLM extraction & schema normalizer
Maps extracted text into strict Pydantic JSON schemas
Claude API with schema-validated output
Human verification queue
Side-by-side review UI with bounding-box highlights
Next.js + PDF.js
Subsystem 01
PDF upload, rasterization and virus scanning
Typical stack
FastAPI + Amazon S3
Subsystem 02
Text coordinates and bounding boxes from scanned pages
Typical stack
Amazon Textract / Google Document AI
Subsystem 03
Maps extracted text into strict Pydantic JSON schemas
Typical stack
Claude API with schema-validated output
Subsystem 04
Side-by-side review UI with bounding-box highlights
Typical stack
Next.js + PDF.js
Data lifecycle
A user or an email bot uploads a multi-page invoice PDF to a secure S3 bucket.
A Lambda dispatcher runs Amazon Textract to produce text coordinate blocks and table arrays.
A layout-aware prompt builder passes the OCR text to the LLM with a strict Pydantic JSON schema.
The LLM returns validated JSON (vendor, tax ID, line items, totals) with a confidence score per field.
Fields above a per-field threshold (such as 95% confidence) sync to the ERP (NetSuite); the rest go to the human verification queue.
Reliability and resilience
Failure mode 01
Mitigation
Preprocess images (deskew, contrast enhancement, binarization) before OCR.
Failure mode 02
Mitigation
Deterministic arithmetic checks (line items must sum exactly to the total) flag any discrepancy.
Failure mode 03
Mitigation
Chunk by logical section, extract each section in parallel, then merge.
Questions
Template OCR breaks whenever a vendor changes its invoice layout. An LLM reads the document's meaning, so it finds the right fields across layout changes, and the validation rules catch the cases where it doesn't.
The review interface shows the PDF beside the extracted fields and highlights each field's bounding box on the original scan, so an operator can confirm or correct it at a glance.
Send us your requirements, expected load and budget. We'll reply within one business day with an honest read on the design, and on whether we're the right team to build it.