AI Document Extraction: How to Extract Structured Data From Documents
How to extract structured data from documents with AI: OCR and layout parsing, templates versus ML versus language and vision models, schema design, structured outputs, tables, confidence, validation and evaluation.
Quick answer
AI document extraction turns documents into structured data in four steps: get the text and layout (native PDF text, OCR or a vision model), extract fields against an explicit schema (with templates, trained models or language and vision models using structured outputs), validate every field (formats, totals, cross-checks, nulls instead of guesses) and route low-confidence or failed fields to human review. Modern language and vision models handle varied layouts well, but they can invent values, so validation and field-level evaluation on your own documents are essential.
Where This Fits
This guide covers the extraction technique. The end-to-end business pipeline (ingestion, review, integration, storage) is in intelligent document processing, an applied example in AI invoice processing, and how extraction fits into automations in AI workflow automation.
Step 1: Get Text and Layout
Native digital PDFs usually contain text you can extract directly, though reading order and tables may need layout analysis. Scans and photos need OCR, ideally layout-aware so tables, columns and key-value pairs are preserved. Vision-capable models can read page images directly, which helps with complex layouts, stamps and handwriting, at a higher cost per page. Keep page and position references so reviewers can see where each value came from.
Step 2: Choose an Extraction Method
| Method | How it works | Best for |
|---|---|---|
| Templates and rules | Fixed positions or anchors per layout | Few, stable layouts at high volume |
| Trained ML extraction | Models trained on labelled examples | Common document types with labelled data |
| Language model on text | Schema prompt over OCR or PDF text | Varied layouts, text-heavy documents |
| Vision-language model | Schema prompt over page images | Complex layouts, tables, stamps, handwriting |
| Hybrid | Different methods per document type | Mixed real-world inputs |
Step 3: Design the Schema
The schema is the contract between the document and your systems. Use clear names and types, formats (dates, currency codes), enums for categorical values, and arrays for line items. Make missing values explicitly nullable and instruct the model to return null rather than guess. Include a source quote or page reference per field where your provider supports it, to make review and validation easier.
{
"type": "object",
"properties": {
"delivery_note_number": { "type": "string" },
"delivery_date": { "type": ["string", "null"], "format": "date" },
"supplier_name": { "type": "string" },
"po_number": { "type": ["string", "null"] },
"lines": {
"type": "array",
"items": {
"type": "object",
"properties": {
"sku": { "type": ["string", "null"] },
"description": { "type": "string" },
"quantity": { "type": "number" },
"unit": { "type": "string", "enum": ["each", "box", "pallet", "kg", "other"] }
},
"required": ["sku", "description", "quantity", "unit"],
"additionalProperties": false
}
}
},
"required": ["delivery_note_number", "delivery_date", "supplier_name", "po_number", "lines"],
"additionalProperties": false
}Step 4: Use Structured Outputs
Major model providers can constrain output to a JSON schema (OpenAI's structured outputs and Anthropic's structured outputs, for example), which eliminates parsing failures. It does not guarantee correct values, so the next step matters more.
Extracting data from documents your systems cannot read?
ZSpace Labs builds extraction pipelines with schemas, validation and review screens, tuned on your own documents.
Step 5: Validate Every Field
- Format checks: dates, IDs, currency codes, check digits
- Arithmetic: line items sum to totals; tax matches rates
- Cross-checks against system records: PO exists, supplier matches
- Presence checks: required fields not null without a reason
- Source checks: extracted value appears in the document text
- Confidence routing: send uncertain fields to review
Evaluating Extraction
Label a representative sample of documents, including poor scans and unusual layouts. Measure accuracy per field (not just per document), separately for header fields and line items, and track how often values are invented versus missed. Compare methods and models on this set before choosing, and rerun it when prompts, models or document sources change.
Costs
Costs come from OCR or document AI services per page, model tokens (images can be token-heavy), review time and infrastructure. Reduce them by routing simple, stable documents to cheaper methods, sending only relevant pages to models, using smaller models where accuracy holds and batching offline work.
Advantages and Limitations
AI extraction handles layout variety that defeats templates and dramatically reduces manual keying. It can invent plausible values, struggles with poor scans and long multi-page tables, and costs more per page than simple OCR rules. Validation, review and evaluation turn it from impressive to dependable.
Tables and Multi-Page Documents
Line items and tables cause most extraction errors. Detect tables with layout analysis or vision models, extract them as arrays with a defined row schema, handle tables that continue across pages by carrying headers forward, and reconcile row totals with document totals. For long documents, classify pages first and send only relevant pages to extraction, which improves accuracy and cuts cost.
Tools and Services
| Category | Use for |
|---|---|
| Cloud document AI services | OCR, layout, prebuilt models for common documents |
| Open-source OCR and layout tools | Self-hosted parsing and data control |
| Language and vision model APIs with structured outputs | Flexible schema extraction across layouts |
| IDP platforms | End-to-end pipelines with review UIs |
| Validation libraries and rules engines | Field checks, cross-checks and business rules |
Worked Example
An illustrative scenario, not a client case: an insurer extracts data from repair estimates in dozens of formats. A vision-language model with a strict schema returns line items and totals; validation checks arithmetic and that each part number appears in the OCR text; mismatches go to adjusters with the relevant region highlighted. Field-level accuracy is tracked weekly by repair shop to spot problem formats.
Common Mistakes
- No nullable fields, so models guess
- Measuring per-document rather than per-field accuracy
- Trusting structured output without value validation
- Ignoring multi-page tables
- Sending whole documents when one page holds the data
Need reliable data from messy documents?
Talk to ZSpace Labs about AI document extraction and integration into your systems.
Conclusion
Good document extraction combines the right text source, a precise schema, structured outputs, rigorous validation and review for uncertain fields, measured on your own documents. Related: intelligent document processing and AI invoice processing.
Common questions
Turning the content of documents such as PDFs, scans and images into structured fields and tables that software can use, using OCR, layout analysis and machine learning or language and vision models.