Skip to content
DocuExtract

Verbatim grounding · 50+ scripts · Human-in-the-loop

Highest-accuracy
document extraction.
No fabricated values.

Visual field-picker. Three handwriting tiers — English, single non-Latin script, mixed multilingual — each routed to the best-in-class engine. Every value links back to its source region.

Reads: printed text · handwriting · tables · multi-column layouts · 50+ scripts including CJK · Arabic · Hindi · Tamil · Telugu · Bengali · Punjabi

500 credits/month on the free tier (~100 pages). No credit card required.

invoice_template · v3
Acme CorporationInvoice
Invoice no.
INV-2026-0042
Date
2026-06-15
Due date
2026-07-21
Bill to
Beta Industries LLC
Consulting (40 hrs)$15,000.00
Materials$617.37
Subtotal$15,617.37
VAT (19%)$2,967.30
Total$18,584.67

Template fields

  • invoice_numbertext
    1.00
  • invoice_datedate
    0.99
  • due_datedate
    0.96
  • vendor_nametext
    1.00
  • bill_totext
    0.98
  • subtotalcurrency
    0.99
  • taxcurrency
    0.97
  • totalcurrency
    0.97

Every value links back to a region in the source — and to the model + version that read it.

How it works

Three steps. One promise: nothing fabricated.

  1. 01

    Define a template

    Upload a sample document. Draw bounding boxes around the fields you want. Name them, set types (text, number, date, currency, table). Save the template.

  2. 02

    Run a batch

    Upload many similar documents — invoices, intake forms, receipts. The engine runs each through the OCR cascade, finds the values, grounds them to source regions.

  3. 03

    Review uncertainties

    Low-confidence fields surface in a review queue with the highlighted source region. Approve or correct. Export the rest as CSV or JSON — every value with its provenance.

Why it's different

What the document actually says. Not what the model guessed.

Most extraction tools either hallucinate confident-but-wrong values, charge enterprise prices for what should be commodity work, or hide the provenance of every value behind a black-box API. DocuExtract is the opposite of all three.

  • Verbatim grounding

    Every extracted value must trace back to a source span in the original document. Values that can’t be grounded are dropped or routed to review — never fabricated by the model.

  • Visual field-picker

    Draw bounding boxes on a sample. No JSON schemas to write, no regex to maintain. Anchors handle position drift between similar documents.

  • Three handwriting tiers

    Handwritten English uses a fast vision model (~5 credits/page). Single non-Latin script handwriting routes to the right engine per script — Gemini for CJK, GPT-4o for Arabic. Mixed-language handwriting runs a dual-engine pipeline for forms with two+ scripts. Pay only for the tier you actually need.

  • Human-in-the-loop, by default

    Low-confidence fields surface for review with the source region highlighted. Reviewer corrections feed back to improve future extractions on that template.

  • Built for batches

    Upload hundreds of similar documents — invoices, intake forms, receipts. The engine runs each through the OCR cascade in parallel. Track progress live; download results as CSV or JSON.

  • Auditable

    Every field carries: which OCR tier ran, which model + version, the confidence score, any human correction. Append-only audit log. Forensic-grade provenance.

  • Credit pricing that scales fairly

    A born-digital invoice costs 1 credit. A handwritten Arabic contract costs 25. Pay for what you actually run — no flat tiers that punish difficult documents or overcharge easy ones.

  • Vertical packs included

    Legal (clause extraction, citation parsing). Logistics (BOL, customs). Immigration (USCIS forms, multi-language IDs). Pick what you need — don't pay Enterprise for one feature. Healthcare (HIPAA BAA + PHI redaction) is on the Q4 2026 roadmap — join the waitlist.

What we read

Handles the documents others can't.

Printed text and born-digital PDFs are table stakes. We also read handwriting, multi-column tables, degraded scans, and 15 languagesacross three honesty tiers. We label by maturity instead of claiming "100+ languages" like most vendors do.

Document types

Printed text & PDFs

Born-digital PDFs read directly from the embedded text layer. Scanned printed text runs through Tesseract / PaddleOCR depending on script.

Handwriting

Signatures, short field values, and printed-form handwriting handled via Tier 3 vision-LLM (Qwen 2.5-VL). 70–85% accuracy on clean handwriting; cursive freeform falls back to human review.

Tables & multi-column

Single-page tables, key/value forms, multi-column layouts. Complex multi-page tables are on the roadmap (tracked in KNOWN_ISSUES).

Degraded scans

Low-resolution photos, faxed documents, stained scans. The cascade escalates to vision-LLM automatically; uncertainty routes to review queue.

Languages (honestly tiered)

Stable

Production-ready. Validated against a comprehensive fixture set.

  • English
  • Spanish
Beta

Works on clean documents. Validate on yours.

  • Chinese (Simplified)
  • Chinese (Traditional)
  • French
  • Vietnamese
  • Korean
  • Tagalog
  • Portuguese
  • Arabic (RTL)
  • Hindi
Experimental

Active development. Accuracy varies.

  • Punjabi
  • Tamil
  • Telugu
  • Bengali

Start free today. Upgrade when you scale. Bring us in when it matters.

The Free tier covers 50 documents/month — enough for a real evaluation, not just a toy. Paid plans scale from $99/mo to $4,999/mo. Hands-on consulting available for custom integrations.