Healthcare · Mortgage · Insurance · Tax

Verifiable document AI for regulated industries.

Foundational models and domain-specialized agents that read the fields, tables, and checkboxes in your most complex documents , every value source-linked and scored, so your team verifies it before it ships.

Every field verifiable · Human-in-the-loop review · Runs on your infrastructure

Built for complex, high-stakes documents
ClaimsLoan filesPoliciesTax formsEOBs
Other AI guesses. Ours proves it.

General-purpose agents hallucinate on the documents your business actually runs on, then leave you to discover where they went wrong.

Our models are purpose-built for document understanding, and every value they return is inspectable against the source. So accuracy isn't a slogan, it's something your team can check , field by field , before it reaches a downstream system.

What you get

Agents that go past OCR , to your schema

Reading the page is table stakes. dextract returns data shaped and validated the way your systems need it.

Beyond OCR

Extract to your schema, not just to text

Define the fields you need and the rules they must satisfy , types, formats, cross-field checks. Every extraction comes back shaped and validated to your spec, not as a wall of raw text.

date_of_birth : date "MM DD YY" # a real date, not raw text
npi : text /^\d{10}$/ # bad IDs caught, never passed on
total = sum(line_charges) # totals must reconcile
rel_to_insured ∈ {self, spouse, child} → self # only valid options, mapped to yours
Verifiable

Proven, field by field

Every value is source-linked and confidence-scored, so you check it, not just trust it.

Private

On your hardware, not a vendor's cloud

Runs inside your compliance boundary. Nothing a document contains ever leaves your network.

Control

Auto-accept the clean, review the rest

Gate on confidence: straight-through for the obvious fields, a short queue for the few that need a human.

Coverage

Fields, tables, and checkboxes

Key–value pairs, dense repeating tables, and marked selections , all in the shape your systems expect.

FieldsTablesCheckboxes
The predictions

Every field carries its proof

Location and confidence on every value , trust the clean ones, route the flagged few to review. Click any field to see how it was verified.

extraction · healthcare claim · 1 flagged
Fields1 to review
▸ click a field for its proof
patient_nameRoberts, Marie H 1.00
date_of_birth09 26 1967 1.00
insured_id123-45-6789 1.00
diagnosis_aK52.9 0.89 review
total_charge$185.67 1.00
Table · service lines1 of 6
▸ click a row for its proof
dateplacecptunitscharge
03 07 2011767051185.67

lines 2–6 blank on this claim , empty rows are reported, never invented.

A real, complex CMS-1500

Read right on the parts that trip up other AI

AI document agents can feel like black magic. Here's ours with the lights on , real crops from a filled claim, every prediction located on the page and independently checked.

birth date and sex, detected
A date and a checkbox, together
date_of_birth1967-09-26
sexF
✓ located✓ date parsed✓ checkbox readconf 1.00
relationship to insured checkboxes, detected
One of four boxes, marked
relationship_to_insuredSelf
✓ 1 of 4 marked✓ located✓ ink + model confirmedconf 0.97
service line table row, detected
A dense table row, cell by cell
date · cpt · charge03-07-20 · 76705 · 185.67
place · pointer · units11 · ABC · 1
✓ 8 cells✓ row reconciled✓ locatedconf 0.99
The team

Built by people who've shipped document AI at scale

dextract is built by AI scientists from Amazon and other top AI labs , the people who've put document extraction into production for some of the largest workloads in the world.

AmazonTop AI labsAI scientistsEnterprise scale
Pricing

Start free. Pay per document. Then stop counting documents.

Paying by the page is fine while the volume is small. Once it is not, you move onto hardware you already pay for and the bill stops tracking how much work you do.

Trial
Free

Bring your own files and see what comes back.

  • 50 documents, 200 pages
  • Fields, tables, and checkboxes
  • Every value source-linked
  • Test data only, no PHI
Start now
Cloud
Per document

Keep running in our cloud and pay for what you actually process.

  • Priced per document and page
  • Failed extractions are not charged
  • Your schema and validation rules
  • Usage visible in your dashboard
Talk to us
Self-hosted
Annual licence

Runs inside your boundary: your cloud account, a box we ship, or your own hardware.

  • Nothing a document contains leaves your network
  • Air-gap capable
  • Priced by capacity, not per document
  • Named support contact
Talk to us

On the licence, volume does not change the price. Capacity does, so the way to spend more with us is to process a great deal more, on more hardware.

Security & deployment

What to demand from document AI

The controls regulated teams need , and where a typical cloud document API leaves you exposed.

CapabilitydextractCloud document APIs
Runs fully on your hardware (on-prem)
Air-gap capable , nothing leaves your network
Source-linked, field-level verificationpartial
Confidence + human-in-the-loop routingpartial
Your schema + validation rules enforcedpartial
Reproducible, audit-ready resultspartial
No per-page cloud fees

Bring us your hardest documents.

Send a stack of your most complex files. We'll show you clean, verifiable extraction , running on your own infrastructure.