Every claim arrives as documents. Extraction is a small-model job.
A 3–7B model fine-tuned on your claim files reads FNOL emails, loss runs and supporting documents into the structured fields your claims system needs — on your infrastructure, at a fraction of frontier API cost, with low-confidence extractions routed to a human.
Insurance claims document processing runs FNOL intake, loss runs and supporting documents through a 3–7B model fine-tuned on your own claim files, deployed on your infrastructure. The model returns validated JSON mapped to your claims platform, routes low-confidence extractions to a human queue, and keeps claimant PII and medical records off third-party APIs. Published small-model migrations cut inference costs 66–80% (Forethought, on AWS).
Claims intake is a document problem before it's a decisions problem.
FNOL emails, ACORD forms, loss runs, adjuster notes and medical reports arrive in formats your claims system can't read. Manual keying is slow; a frontier API means paying open-ended reasoning prices for a closed, repetitive task — and sending claimants' personal and medical data to a third party. Extraction from known document types is precisely the narrow work where a fine-tuned small model wins.
A worked example, our assumptions stated. A carrier or MGA processing 10,000 claims a month, at 15–25K tokens of documents per claim, runs 150–250M tokens a month. The Crit recomputes this on your actual claim mix before you commit to anything.
| Frontier API | Fine-tuned model, yours | |
|---|---|---|
| Cost shape | Per token, linear with claim volume | Flat GPU infrastructure — marginal cost per claim is a rounding error |
| 10,000 claims a month | A five-figure monthly API bill at public list prices, August 2026 | $2,000–5,000/mo GPU infrastructure — our cost model |
| Claimant PII and medical records | Transit a third-party API | Stay on your hardware |
| Accuracy on your document types | General model guessing at your forms | Trained on your back catalogue; measured field by field before go-live |
Published benchmarks, not our claims. Forethought reports 66–80% lower inference costs from moving routine workloads to fine-tuned small models (published on AWS). distil labs reports 50–68% cost reduction with accuracy rising from 81% to 93% at Knowunity, and ~50% savings with Cerebrium (vendor-reported). We haven't shipped this exact vertical yet; the published evidence says the economics hold, and we prove it on your claims before you commit.
If your claim volume doesn't clear the threshold where ownership beats the API — or your documents are genuinely novel every time, with no back catalogue to train on — we'll say so in the audit, and you keep the numbers. And extraction is not adjudication: coverage decisions stay with your adjusters, by design.
The pipeline, measured. Yours.
- A fine-tuned extraction model for your document types, returning validated JSON mapped to your claims platform's fields
- A claim-type and severity classifier for queue routing
- A frozen eval set from your historical claims; field-level accuracy measured before go-live
- A human-review queue for low-confidence extractions, with thresholds your ops team controls
- Deployment on your hardware or private cloud — no claimant data in third-party APIs
- Production monitoring and retraining under an Assurance retainer
Can AI extract data from insurance claims documents accurately?
On known document types with a defined field schema — yes, and this is where fine-tuned small models specifically beat general-purpose ones: the task is narrow and the training data is your own back catalogue of claims. We don't ask you to take that on faith; we measure accuracy field by field against a frozen eval set of your historical claims before the system touches production, and low-confidence extractions go to a human queue.
Do claimant medical records leave our infrastructure?
No. The extraction and classification models run on your hardware or in your private cloud tenancy. There is no third-party API in the path, which keeps your data-protection assessment short and your subprocessor list unchanged.
What happens when document formats change?
The system tells you before it fails silently: we monitor production accuracy against the eval set, and drift — a new ACORD version, a new loss-run layout, a new line of business — triggers retraining under the Assurance retainer. Format change is an operations event, not an outage.
Special-category data stays home.
- Medical records never leave
- No external AI vendors in claims
- Human review lanes
- Full audit trail
We’re engineers, not your compliance advisers — the architecture keeps your existing compliance intact instead of adding a vendor to it.
Your claim files already wrote the spec.
Two weeks inside one document flow — FNOL or loss runs — and you'll know field-level accuracy, projected cost per claim, and whether ownership beats the API at your volume. The numbers are yours either way.