Your documents, read in your perimeter — not in someone's cloud.
A small model fine-tuned on your document types extracts structured data from invoices, contracts, and application forms — running on-prem or in your VPC, at a fraction of per-page API pricing, with per-field accuracy measured before rollout.
On-prem document processing runs a compact vision-language model — 2–8B class, fine-tuned on your specific document types — inside your network, turning invoices, contracts, and application forms into validated JSON for your accounting, ERP, or document management system. In distil labs' published work with Cerebrium, moving extraction to fine-tuned small models cut inference costs by about 50% (vendor-reported) — and no page ever leaves your perimeter.
Manual re-keying, or a compliance conversation nobody wants.
Someone in your back office is still re-keying supplier invoices into the accounting system, and the alternative on offer is worse: cloud document AI that charges per page and requires shipping client contracts to a third party's servers. For firms whose clients — or regulators — forbid exactly that, both options fail.
Per-page math, from public list prices, August 2026. Frontier multimodal APIs price document extraction at roughly $0.01–0.03 per page in tokens; at 20K pages a month the bill climbs with volume — and every page transits a third party. A fine-tuned small model on a single GPU processes the same volume for infrastructure cost on the order of $0.001 per page at steady load (our cost model). In distil labs' published work with Cerebrium, moving extraction workloads to fine-tuned small models cut inference costs by about 50% (vendor-reported) — and self-hosting compounds that at volume.
When privacy, not cost, is the trigger. For legal, healthcare, and finance back offices, the economics are often secondary: the constraint is that client documents cannot be processed by a third party at all. An owned model inside your perimeter is the version of document AI your engagement letters already permit.
| Cloud document AI | Owned extraction model | |
|---|---|---|
| Per-page cost | $0.01–0.03 per page (public list prices, Aug 2026) | ~$0.001 per page at steady load — our cost model |
| Cost shape | Per page, forever — the meter runs with volume | One-time build plus a single GPU box |
| Data path | Every page transits a third party's servers | No page leaves your network |
| Published result | — | ~50% lower inference cost — distil labs × Cerebrium (vendor-reported) |
Below roughly 10K pages a month with no confidentiality constraint, a cloud API is cheaper than a build — and The Crit will tell you so with your numbers in it.
The pipeline. Inside your walls.
- Target schemas per document type, agreed with your ops team
- Labeled eval set from your real documents — the accuracy contract, yours to keep
- Fine-tuned extraction model and weights — you own them
- Deployment: on-prem hardware spec or VPC infrastructure-as-code, plus intake (folder, inbox, API) and export into your accounting, ERP, or DMS
- Review flow for low-confidence fields, and a monthly per-field accuracy report
Can small models read scanned documents and PDFs accurately?
Yes — compact vision-language models in the 2–8B range handle scans, photos, and PDFs, and fine-tuning them on your specific document types is what makes them reliable: your suppliers' layouts, your forms, your field conventions. We measure per-field accuracy on a frozen eval set built from your real documents before rollout, and low-confidence fields are flagged for human review rather than guessed.
Do our documents leave our network?
No. The model runs on-prem or in your VPC, watches your folder, inbox, or API, and writes structured output directly into your systems. No third-party processor touches the documents, no vendor trains on your data, and your existing security controls keep applying because the whole system lives inside your perimeter.
What does AI document processing cost compared to cloud APIs?
Cloud multimodal APIs run roughly $0.01–0.03 per page (public list prices, August 2026), scaling forever with volume; a fine-tuned small model on your own hardware processes pages for on the order of $0.001 each at steady load (our cost model), with cost concentrated in a one-time build. In distil labs' published work with Cerebrium, moving to fine-tuned small models cut inference costs by about 50% (vendor-reported). Below roughly 10K pages a month without a privacy constraint, the cloud API usually wins — we'll tell you which side of that line you're on.
Cut the bill. Keep the custody.
- Data stays in your VPC
- No new subprocessor
- Escalation under your rules
- Keys and logs are yours
Paper in. JSON out. Nothing leaves.
A document census and a labeled eval set tell you the accuracy a model reaches on your paperwork — and what it takes to run it inside your walls. The numbers are yours either way.