Use case · Back-office documents

Your documents, read in your perimeter — not in someone's cloud.

A small model fine-tuned on your document types extracts structured data from invoices, contracts, and application forms — running on-prem or in your VPC, at a fraction of per-page API pricing, with per-field accuracy measured before rollout.

On-prem document processing runs a compact vision-language model — 2–8B class, fine-tuned on your specific document types — inside your network, turning invoices, contracts, and application forms into validated JSON for your accounting, ERP, or document management system. In distil labs' published work with Cerebrium, moving extraction to fine-tuned small models cut inference costs by about 50% (vendor-reported) — and no page ever leaves your perimeter.

Manual re-keying, or a compliance conversation nobody wants.

Someone in your back office is still re-keying supplier invoices into the accounting system, and the alternative on offer is worse: cloud document AI that charges per page and requires shipping client contracts to a third party's servers. For firms whose clients — or regulators — forbid exactly that, both options fail.

How it works · The Crit Path
01Census — schema and eval setWe sample your real document types — invoices, contracts, application forms — define the target schema (which fields, which formats), and build an eval set from actual documents with human-verified labels. This set is the accuracy contract for everything that follows.
02Fine-tune — the extraction modelA compact vision-language model (2–8B class, Qwen-VL family) for scans and photos; a text SLM for digital documents. Trained on your document types specifically — your suppliers' invoice layouts, your intake forms — not documents in general.
03Deploy — where the documents liveA single GPU box on-prem, or a modest instance in your VPC. It watches a folder, an inbox, or an API endpoint; outputs validated JSON straight into your accounting, ERP, or document management system. No page ever leaves your network.
04Assure — with a human in the loopLow-confidence fields are flagged for review rather than guessed; corrections feed re-training. You get a monthly per-field accuracy report against the frozen eval set.
The economics

Per-page math, from public list prices, August 2026. Frontier multimodal APIs price document extraction at roughly $0.01–0.03 per page in tokens; at 20K pages a month the bill climbs with volume — and every page transits a third party. A fine-tuned small model on a single GPU processes the same volume for infrastructure cost on the order of $0.001 per page at steady load (our cost model). In distil labs' published work with Cerebrium, moving extraction workloads to fine-tuned small models cut inference costs by about 50% (vendor-reported) — and self-hosting compounds that at volume.

When privacy, not cost, is the trigger. For legal, healthcare, and finance back offices, the economics are often secondary: the constraint is that client documents cannot be processed by a third party at all. An owned model inside your perimeter is the version of document AI your engagement letters already permit.

Cloud document AIOwned extraction model
Per-page cost$0.01–0.03 per page (public list prices, Aug 2026)~$0.001 per page at steady load — our cost model
Cost shapePer page, forever — the meter runs with volumeOne-time build plus a single GPU box
Data pathEvery page transits a third party's serversNo page leaves your network
Published result~50% lower inference cost — distil labs × Cerebrium (vendor-reported)
Where this doesn't work

Below roughly 10K pages a month with no confidentiality constraint, a cloud API is cheaper than a build — and The Crit will tell you so with your numbers in it.

What you get

The pipeline. Inside your walls.

  • Target schemas per document type, agreed with your ops team
  • Labeled eval set from your real documents — the accuracy contract, yours to keep
  • Fine-tuned extraction model and weights — you own them
  • Deployment: on-prem hardware spec or VPC infrastructure-as-code, plus intake (folder, inbox, API) and export into your accounting, ERP, or DMS
  • Review flow for low-confidence fields, and a monthly per-field accuracy report
Who this is for
SMB back offices buried in paperThousands of supplier invoices, contracts, or applications a month, still processed by hand.
Firms whose documents can't go outClients or regulators forbid third-party processing — legal, healthcare, financial services.
Ops teams watching the per-page meterAlready paying cloud document AI per page, with volume climbing every quarter.
Questions, answered straight

Can small models read scanned documents and PDFs accurately?

Yes — compact vision-language models in the 2–8B range handle scans, photos, and PDFs, and fine-tuning them on your specific document types is what makes them reliable: your suppliers' layouts, your forms, your field conventions. We measure per-field accuracy on a frozen eval set built from your real documents before rollout, and low-confidence fields are flagged for human review rather than guessed.

Do our documents leave our network?

No. The model runs on-prem or in your VPC, watches your folder, inbox, or API, and writes structured output directly into your systems. No third-party processor touches the documents, no vendor trains on your data, and your existing security controls keep applying because the whole system lives inside your perimeter.

What does AI document processing cost compared to cloud APIs?

Cloud multimodal APIs run roughly $0.01–0.03 per page (public list prices, August 2026), scaling forever with volume; a fine-tuned small model on your own hardware processes pages for on the order of $0.001 each at steady load (our cost model), with cost concentrated in a one-time build. In distil labs' published work with Cerebrium, moving to fine-tuned small models cut inference costs by about 50% (vendor-reported). Below roughly 10K pages a month without a privacy constraint, the cloud API usually wins — we'll tell you which side of that line you're on.

Security · by architecture

Cut the bill. Keep the custody.

  • Data stays in your VPC
  • No new subprocessor
  • Escalation under your rules
  • Keys and logs are yours
Architected to deploy inside your
SOC 2 GDPR ISO 27001
Small models · critical hits

Paper in. JSON out. Nothing leaves.

A document census and a labeled eval set tell you the accuracy a model reaches on your paperwork — and what it takes to run it inside your walls. The numbers are yours either way.