Models your risk team can put in a file.
Small, fine-tuned models for KYC screening, onboarding extraction, and transaction classification — deployed inside your perimeter, gated behind a frozen eval set your second line can actually review.
Financial-services AI works when it can be governed. A fine-tuned small model runs in your VPC or on-prem, stays versioned and frozen so it doesn't change under you, and ships only after it clears a frozen eval set your risk team reviews. Published migrations report 66–80% lower inference costs on routine workloads (Forethought, on AWS) — and the ambiguous cases escalate by policy you set.
Financial services runs on exactly the kind of work small models are good at: reading documents, checking them against rules, classifying transactions, extracting fields from onboarding packets. These are narrow, high-volume tasks with definable right answers — which means they can be measured, and what can be measured can be governed. That's the part frontier APIs make hard: a third-party model you can't inspect, on infrastructure you don't control, with behavior that changes when the vendor ships an update.
A fine-tuned small model inverts each of those properties. It runs in your VPC or on-prem, so no customer data crosses your boundary. It's versioned and frozen, so it doesn't change under you. And because we gate every deployment behind a fixed eval set, your risk and compliance functions get an artifact — measured accuracy on your data, per document type — instead of a vendor's assurances. Genuinely ambiguous cases escalate to a frontier model or a human, by policy you set.
We haven't shipped inside a bank or fintech yet, and we won't pretend otherwise. The published evidence below says the economics hold; The Crit is how we'd prove it on your documents before you commit to anything.
Narrow tasks, measured answers.
The common thread across FCA expectations in the UK, DORA in the EU, and model-risk practice in the US is the same: know your third parties, know your models, and be able to show your work. A small model deployed on your infrastructure simplifies all three at once — there's no new external processor in the data flow, and the frozen eval set plus production monitoring give your model-risk documentation something concrete to reference. We're engineers, not your compliance advisers; what we provide is a system your compliance function can examine.
We haven't shipped this vertical yet. Here is the published evidence the economics hold — and here is how we'd prove it on your documents, against a frozen eval set, before you commit.
| Published result | Source |
|---|---|
| Inference costs fell 66–80% after moving routine workloads to fine-tuned small models | Forethought, published on AWS |
| Task accuracy rose from 81% to 93% while inference costs fell 50–68% | distil labs × Knowunity, vendor-reported |
| Serving costs fell roughly 50% on dedicated small-model infrastructure | distil labs on Cerebrium, vendor-reported |
| Task-specific small models to see 3× the adoption of general-purpose LLMs by 2027 | Gartner forecast |
How does a small model fit our model-risk framework?
Better than a rented one. The model is versioned and frozen, so it doesn't change when a vendor ships an update, and every deployment carries a fixed eval set with measured accuracy on your data, per document type. Your model-risk documentation references an artifact your second line can examine — not a vendor's assurances.
Does customer data leave our infrastructure?
No. The model runs in your VPC or on-prem, so no customer data crosses your boundary and no new external processor enters the data flow. Genuinely ambiguous cases escalate to a frontier model or a human — by policy you set, not a vendor's default.
Have you deployed inside a bank or fintech?
Not yet, and we won't pretend otherwise. The published evidence — 66–80% lower inference costs (Forethought, on AWS), 81% to 93% accuracy with costs down 50–68% (distil labs × Knowunity, vendor-reported) — says the economics hold. The Crit proves it on your documents, against a frozen eval set, before you commit to anything.
PII never crosses your boundary.
- Screening runs in your VPC
- No new third party to file
- Frozen, versioned models
- A per-decision audit log
We’re engineers, not your compliance advisers — the architecture keeps your existing compliance intact instead of adding a vendor to it.
See if the numbers hold in your workflows.
The Crit is a teardown of your AI spend and workflows: where a small model wins, where an API is fine, and where AI shouldn't be used at all. You get the numbers either way.
If your workload doesn't clear roughly 50M tokens a month and you have no privacy constraint, a frontier API is probably fine — and we'll tell you so.