Screening alerts are routine work. Stop paying frontier prices for them.
A small model fine-tuned on your historical adjudications triages sanctions, PEP and adverse-media alerts inside your VPC — cheaper per alert, with customer PII never leaving your infrastructure, and every decision logged for audit.
KYC and sanctions-screening triage runs on a 3–7B model fine-tuned on your own historical adjudications, deployed inside your VPC. The model clears routine false positives with a structured rationale, escalates ambiguous alerts to a human queue, and writes a per-decision audit log. Published migrations of routine workloads to small models cut inference costs 66–80% (Forethought, on AWS) — and no customer PII reaches a third party.
Most alerts are false positives. Someone still clears every one.
An analyst reads the match, checks the entity and writes a rationale — thousands of times a month. Routing that work through a frontier API means paying general-purpose prices for a narrow, repetitive task, and sending customer PII to a third party your DPA and your regulator both have questions about. The task is exactly what small models are built for: narrow, high-volume, and defined by your own historical decisions.
A worked example, our assumptions stated. A mid-size fintech clearing ~10,000 alerts a month at 8–12K tokens per adjudication runs roughly 100M tokens a month. The Crit recomputes every line of this on your actual volumes before you sign anything.
| Frontier API | Fine-tuned small model, yours | |
|---|---|---|
| Cost shape | Per token, linear with alert volume | Flat VPC compute — marginal cost per alert is a rounding error |
| ~10,000 alerts a month | Tens of thousands of dollars monthly at public list prices, August 2026 | $1,500–4,000/mo GPU infrastructure — our cost model |
| Customer PII | Processed by a third party | Never leaves your VPC |
| Audit trail | Vendor logs, on request | Per-decision: input, evidence, rationale, model version |
Published benchmarks, not our claims. Forethought reports 66–80% lower inference costs after replacing routine frontier calls with fine-tuned small models (published on AWS). distil labs reports accuracy rising from 81% to 93% on a narrow task while costs fell 50–68% in their Knowunity work, and a ~50% cost reduction with Cerebrium (both vendor-reported). Gartner expects task-specific small models to be adopted three times more than general-purpose LLMs by 2027.
Below roughly 50M tokens a month with no privacy constraint, a frontier API is probably fine — and The Crit will say so in writing. And no model replaces your compliance judgement: ambiguous alerts escalate to humans by design, and the thresholds stay in your team's hands.
The model, the evals, the trail. Yours.
- A fine-tuned alert-adjudication model deployed in your VPC, with structured JSON output mapped to your case-management fields
- A frozen eval set from your historical decisions; precision and recall measured against it before deployment
- An escalation router with thresholds your compliance team sets — ambiguous alerts go to a human, not to a guess
- A per-decision audit log designed for MLRO and regulator requests
- KYC document extraction — onboarding documents to structured fields — as an optional second workflow
- Monitoring, drift detection and re-training under an Assurance retainer
Can we use an LLM for KYC and sanctions screening without sending customer data to a third party?
Yes — by running the model inside your own infrastructure. A fine-tuned 3–7B open-weight model deployed in your VPC processes the alert where the data already lives. No third-party processor joins your stack, your existing security controls keep applying, and the data-residency answer stays one sentence long.
Is a small model accurate enough for compliance decisions?
On a narrow, well-defined task, a fine-tuned small model typically beats a general-purpose model guessing at it — distil labs' published Knowunity work measured accuracy rising from 81% to 93% (vendor-reported). More importantly, we don't ask you to trust the benchmark: we measure the model against a frozen eval set built from your own historical adjudications before it touches production, and ambiguous cases escalate to humans by design.
What does the regulator see when they ask about a cleared alert?
A complete trail: the alert input, the evidence the model retrieved, the structured rationale it produced, the model version that produced it, and the reviewing analyst where a human made the call. The system exists so the answer to "show us your process" is an export, not a reconstruction.
PII never crosses your boundary.
- Screening runs in your VPC
- No new third party to file
- Frozen, versioned models
- A per-decision audit log
We’re engineers, not your compliance advisers — the architecture keeps your existing compliance intact instead of adding a vendor to it.
Your adjudications are the training data.
Two weeks inside your screening stack: cost per cleared alert, projected savings on your volumes, and a frozen eval set built from your own decisions. The numbers are yours either way.