Use case · Compliance & KYC screening

Screening alerts are routine work. Stop paying frontier prices for them.

A small model fine-tuned on your historical adjudications triages sanctions, PEP and adverse-media alerts inside your VPC — cheaper per alert, with customer PII never leaving your infrastructure, and every decision logged for audit.

KYC and sanctions-screening triage runs on a 3–7B model fine-tuned on your own historical adjudications, deployed inside your VPC. The model clears routine false positives with a structured rationale, escalates ambiguous alerts to a human queue, and writes a per-decision audit log. Published migrations of routine workloads to small models cut inference costs 66–80% (Forethought, on AWS) — and no customer PII reaches a third party.

Most alerts are false positives. Someone still clears every one.

An analyst reads the match, checks the entity and writes a rationale — thousands of times a month. Routing that work through a frontier API means paying general-purpose prices for a narrow, repetitive task, and sending customer PII to a third party your DPA and your regulator both have questions about. The task is exactly what small models are built for: narrow, high-volume, and defined by your own historical decisions.

How it works · The Crit Path
01Scope — The Crit on your screening stackAn audit of volumes, false-positive rates and cost per cleared alert. We build a frozen eval set from your historical adjudications — the ground truth the model must beat before anything ships.
02Prove — fine-tune on your decisionsWe train a 3–7B open-weights model on how your analysts actually adjudicate — entity resolution against the match record, adverse-media summarisation, and a structured risk rationale in JSON — not on generic compliance text. It ships only if it beats the eval set.
03Deploy — inside your VPCThe model runs in a private subnet next to your data — AWS, GCP, Azure or on-prem. A router escalates genuinely ambiguous alerts either to a frontier model through a redaction layer or straight to a human queue; your compliance team chooses which.
04Assure — log every decisionModel version, input, retrieved evidence, output, reviewer — the complete trail your MLRO produces when the auditor or the regulator asks who cleared an alert and why.
The economics

A worked example, our assumptions stated. A mid-size fintech clearing ~10,000 alerts a month at 8–12K tokens per adjudication runs roughly 100M tokens a month. The Crit recomputes every line of this on your actual volumes before you sign anything.

Frontier APIFine-tuned small model, yours
Cost shapePer token, linear with alert volumeFlat VPC compute — marginal cost per alert is a rounding error
~10,000 alerts a monthTens of thousands of dollars monthly at public list prices, August 2026$1,500–4,000/mo GPU infrastructure — our cost model
Customer PIIProcessed by a third partyNever leaves your VPC
Audit trailVendor logs, on requestPer-decision: input, evidence, rationale, model version

Published benchmarks, not our claims. Forethought reports 66–80% lower inference costs after replacing routine frontier calls with fine-tuned small models (published on AWS). distil labs reports accuracy rising from 81% to 93% on a narrow task while costs fell 50–68% in their Knowunity work, and a ~50% cost reduction with Cerebrium (both vendor-reported). Gartner expects task-specific small models to be adopted three times more than general-purpose LLMs by 2027.

Where this doesn't work

Below roughly 50M tokens a month with no privacy constraint, a frontier API is probably fine — and The Crit will say so in writing. And no model replaces your compliance judgement: ambiguous alerts escalate to humans by design, and the thresholds stay in your team's hands.

What you get

The model, the evals, the trail. Yours.

  • A fine-tuned alert-adjudication model deployed in your VPC, with structured JSON output mapped to your case-management fields
  • A frozen eval set from your historical decisions; precision and recall measured against it before deployment
  • An escalation router with thresholds your compliance team sets — ambiguous alerts go to a human, not to a guess
  • A per-decision audit log designed for MLRO and regulator requests
  • KYC document extraction — onboarding documents to structured fields — as an optional second workflow
  • Monitoring, drift detection and re-training under an Assurance retainer
Who this is for
Fintechs and EMIs past analyst headcountScreening volume has outgrown the team, but DPAs and data-residency terms rule out sending PII to third-party APIs.
Teams that already hit the wallYou tried a frontier model on alert triage and hit either the bill or the data-processing question.
Compliance leaders who answer to auditorsEvery automated decision must stay explainable and reproducible — model version, evidence and rationale on file.
Questions, answered straight

Can we use an LLM for KYC and sanctions screening without sending customer data to a third party?

Yes — by running the model inside your own infrastructure. A fine-tuned 3–7B open-weight model deployed in your VPC processes the alert where the data already lives. No third-party processor joins your stack, your existing security controls keep applying, and the data-residency answer stays one sentence long.

Is a small model accurate enough for compliance decisions?

On a narrow, well-defined task, a fine-tuned small model typically beats a general-purpose model guessing at it — distil labs' published Knowunity work measured accuracy rising from 81% to 93% (vendor-reported). More importantly, we don't ask you to trust the benchmark: we measure the model against a frozen eval set built from your own historical adjudications before it touches production, and ambiguous cases escalate to humans by design.

What does the regulator see when they ask about a cleared alert?

A complete trail: the alert input, the evidence the model retrieved, the structured rationale it produced, the model version that produced it, and the reviewing analyst where a human made the call. The system exists so the answer to "show us your process" is an export, not a reconstruction.

Security · by architecture

PII never crosses your boundary.

  • Screening runs in your VPC
  • No new third party to file
  • Frozen, versioned models
  • A per-decision audit log
Built for review under
DORA FCA SYSC GDPR SOC 2

We’re engineers, not your compliance advisers — the architecture keeps your existing compliance intact instead of adding a vendor to it.

Small models · critical hits

Your adjudications are the training data.

Two weeks inside your screening stack: cost per cleared alert, projected savings on your volumes, and a frozen eval set built from your own decisions. The numbers are yours either way.