Audit. Build. Assure.
Five engagements, one method, no faith required. Every one starts with The Crit — a teardown that puts numbers on the table before you commit to a build.
Every build runs on The Crit Path: Scope the workflow, Prove the model on your data against a frozen eval set, Deploy inside your perimeter, Assure with monitoring and re-training. The Prove step is a gate, not a formality — if the model doesn't clear the pass mark you set, the project stops and you keep the artifacts. You never pay for a full build on faith.
The Crit. The teardown before you commit to anything.
In design, a crit is the honest review before anything gets built. That's the product: we take your AI workflows apart in front of you, measure them, and tell you what holds. We work read-only, under NDA — your API logs and invoices, token volumes per workflow, prompt chains, latency and privacy constraints. Sensitive data stays onsite; we can work from redacted samples. We credit the fee in full against any build that starts within 90 days.
- Workload map — every AI-touching workflow, ranked by spend and volume
- Frozen eval set — 100–300 graded examples for the top candidate workflows, yours to keep
- Benchmark pass — two or three open-weights small models scored against your eval sample; measurements, not projections
- Cost model — current spend vs. projected small-model cost, break-even point, 12-month view
- A verdict per workflow — migrate, keep the frontier API, or don't automate at all. In writing
- A fixed-price build proposal — or a written recommendation not to build. Both outcomes cost the same
- Not a strategy deck. Every page of the readout is a measurement or a decision, not a framework.
- Not a sales document in disguise. The verdicts stand alone; the build proposal is a separate attachment you're free to ignore.
- Not a free pilot, and not a discount audit padded to look thorough. Fixed scope, three weeks.
- Not a compliance certification. We assess your workflows, not your controls.
What if you find nothing worth building?
Then the readout says so, workflow by workflow, and recommends you keep what you have. Roughly, if you run under 50M tokens a month with no privacy constraint, that's the likely verdict — we'd rather put it in the readout than bill you to hear it later. The audit stands alone as a document worth having either way.
Who actually sees our data?
One senior operator, under NDA. We work from read-only access or redacted samples, we train nothing outside your project, and nothing leaves your environment without written agreement.
Model Distillation. Replace routine frontier calls with a model you own.
Most of what companies send to frontier models is routine, narrow work — classification, extraction, drafting, scoring. At 100M tokens a month, a frontier API runs roughly $50K a month; a distilled 7B model serving the same routine traffic runs $1.5K–$4K in infrastructure (our cost model, from published API and GPU list prices, August 2026). Frontier models are for the hard 20%. We build the split.
- Frozen eval set built from your real traffic — the pass mark is set here, before any training
- Training data distilled from your pipeline's outputs, validated by humans, not vibes
- A fine-tuned 3–7B open-weights model, trained on your task and nothing else
- Hitbox, our eval-and-routing harness — it gates deployment behind measured accuracy and decides, per request, whether the small model answers or the task escalates
- Deployment in your VPC or on-prem, with observability, runbooks, and a rollback path
- Full ownership — weights, training data, eval set, and harness config are yours; no proprietary runtime, nothing that dies if we do
What moves the scope. Every quote decomposes into the same six components: eval-set construction and labeling, teacher-model inference for training-data generation, human validation of training data, fine-tuning compute, deployment with routing and observability, and evaluation cycles to the pass mark. Multiple task types, sensitive data that needs more human validation, on-prem rather than VPC deployment, and compliance documentation move you up the range. Negotiation doesn't.
| Result | Number | Source |
|---|---|---|
| Inference cost after moving routine support workloads to fine-tuned small models | −66–80% | Forethought, published on AWS |
| Task accuracy after specialization, with inference costs down 50–68% | 81% → 93% | distil labs × Knowunity, vendor-reported |
| Serving cost after a small-model deployment | ≈ −50% | distil labs on Cerebrium, vendor-reported |
| Adoption of task-specific small models vs. general-purpose LLMs by 2027 | 3× | Gartner, forecast |
Published results from other teams, not our client numbers — which is exactly why every project proves the economics on your data, against your eval set, before you commit to the build.
- Not a chatbot project, and not "AI transformation." One workflow, one model, one measured result.
- Not fine-tuning as magic. If the eval says the frontier model keeps winning on your task, the readout says that, and the gate does its job.
- Not lock-in. No proprietary serving layer; your team can run the system without us.
- Not research. We ship to production or stop at the gate — there is no third mode.
Do we lose quality by moving off a frontier model?
On the narrow task, published evidence points the other way — in distil labs' work with Knowunity, task accuracy rose from 81% to 93% after specialization (vendor-reported). The router keeps a frontier model in the loop for cases that genuinely need one. The eval set decides, not opinion.
Base models improve every few months. Does our model go stale?
It drifts, which is why re-training exists: a re-distillation cycle runs as a scoped project when the eval numbers say it's due, or folds into an Assurance retainer. Even with that recurring work, a task-specific small model serves routine volume at a fraction of frontier API list prices.
Private AI Deployment. Your data never leaves — that's the whole design.
We deploy small language models — plain or fine-tuned — inside your VPC or on your own hardware. No third-party processing, no vendor training on your data, and your existing security controls keep applying, because the system lives inside your perimeter. Who buys this: firms whose clients contractually forbid external data processing, regulated mid-market teams whose DPAs stall every cloud-AI procurement, and anyone who has concluded that the simplest compliance argument is architectural — nothing leaves, so there's nothing to argue about. For law firms in England & Wales we run a dedicated practice: Legal.
- An inference server (vLLM-class) sized to your workload
- The model — open-weights as-is, or fine-tuned on your task through a distillation project
- Retrieval over your documents, where the workflow calls for it
- Access control wired into your SSO and existing permissions
- Audit logging, monitoring, runbooks, and a handover your IT team can actually operate
- Hardware spec, if you need one — you buy it from your vendor; we take no margin on boxes
Security posture, stated plainly. We're a studio, not a certified vendor: there is no SOC 2 badge here. We offer an architecture where certifying us stops mattering — the system runs inside your perimeter, under your controls, and we hand over every config and runbook for your team to inspect line by line. If your procurement process requires a certified external processor, we're honestly the wrong shape. If it requires that nothing leaves your network, we're exactly the right one.
- Not a managed cloud service. You host it; we build it, document it, and optionally keep it measured under Assurance.
- Not a "private ChatGPT for everything" portal. We deploy against named workflows with defined success measures — broad assistants without an eval are how private AI projects quietly die.
- Not a compliance certification, and not legal advice on your regulatory position.
Can our IT team run this after handover?
Yes — that's a design constraint, not an afterthought. Standard open components, runbooks, monitoring dashboards, and a documented update path. Assurance exists for teams that don't want to own the rota, not because the handover is incomplete.
Is a private model just a worse model?
On open-ended reasoning, frontier APIs win. On the narrow workflow you'd deploy this for, published work shows fine-tuned small models matching or beating them — distil labs' Knowunity project moved task accuracy from 81% to 93% (vendor-reported). Hard cases escalate through the router, and air-gapped deployments can disable escalation entirely. The trade-off is real, and we show it to you in numbers.
GTM & RevOps AI Systems. Revenue systems that don't ship your pipeline to an API.
This is the vertical we operate in ourselves. The person building your system spends every working day inside Salesforce, Pardot, Clay, Lemlist, and ZoomInfo as a GTM engineer at a B2B SaaS company. You're not paying an agency to discover your stack on your budget. Everything runs in your VPC, inside your Salesforce org's trust boundary, inside your own automation layer — your prospect and customer data is the asset these systems run on, and the entire point is that it stays yours.
- Private lead & intent scoring — a fine-tuned classifier scores leads inside your perimeter; CRM data never leaves. Published work shows task-specific small models beating general ones on exactly this kind of narrow classification (distil labs × Knowunity, 81% → 93%, vendor-reported)
- CRM data extraction — a local model parses email and call notes into structured Salesforce fields, instead of a third-party processor reading your pipeline
- Enrichment engine — job titles, industries, and company names normalized at fractions of a cent per row; drops into Clay as a node
- Outbound personalization at scale — 100K personalized emails on a locally served 7B model. Our estimate from published list prices, August 2026: $0.00004–$0.0001 per call against $0.025–$0.15 on a frontier API — roughly 250×, which is what makes personalization at real volume affordable at all
- Not a Clay agency. If you need tables configured and no model work, a tables-and-workflows GTM agency is cheaper — and we'll tell you so on the first call.
- Not an outbound-sending service. We build the system; your team runs the sends, owns the domain reputation, and owns the results.
- Not a data vendor. Bring your own ZoomInfo, Clay, and Lemlist licenses — we don't resell data or seats.
We already run GPT inside Clay. Why change anything?
Two reasons, both measurable: per-row cost at volume, and every enrichment call shipping your CRM data to a third party. If your volume is low and your data policy allows it, your current setup is fine — that's a legitimate Crit verdict.
Does this replace SDRs?
No. It removes the routine layer — scoring, parsing, normalizing, drafting — that eats their hours or your API budget. Decisions and conversations stay human; anyone promising otherwise sells you a headcount story, not a system.
Assurance. A deployed model is a system, not a finish line.
Data drifts. Upstream tools change their outputs. Base models get deprecated. Router thresholds set in March are wrong by September. None of this is a crisis if someone measures — all of it is, quietly, if no one does. You can run this yourself: every build ships with runbooks, and the handover is complete without us. Assurance is for teams that don't want to own the rota. No lock-in mechanics: the weights, eval set, and configs were yours from the day the build shipped, so leaving Assurance costs you nothing but the rota.
- Monthly eval re-runs against your frozen set, with scores against the pass mark
- Drift and cost report — one page of numbers, not a slide deck
- Monitoring and alerts on accuracy, cost per call, and escalation rate
- Router threshold tuning as your traffic shifts
- Re-training cycles when the eval numbers say they're due — not on a calendar, not on a hunch
- Workflow change requests scoped into the retainer as your systems evolve
- Not a general AI helpdesk or advisory line. Assurance covers the systems we deployed, against the evals we froze.
- Not infrastructure operations. Your team owns the servers and the network; we own the model system's health on top of them.
- Not mandatory. Some teams take the runbooks and run everything themselves — the handover is designed for exactly that.
What does the monthly report actually show?
Accuracy against the pass mark, cost per call, escalation rate to the frontier model, drift indicators on inputs, and what changed since last month. One page of numbers, not a slide deck.
What happens if accuracy drops below the pass mark?
A re-training cycle triggers under your Assurance scope. If the drop is structural — the workflow itself changed shape — the report says so and we scope the options honestly, including the option that this workflow no longer suits a small model.
What skipping The Crit costs you. All public numbers.
Un-audited AI stacks drift toward numbers like these. None of them are our fees — they're what AI spend looks like from the outside when nobody measured first.
| Line item | Public number | Source |
|---|---|---|
| Clay at production scale — teams budget from the list price, then credits, seats, and enrichment overages land | $6K–$30K listed → $75K–$120K/yr actual | Amplemarket, public analysis |
| Inference as a share of revenue, average across AI-SaaS companies | 23% of revenue | ICONIQ, published survey |
| One 7-step Zapier workflow at 100 tickets a day | ≈ $300/month | Public list prices |
Public sources and list prices, August 2026. The Crit exists so numbers like these show up in a readout before they show up in your invoice.
Your workload runs under roughly 50M tokens a month and nothing about your data is private. A frontier API is probably fine — and The Crit will say so in writing.
You need a strategy deck, a team of twenty, or a vendor with a SOC 2 badge for a procurement file. We're the wrong studio, and we'd rather say it here than on a call.
You want a chatbot for everything rather than a measured result on one workflow. Broad assistants without an eval are how AI projects quietly die.
We qualify hard. It keeps your budget and our calendar honest.
Zero data risk before day one.
- NDA before anything moves
- Read-only, least access
- Redacted samples welcome
- Your perimeter from the start
We’re engineers, not your compliance advisers — the architecture keeps your existing compliance intact instead of adding a vendor to it.
Start with the numbers.
Two to three weeks inside your workflows and invoices, and you know where a small model wins, where the API is fine, and where AI doesn't belong. The numbers are yours either way.