Score every lead without sending your CRM to someone else's model.
A fine-tuned 3–8B model, deployed in your VPC, classifies leads and buying intent from the data already in Salesforce or HubSpot. In published fine-tuning work, task-specific small models match or beat general models on narrow classification (distil labs × Knowunity: 81→93%).
Private lead scoring runs a fine-tuned 3–8B model inside your VPC, reads the signals already in your CRM, and writes a score plus reason codes back to the fields your team routes on. No lead record touches a third-party API. We measure accuracy on a frozen eval set built from your historical outcomes before the model ships — you see the number first.
SDRs work lists sorted by guesswork.
Most mid-market revenue teams either score leads with static rules written two RevOps hires ago, or don't score at all — because piping firmographic and behavioral CRM data through a third-party AI API is exactly what the security review said no to. Either way, the signals sitting in your own CRM go unread. The data never has to leave your infrastructure to be useful.
A worked example, from public list prices. At 100,000 lead scores a month — roughly 1,500 tokens of context per lead — a frontier API bills every score and grows linearly with volume. The same workload on a fine-tuned 3–8B model runs on one GPU instance at a flat cloud cost. Our cost model at public list prices, August 2026; The Crit replaces it with your numbers. Gartner expects task-specific small models to be adopted three times more than general-purpose LLMs by 2027.
| Frontier API | Fine-tuned small model, yours | |
|---|---|---|
| Monthly inference | ~$400–1,500 at published list prices, growing linearly with volume | Flat: one GPU instance, ~$400–700/mo at cloud list prices; a score costs a fraction of a cent |
| Data exposure | Every lead record processed by a third party | Nothing leaves your security perimeter |
| Accuracy on your task | General model guessing at your ICP | Trained on your outcomes: 81→93% in distil labs × Knowunity (vendor-reported) |
| Published result | — | 66–80% lower inference cost on routine workloads — Forethought, on AWS |
If your volume is low and nothing stops you from sending CRM data to a third party, a frontier API — or a well-built rules layer — is probably fine. The Crit will tell you so before you spend a build budget finding out.
Scoring you explain. And own.
- A frozen eval set built from your historical lead outcomes — yours to keep, whoever you work with next
- A fine-tuned 3–8B scoring model with measured accuracy against that eval set, reported before deployment
- Deployment in your VPC or on-prem: scoring API, CRM write-back via native integration
- Score fields plus reason codes — every score is explainable to the SDR reading it
- A monitoring dashboard: score distribution, drift against outcomes, per-segment precision
- Retraining runbook and cadence your team can operate
Can lead scoring run without sending data to a third-party API?
Yes. A small open-weights model (3–8B parameters) fine-tuned on your historical lead outcomes runs entirely inside your VPC or on-prem, so no third-party API ever processes a CRM record. In one vendor-reported benchmark, a fine-tuned 3.8B model outperformed GPT-4o on enterprise text classification — for a narrow, well-defined task like lead scoring, small models are not a downgrade.
How accurate is a small model at lead scoring compared to GPT-4-class models?
On a narrow classification task with good training data, a fine-tuned small model typically matches or beats a general-purpose frontier model. In distil labs' published work with Knowunity, task accuracy rose from 81% to 93% after fine-tuning a small model, while inference costs fell 50–68% (vendor-reported). We measure accuracy on a frozen eval set built from your own outcomes before anything deploys.
What does on-prem lead scoring cost to run?
A single GPU instance — roughly $400–700 a month at cloud list prices — handles hundreds of thousands of lead scores monthly, at a marginal cost per score of a fraction of a cent (our cost model, August 2026). The Crit tells you whether the economics work for your volume before you commit.
Your CRM never feeds Big AI.
- Pipeline data stays in your VPC
- No third-party enrichment APIs
- Your DPA stays unchanged
- Prospect data never resold
Your CRM already holds the training set.
Two weeks inside your lead history, and you'll know what a model trained on your own outcomes can do — and what it can't. The numbers are yours either way.