Use case · Private lead scoring

Score every lead without sending your CRM to someone else's model.

A fine-tuned 3–8B model, deployed in your VPC, classifies leads and buying intent from the data already in Salesforce or HubSpot. In published fine-tuning work, task-specific small models match or beat general models on narrow classification (distil labs × Knowunity: 81→93%).

Private lead scoring runs a fine-tuned 3–8B model inside your VPC, reads the signals already in your CRM, and writes a score plus reason codes back to the fields your team routes on. No lead record touches a third-party API. We measure accuracy on a frozen eval set built from your historical outcomes before the model ships — you see the number first.

SDRs work lists sorted by guesswork.

Most mid-market revenue teams either score leads with static rules written two RevOps hires ago, or don't score at all — because piping firmographic and behavioral CRM data through a third-party AI API is exactly what the security review said no to. Either way, the signals sitting in your own CRM go unread. The data never has to leave your infrastructure to be useful.

How it works · The Crit Path
01Scope — freeze what a good lead meansWe pull 1,000–2,000 historical leads with known outcomes from your Salesforce or HubSpot — closed-won, closed-lost, no-show, ghosted — and freeze an eval set that defines what a good lead means for your business. Nothing ships until the model clears it.
02Prove — fine-tune on your definition of intentA 3–8B open-weights model (Llama or Qwen class) trains on your labeled leads and signal history: title patterns, engagement sequences, firmographic fit, product usage where you have it. It must beat your current baseline on the frozen set — or the project stops there, and you keep the eval set.
03Deploy — inside your perimeterThe model runs on a single GPU instance in your VPC or on-prem, behind a plain scoring API. A native integration reads new and updated leads, scores them, and writes the score and reason codes back to the lead fields your team already routes on.
04Assure — retrain when reality movesWe track scores against actual outcomes. When drift shows up — new ICP, new product line, new market — the model retrains against the updated eval set. Quarterly by default, on demand when the data says so.
The economics

A worked example, from public list prices. At 100,000 lead scores a month — roughly 1,500 tokens of context per lead — a frontier API bills every score and grows linearly with volume. The same workload on a fine-tuned 3–8B model runs on one GPU instance at a flat cloud cost. Our cost model at public list prices, August 2026; The Crit replaces it with your numbers. Gartner expects task-specific small models to be adopted three times more than general-purpose LLMs by 2027.

Frontier APIFine-tuned small model, yours
Monthly inference~$400–1,500 at published list prices, growing linearly with volumeFlat: one GPU instance, ~$400–700/mo at cloud list prices; a score costs a fraction of a cent
Data exposureEvery lead record processed by a third partyNothing leaves your security perimeter
Accuracy on your taskGeneral model guessing at your ICPTrained on your outcomes: 81→93% in distil labs × Knowunity (vendor-reported)
Published result66–80% lower inference cost on routine workloads — Forethought, on AWS
Where this doesn't work

If your volume is low and nothing stops you from sending CRM data to a third party, a frontier API — or a well-built rules layer — is probably fine. The Crit will tell you so before you spend a build budget finding out.

What you get

Scoring you explain. And own.

  • A frozen eval set built from your historical lead outcomes — yours to keep, whoever you work with next
  • A fine-tuned 3–8B scoring model with measured accuracy against that eval set, reported before deployment
  • Deployment in your VPC or on-prem: scoring API, CRM write-back via native integration
  • Score fields plus reason codes — every score is explainable to the SDR reading it
  • A monitoring dashboard: score distribution, drift against outcomes, per-segment precision
  • Retraining runbook and cadence your team can operate
Who this is for
Teams blocked by security reviewRevenue teams on Salesforce or HubSpot whose security or legal review has said no to third-party AI processing of CRM data.
Volume that rules can't keep up withMid-market companies scoring 10,000+ leads or accounts a month where rules-based scoring has visibly stopped matching reality.
RevOps that wants an audit trailLeaders who want scoring they can explain — a model they own, with reason codes, not a black-box SaaS score.
Questions, answered straight

Can lead scoring run without sending data to a third-party API?

Yes. A small open-weights model (3–8B parameters) fine-tuned on your historical lead outcomes runs entirely inside your VPC or on-prem, so no third-party API ever processes a CRM record. In one vendor-reported benchmark, a fine-tuned 3.8B model outperformed GPT-4o on enterprise text classification — for a narrow, well-defined task like lead scoring, small models are not a downgrade.

How accurate is a small model at lead scoring compared to GPT-4-class models?

On a narrow classification task with good training data, a fine-tuned small model typically matches or beats a general-purpose frontier model. In distil labs' published work with Knowunity, task accuracy rose from 81% to 93% after fine-tuning a small model, while inference costs fell 50–68% (vendor-reported). We measure accuracy on a frozen eval set built from your own outcomes before anything deploys.

What does on-prem lead scoring cost to run?

A single GPU instance — roughly $400–700 a month at cloud list prices — handles hundreds of thousands of lead scores monthly, at a marginal cost per score of a fraction of a cent (our cost model, August 2026). The Crit tells you whether the economics work for your volume before you commit.

Security · by architecture

Your CRM never feeds Big AI.

  • Pipeline data stays in your VPC
  • No third-party enrichment APIs
  • Your DPA stays unchanged
  • Prospect data never resold
Architected to deploy inside your
GDPR CCPA SOC 2
Small models · critical hits

Your CRM already holds the training set.

Two weeks inside your lead history, and you'll know what a model trained on your own outcomes can do — and what it can't. The numbers are yours either way.