The Crit Path.
Four stages. One rule: nothing ships until it beats the benchmark on your data. The name plays on the critical path — the shortest sequence that determines the result — and on the design crit: the honest review before anything gets built.
The Crit Path is how we ship: Scope the workflow, Prove the model on your data against a frozen eval set, Deploy inside your perimeter, Assure with monitoring and re-training. The Prove step is a gate, not a formality — if the model doesn't clear the pass mark you set, the project stops and you keep the artifacts. Every build, in every vertical, runs on this rail.
Scope. Freeze the target before anyone trains anything.
Prove. Beat the benchmark or stop.
Deploy. Inside your perimeter, behind your interface.
Assure. Keep the system honest after launch.
Hitbox. The part that decides who answers.
Hitbox is our eval-and-routing harness, built and operated by Crit Studio. It does two jobs. Before launch, it gates deployment behind measured accuracy: no model ships until it clears the pass mark on the frozen eval set. In production, it decides per request whether the small model answers or the task escalates to a frontier model.
Escalation is a design target, not an embarrassment. We set the expected escalation share at Scope, and the monthly report tracks it against that target. A router that never escalates isn't confident — it's unmeasured.
The gate is the point. A stopped project with honest measurements beats a deployed one that quietly underperforms.
The pass mark freezes at Scope, before anyone trains anything — so the bar can't quietly move to meet the model. If the model doesn't clear it after the agreed evaluation cycles, the project stops. You pay for the phases delivered and keep everything produced: the eval set, the training data, the numbers.
The eval set outlasts the project. It's the yardstick for any build — by us or by anyone else — and walking away with honest numbers is a legitimate outcome, not a consolation prize. We price the audit so its readout is worth having on its own.
What happens if the model fails the Prove gate?
The project stops. You pay for the phases delivered and keep everything produced — the frozen eval set, the human-validated training data, and the numbers. A stopped project with honest measurements beats a deployed one that quietly underperforms.
Who sets the pass mark?
You do, with us, at Scope — anchored to your current system's measured accuracy on the frozen eval set. We write it into the statement of work before any training starts, and it doesn't move afterwards. A bar that moves to meet the model isn't a bar.
How is escalation to a frontier model decided?
Hitbox, our eval-and-routing harness, scores each request's confidence; below the threshold, the task escalates to your existing frontier API. We set the expected escalation share as a design target at Scope and track it in the monthly Assure report. Air-gapped deployments can disable escalation entirely.
One small model. Then a workforce.
Every system we ship shares one core: your model, your router, your eval harness. The first deployment is the expensive one — every workflow after it reuses the same infrastructure, the same monitoring, the same lessons. Owned models compound; rented APIs just meter.
Scale is a property of ownership. Rented APIs meter every workflow you add — an owned core amortizes them.
Custody holds at every stage.
- Data stays in your cloud
- Access expires with the project
- Every deployment logged
- Escalation only if you allow it
Put your workflow on the rail.
One scoping call tells you whether your workflow belongs on the Crit Path — and the Prove gate makes sure you never find out the expensive way.