Clinical notes, drafted where the data already lives.
Local speech-to-text and a fine-tuned small model draft visit notes on hardware inside your perimeter. No audio or patient records transit a third-party API, and no note enters the record without a clinician's sign-off.
On-prem healthcare documentation runs a Whisper-class speech model and a fine-tuned 7B note-drafting model entirely on hardware you control. Audio becomes a transcript inside your perimeter, the model drafts the note against your own templates, and a clinician reviews and signs before anything enters the record. The system drafts; it does not diagnose.
Ambient scribes proved the demand. Their data flow stalls the sign-off.
Documentation time is a known driver of clinician workload, and ambient scribes have proven the demand — but nearly all of them stream consultation audio and patient identifiers to a vendor's cloud, which is exactly what many data-protection officers, NHS information governance teams and US compliance officers cannot sign off. The underlying tasks — transcription, summarisation into a note template — are narrow enough for models that run entirely on your own hardware.
Per-seat versus owned — a worked example, our assumptions. Commercial ambient scribes list at roughly $100–250 per clinician per month (vendor-published pricing, 2025). An owned system inverts the shape: a fixed build, modest on-prem hardware, and a monthly assurance retainer — with cost independent of seat count.
| Cloud ambient scribe | A system you own | |
|---|---|---|
| Pricing shape | Per clinician, per month — roughly $100–250 at vendor-published prices, 2025 | Fixed build plus assurance retainer — independent of seat count |
| Fifty clinicians, a year | $60,000–150,000, recurring, plus the data-processing agreement | Modest on-prem hardware, owned |
| Consultation audio | Streams to the vendor's cloud | Processed and discarded inside your perimeter |
| Governance review | Review the vendor's cloud | Review a system you run |
Published benchmarks, not our claims. Fine-tuned small models on narrow tasks have shown accuracy gains from 81% to 93% with 50–68% lower inference costs (distil labs / Knowunity, vendor-reported), and Gartner expects task-specific small models to be adopted three times more than general-purpose LLMs by 2027.
If a cloud scribe has already cleared your information-governance review and per-seat pricing works at your headcount, buy it — this system earns its place where governance blocks the cloud path or seat counts compound. And nothing here makes clinical decisions: documentation is a deliberate boundary, not a temporary one.
We haven't shipped this vertical yet.
Healthcare deployments carry real integration and governance cost — this is the most capital-intensive workflow on this site. That's why the engagement starts with a paid audit and a measured pilot on your own data, not a contract for a platform. The audit produces the data-flow map, the hardware spec and the accuracy numbers; you decide with those in hand.
The pipeline, inside your governance. Yours.
- Local ASR plus a note-drafting model deployed on your hardware or private cloud
- Note templates tuned to your specialty, with mandatory clinician sign-off in the workflow
- A frozen eval set from your own consented, governed documentation; accuracy measured before rollout
- A data-flow map and governance documentation pack for your DPO / information governance review
- A retention design where audio is processed and discarded on your terms
- Monitoring and retraining under an Assurance retainer
Can ambient AI scribes work without sending audio to the cloud?
Yes. Speech-to-text models in the Whisper class run well on local hardware, and a fine-tuned 7B model drafts the note from the transcript on the same infrastructure. The entire pipeline — audio, transcript, draft — stays inside your perimeter, which changes the information-governance conversation from "review this vendor's cloud" to "review this system we run."
Is this a medical device? Does it make clinical decisions?
No. The system transcribes and drafts documentation; it does not diagnose, recommend treatment or triage patients. Every note requires clinician review and sign-off before it enters the record. Scoping the system to documentation is a deliberate boundary, not a temporary one.
What does deployment require from our IT team?
Less than a typical clinical system: one GPU server (or equivalent private-cloud tenancy) inside your network, access to your note templates, and a governed sample of historical documentation to build the eval set. The audit produces the full data-flow map and hardware spec before you commit to anything.
Nothing leaves the building.
- Audio never reaches a vendor
- Records stay on premises
- Clinicians control the data
- DPO-ready documentation
We’re engineers, not your compliance advisers — the architecture keeps your existing compliance intact instead of adding a vendor to it.
Start with the audit, not the platform.
The Crit produces the data-flow map, the hardware spec and a measured pilot plan on your own documentation — before you commit to anything. The numbers are yours either way.