Your case files, answerable. Without leaving the building.
A retrieval system and a fine-tuned small model, deployed on infrastructure your firm controls. It answers only from your own documents, cites every source, and logs every question — the trail your COLP can produce on request.
Legal document intelligence puts case-file search and contract pre-review on models that run inside your firm's perimeter. A local index reads your DMS and drives in place, a 7–8B open-weight model answers only from retrieved passages with a citation on every sentence, and an audit log records every query. After [2026] UKUT 81 — where the Upper Tribunal held that uploading confidential documents to a public AI tool can waive privilege — the perimeter is the argument.
The choice isn't whether AI touches client work. It's where.
In 2025 the High Court referred the lawyers in Ayinde v Haringey to their regulators after invented, AI-generated citations reached the court. In 2026 the Upper Tribunal went further: uploading confidential client documents to a public AI tool can place them in the public domain and waive legal professional privilege for good ([2026] UKUT 81). Most fee earners use these tools anyway — so the real choice is whether AI does client work inside your perimeter, with citations and a log, or outside it, with neither.
For a firm of 10–50 fee earners, cost is usually the second argument — confidentiality is the binding constraint. The numbers still matter, so here they are, from public list prices, August 2026:
| Per-seat legal AI | A system you own | |
|---|---|---|
| Cost shape | Per user, per month, forever | Flat — the same for five fee earners or fifty |
| Thirty seats, a year | £36,000–54,000 at roughly £100–150/user/month — public list prices, August 2026 | One GPU server plus an assurance retainer |
| Restricted client files | Excluded — third-party processing | Included — nothing leaves your perimeter |
| Archives outside your tenant | Invisible to cloud tools | Indexed in place by connectors |
Per-document math — a worked example, our assumptions. A first-pass review of a 40-page agreement is roughly 30–35K tokens: a few tens of cents per pass through a frontier API at published list prices — cheap at low volume, and we'll tell you so. The economics flip at volume. Disclosure review, due-diligence data rooms and archive-wide search run thousands of documents, where a self-hosted small model's marginal cost approaches electricity.
Published benchmarks, not our claims. In distil labs' published work with Knowunity, a fine-tuned small model raised task accuracy from 81% to 93% while cutting inference costs 50–68% (vendor-reported). Forethought reports 66–80% lower inference costs after moving routine workloads to fine-tuned small models (published on AWS). We haven't shipped this exact configuration for an external law firm yet; we prove it on your matters before you commit.
At low document volume with no confidentiality constraint, a per-seat tool or a frontier API is probably fine — and The Crit will say so in writing. And no system removes the lawyer: the SRA is clear that responsibility for AI outputs stays with you, so every answer still needs professional verification before it leaves the building. We build for that workflow, not around it.
The system, the evals, the log. Yours.
- A working retrieval and Q&A system over your matter archive, deployed in your office or private cloud
- A contract pre-review model fine-tuned on your playbook, with confidence thresholds and mandatory human sign-off
- A frozen eval set built from your own documents — accuracy measured before deployment, not asserted after
- Source-pinned citations on every answer; a complete query log exportable for your COLP or insurer
- Drafted answers to the AI questions on your professional indemnity proposal form
- Monitoring, re-indexing and re-training under an Assurance retainer
Is it safe to cite AI answers in court work?
Not without verification, and we don't pretend otherwise. Our system reduces the risk structurally: it answers only from your own retrieved documents, every claim links to its source paragraph, and the workflow requires a lawyer's sign-off before anything leaves the building. The SRA's guidance is blunt: “You remain responsible and accountable for the outputs from AI you are using.” Our job is to make that responsibility easy to carry and easy to evidence.
Does client data ever leave our infrastructure?
No. The embedding model, the index and the language model all run on hardware you control — an on-prem GPU server or your private cloud tenancy. There is no third-party processing, no vendor training on your data, and no new entry on your client-facing subprocessor list. That matters after [2026] UKUT 81, where the Upper Tribunal held that uploading confidential documents to a public AI tool can waive privilege.
Why not just Copilot from our MSP?
Sometimes that is the right answer, and we'll say so in the audit. What Copilot doesn't do: index the matters sitting outside your tenant, satisfy engagement terms that prohibit third-party processing outright, or give you citation-verified answers over your archive with a log your COLP can produce. If any of those describes your firm, the comparison stops being close.
Privilege survives the AI.
- Files never leave the firm
- No public AI in the chain
- Confidentiality by custody
- COLP-ready audit log
We’re engineers, not your compliance advisers — the architecture keeps your existing compliance intact instead of adding a vendor to it.
Your archive already knows the answer.
Two weeks inside a governed sample of your matters, and you'll know what a private system can answer, what it can't, and what running one inside your walls takes. The numbers are yours either way.