Use case · Operations & Edge

AI that works where the signal doesn't.

We fine-tune small models to run on the rugged tablets and phones your field crews already carry — construction sites, warehouses, farms. No connection required. No data leaving the device.

Offline field AI runs a small model fine-tuned on one task directly on the phones and rugged tablets your crews already carry. A 1–3B model, quantized to 4-bit, answers with no network round trip and no data leaving the device. In distil labs' published work with Knowunity, the same approach raised task accuracy from 81% to 93% while cutting inference costs 50–68% (vendor-reported) — we prove it on your field data before full deployment.

Cloud AI assumes a connection. Field work doesn't offer one.

Basements, remote sites, warehouses with metal shelving, farmland with no coverage — the AI tools your office runs on die exactly where your margin is made. And even where there is signal, site photos, client documents, and defect logs often shouldn't leave the device at all.

How it works · The Crit Path
01Scope — pick the one workflowDaily reports, defect logging, spec lookup, inventory counts — The Crit picks a single task with measurable time cost and builds a frozen eval set from your real field data before anything else.
02Prove — fine-tune a 1–3B model for that taskWe train a Llama 3.2 3B or Qwen 2.5 3B-class model on your forms, terminology, and photo metadata, then quantize it to 4-bit so it runs on a phone or rugged tablet NPU — via llama.cpp, MLC, or ExecuTorch, depending on your device fleet. It ships only if it clears the frozen eval set.
03Ship — inside your existing field appThe model answers on-device with no round trip. When connectivity returns, the app syncs results to your backend; genuinely hard cases can escalate to a frontier model — online only, and only if you want them to.
04Assure — monitor and retrainTelemetry syncs opportunistically; we track accuracy against the frozen eval set and retrain when your forms, specs, or workflows change.
The economics

Connectivity is the headline here, not the API bill. A cloud model at any price returns nothing in a dead zone. The numbers still point the same way — illustrative math from our cost model and public list prices, August 2026; The Crit replaces it with your volumes: 100 field workers making 30 AI-assisted actions a day over 22 working days is roughly 66,000 requests a month. Through a frontier API that runs $660–1,300 a month, grows linearly with headcount, and still covers nothing offline. A quantized 3B model on hardware your crews already carry adds no marginal inference cost — the remaining sync and optional escalation traffic is a rounding error.

Cloud APIFine-tuned model on the device
CoverageDies with the signal — zero answers in a dead zoneAnswers in a basement, a tunnel, a field with no bars
Cost shapePer request, linear with headcountNo marginal inference cost — hardware your crews already carry
Data exposureSite photos and client documents stream to a third partyNothing leaves the device
Accuracy on the narrow taskGeneral model guessing at your forms81→93% in distil labs × Knowunity (vendor-reported)
Where this doesn't work

If your teams always have connectivity and no client contract restricts where data goes, a frontier API is probably fine — and The Crit will say so. Same if the workflow changes weekly: a fine-tuned model wants a stable, repetitive task.

What you get

Yours to run, on hardware you already own.

  • Own the fine-tuned, quantized weights — no per-token licence, no vendor to outgrow
  • Ship the model inside your existing field app, iOS and Android
  • Keep the frozen eval set and a written accuracy report from your real field data
  • Run offline-first: an on-device sync layer plus optional online escalation to a frontier model
  • Operate it yourself with a monitoring dashboard and a retraining runbook
Who this is for
Crews that work where coverage diesConstruction, logistics, and agriculture operators whose sites, warehouses, and fields sit outside reliable signal — or whose client contracts restrict what goes to the cloud.
Teams with a field app already in handAn existing iOS or Android field app — or room to add one screen to it — and a repetitive documentation task eating crew hours.
Data that shouldn't travelSite photos, defect logs, and client documents that stay on the device by default, not by exception.
Questions, answered straight

Can an AI model run fully offline on a phone or tablet?

Yes. A 1–3B parameter model, quantized to 4-bit, runs on a modern phone or rugged tablet NPU with no network connection — using runtimes like llama.cpp, MLC, or ExecuTorch. It won't write essays like a frontier model, but fine-tuned on one narrow task — logging defects, filling reports, answering spec questions — it matches or beats large models on that task.

How accurate is a small on-device model compared to GPT-4o-class models?

On the one task it was fine-tuned for, a small model is usually more accurate, not less. In distil labs' published work with Knowunity, accuracy rose from 81% to 93% after moving from a large general model to a fine-tuned small one (vendor-reported). We measure this on your data against a frozen eval set before deployment — if the small model doesn't clear the bar, we tell you.

What happens when the device comes back online?

The app syncs completed work to your backend and pulls model updates if any. Optionally, cases the on-device model flagged as hard can escalate to a frontier model while connected. Escalation is a config choice, not a requirement — some clients keep everything on-device for data reasons.

Security · by architecture

Air-gap friendly by design.

  • Works fully offline
  • Data lives on the device
  • Sync only to your servers
  • Cloud optional — never required
Architected to deploy inside your
GDPR ISO 27001
Small models · critical hits

The signal dies. The model doesn't.

One workflow, one fine-tuned model on hardware your crews already carry — proven against your real field data before it ships. If the numbers don't clear the bar, you keep the eval set and the report.