Tier-1 tickets, resolved by a model you own.
Order status, returns, shipping, sizing — the repetitive bulk of e-commerce support, handled by a small model fine-tuned on your resolved tickets and running in your infrastructure. The resolution rate is verified against your ticket history before rollout, not after.
Tier-1 support automation fine-tunes a small model on your own resolved tickets and runs it inside your infrastructure — auto-resolving routine categories, drafting replies for agent review, and handing the rest to humans with a summary. Forethought, a customer-support AI company, reports 66–80% lower inference costs after moving routine support workloads to fine-tuned small models (vendor-reported, on AWS).
Every promotion is a wave of tickets. Every wave is rent, per ticket, forever.
Tier-1 volume scales with sales: each campaign either burns out agents or inflates a per-resolution AI vendor bill. Meanwhile the tickets themselves — and the customer data inside them — flow through a third party's cloud, on a problem that is overwhelmingly repetitive.
Per-ticket math, from public list prices, August 2026. A tier-1 ticket runs 1–3K tokens end to end. Per-resolution AI support vendors typically list around $1 per automated resolution; a frontier API costs $0.01–0.05 per ticket in tokens; a self-hosted 7B model handles the same ticket for a fraction of a cent, with costs that are mostly fixed infrastructure rather than per-ticket rent. At 30K tickets a month, that's the difference between a five-figure monthly vendor bill and a GPU instance. Our cost model; The Crit replaces it with your numbers.
| Per-resolution vendor / frontier API | Fine-tuned small model, yours | |
|---|---|---|
| Cost shape | Per resolution or per token — scales with every campaign | Fixed build plus modest infrastructure; marginal cost near zero |
| Per tier-1 ticket | ~$1 per automated resolution (public list prices, Aug 2026) | A fraction of a cent in marginal cost |
| Customer data | Transits a third party's cloud | Stays in your VPC — read-only access to your OMS |
| Published result | — | 66–80% lower inference cost — Forethought, on AWS |
Below a few thousand tickets a month, an off-the-shelf tool is probably the right answer, and we'll say so in the audit. The owned-model economics need volume — or a hard privacy constraint — to beat per-resolution rent.
The support stack. Yours.
- Ticket taxonomy and automation map: what automates, what drafts, what stays human
- Fine-tuned model and weights — trained on your tickets, owned by you
- Helpdesk and OMS integration: Zendesk, Gorgias, or Intercom; Shopify or your order system, read-only
- Human-handoff flows with case summaries, and per-category confidence thresholds you control
- Eval set built from your ticket history, plus a dashboard: resolution rate, escalation rate, CSAT on automated replies
Can a small language model really handle customer support tickets?
Yes, for tier-1 — the repetitive majority: order status, returns, shipping, product questions. A 3–7B model fine-tuned on your resolved tickets and grounded in your policies handles these reliably because the task is narrow; we verify the resolution rate against your own ticket history before anything goes live. Complex and sensitive tickets are routed to humans by design.
What happens to tickets the model can't handle?
They go to your agents, with a summary attached. Every category has a confidence threshold: above it the model resolves or drafts, below it the ticket escalates to a human. The escalation rate is a tracked metric, and thresholds are tightened or loosened per category based on measured accuracy — not set once and forgotten.
How does an owned support model compare to per-resolution AI vendors on cost?
Per-resolution vendors typically list around $1 per automated resolution (public list prices, August 2026), so the bill scales with your ticket volume forever. A fine-tuned small model running in your infrastructure resolves the same tier-1 ticket for a fraction of a cent in marginal cost — the spend shifts from per-ticket rent to a fixed build plus modest infrastructure, which favors ownership from roughly 10K tickets a month.
Cut the bill. Keep the custody.
- Data stays in your VPC
- No new subprocessor
- Escalation under your rules
- Keys and logs are yours
Your resolved tickets are the training set.
Six months of ticket history tells us what a model can resolve, what it should draft, and what stays human — measured before rollout, not promised after. The numbers are yours either way.