We ran the same support-ticket workload across open and proprietary models and compared answer quality, latency and cost per resolved ticket.
Method
Each model answered the same anonymised set of support tickets using an identical retrieval setup. Answers were scored by trained reviewers against a shared rubric.
How to use the results
The cheapest model per token is rarely the cheapest per resolved ticket. The report shows where smaller models are good enough, and where escalation to a larger model pays for itself.
Methodology and dataset
We used an anonymised set of support tickets with known resolutions. Every model received the same retrieved context and system prompt.
Quality scores by model
Larger models scored highest on complex multi-step tickets. On routine tickets the gap between large and small models was small.
Latency at p50 and p95
Median latency varied widely between providers. Tail latency at p95 mattered more for user experience than the median.
Cost per resolved ticket
Cost per resolved ticket combines token cost with the rate of escalation to a human. A cheaper model that escalates more often can cost more overall.
Recommendations by volume tier
At low volume, use the strongest model and optimise later. At high volume, route routine tickets to a smaller model and escalate the rest.