Back
Reports & Research

Inference Cost Benchmarks for Customer Support Workloads

We ran the same support-ticket workload across open and proprietary models and compared answer quality, latency and cost per resolved ticket.

Format
Research · PDF
Length
10 min read
Updated
Aug 2026
Topics
Generative AI, Enterprise AI

Method

Each model answered the same anonymised set of support tickets using an identical retrieval setup. Answers were scored by trained reviewers against a shared rubric.

How to use the results

The cheapest model per token is rarely the cheapest per resolved ticket. The report shows where smaller models are good enough, and where escalation to a larger model pays for itself.

View full reportResearch · PDF · 10 min read
Open report