Method
Each model answered the same anonymised set of support tickets using an identical retrieval setup. Answers were scored by trained reviewers against a shared rubric.
How to use the results
The cheapest model per token is rarely the cheapest per resolved ticket. The report shows where smaller models are good enough, and where escalation to a larger model pays for itself.
View full reportResearch · PDF · 10 min read
Open report