How to build an evaluation process that tells you whether a prompt, model or retrieval change made your AI feature better or worse.
Why evaluation comes first
LLM outputs vary, and small changes to a prompt or model can improve one case while breaking another. Without a repeatable evaluation, every change is a guess. A modest evaluation set turns those decisions into measurable comparisons.
Build an evaluation dataset
Collect real or realistic inputs, covering easy, typical and difficult cases, including ones the system should refuse or escalate. For each, write down what a good answer must contain or do. Start with a few dozen examples and grow the set from production failures over time.
Define clear criteria
Good criteria describe observable behaviour, such as "states the refund window from the policy" or "does not invent an order number", rather than "high quality". Separate criteria for correctness, grounding in sources, completeness, format and tone make results easier to act on.
Automated checks
Some criteria can be checked with code: valid JSON, required fields present, length limits, or whether a cited source was actually retrieved. These checks are fast and cheap, so run them on every change.
LLM-as-a-judge
A second model can score outputs against a rubric, which scales better than human review. Research shows strong models often agree with human raters, but judges can favour longer answers or their own style. Calibrate the judge against a sample of human scores and use specific rubrics, ideally with pass or fail decisions per criterion.
Human review
People remain the reference point, especially for subjective or high-risk outputs. Review a sample regularly, use two reviewers on a subset to check agreement, and feed disagreements back into clearer criteria.
Checking for hallucinations
For grounded tasks, check whether every claim in the answer is supported by the retrieved context. This can be done with a judge model prompted to compare claims with sources, or with dedicated evaluation tools. Track unsupported claims as their own metric.
Evaluation in production
Log inputs, outputs, retrieved context and user feedback. Sample conversations for review, track metrics over time, and add new failure cases to the dataset. Run the full evaluation before every release so regressions are caught before users see them.
Key takeaways
- A small, real evaluation set beats none at all.
- Write criteria as observable behaviours.
- Calibrate LLM judges against human scores.
- Keep evaluating after launch and grow the dataset from real failures.
If you need an evaluation process for an AI feature before launch, our engineers can help you set one up.
Generative AI Development- Judging LLM-as-a-Judge with MT-Bench and Chatbot ArenaZheng et al. · 2023 · Research paperarxiv.org
- Evals guideOpenAI · Documentationplatform.openai.com
- Define success criteria and build evaluationsAnthropic, Claude Docs · Documentationplatform.claude.com