How to evaluate LLM outputs: test sets, rubrics, and automated checks
By TechlyUpUpdated 2 min readDevelopers and AI engineers
Quick answer
Evaluate LLM features like any other software: a representative test set, clear pass criteria, and automated runs on every change. Use exact checks where outputs are structured, rubric-based scoring (by people or a carefully validated model grader) for open-ended text, and track results over time. Evaluation is what lets you change prompts and models without guessing.
Build a representative test set
Sample real (or realistic synthetic) inputs across the variety you expect, including hard and adversarial cases. Label expected outcomes or acceptance criteria for each.
Choose the right check for each output
Match the check to the output type.
- Structured output: schema validation and exact-match fields.
- Classification: accuracy, precision/recall per class, confusion matrix.
- Extraction: field-level correctness against labelled data.
- Free text: rubric scores for correctness, completeness, groundedness, and tone.
Model-graded evaluation with care
Using a model to grade outputs scales well but must be validated against human judgements on a sample. Give the grader a specific rubric and examples, and watch for bias toward longer or more confident answers.
Make it continuous
Run evaluations in CI when prompts, models, or retrieval change. Keep a history so you can see regressions and improvements.
eval results — prompt v7 vs v8 schema_valid: 100% → 100% category_accuracy: 91% → 94% hard_cases: 12/20 → 15/20 cost_per_100: ₹X → ₹Y (fill from your own billing)
Evaluation mistakes to avoid
These make evaluation results misleading.
- Building the test set only from easy, typical cases.
- Changing the test set every time, so results aren't comparable.
- Trusting a model grader without checking it against human judgement.
- Reporting a single average that hides serious failures in important categories.
Worked example: evaluating a summariser
A team builds 40 test documents with human-written reference points: key facts that must appear and statements that must not. Each summary is scored on coverage of key facts, absence of unsupported claims, and length limits.
They validate a model grader on 20 summaries by comparing with two human reviewers, adjust the rubric wording until agreement is good, and then run it in CI. When a model update is released, they know within an hour whether quality moved — rather than finding out from users.
Try it yourself
Write a rubric with four criteria for a summarisation feature, score 20 outputs yourself, then compare with a model grader's scores.
Frequently asked questions
How big should an evaluation set be?
Start with dozens of well-chosen cases and grow it as you find failures in production.
Can I trust LLM-as-judge?
Only after checking its agreement with human reviewers on your task. Use it to scale, not to replace human judgement entirely.
What should I evaluate besides accuracy?
Latency, cost, safety behaviour, refusal rates, and consistency across runs.
Want a suggested next step for your situation?
Share a few details and someone from TechlyUp will get back to you. No automated sequences.
Sources and further reading
- NIST AI Risk Management Framework
- Hugging Face LLM Course
- Microsoft: Red teaming large language models
Examples are authored practice material, not measured learner outcomes. Tool behavior can change. Found an error? Contact TechlyUp with the page URL and correction.