How to Evaluate LLM Applications: A Practical Guide
How to evaluate LLM applications: choosing metrics, building test sets, using LLM-as-judge safely, picking eval tools, and running regression tests in CI.

To evaluate LLM applications well, you need three things: a test set of realistic inputs, clear pass/fail criteria for each one, and an automated way to run them every time you change a prompt, model or retrieval setting. Public benchmarks tell you little about your chatbot or agent. A few hundred well-chosen cases from your own domain tell you much more. This guide shows how we set that up.
Why LLM evaluation is different from normal testing
Traditional software tests check deterministic outputs: add(2, 2) must return 4. LLM outputs are variable, open-ended and often have several acceptable answers. That leads to three problems:
- You cannot string-match most answers. “Your order ships Monday” and “It will be dispatched on Monday” are both correct.
- Small changes have wide effects. A prompt tweak that fixes one case can quietly break ten others.
- Quality has several dimensions. An answer can be accurate but rude, or friendly but wrong.
A good evaluation setup mixes cheap deterministic checks with model-based grading and a small amount of human review.
Step 1: Define what “good” means
Before choosing tools, write down success criteria that someone could actually check. Anthropic’s guide on defining success criteria makes the same point: criteria should be specific and measurable.
| Vague goal | Measurable criterion |
|---|---|
| “Answers should be accurate” | Answer is supported by the retrieved documents; no unsupported claims |
| “Be helpful” | Resolves the user’s question without asking for info already given |
| “Stay on brand” | No competitor recommendations; no pricing promises outside the price list |
| “Return valid data” | Output parses against the JSON schema; all required fields present |
| “Be fast” | 95th percentile latency under a target you set |
| “Be safe” | Refuses out-of-scope requests (legal advice, medical dosing) politely |
Step 2: Build a test set from real traffic
Your eval set is the most valuable asset in the whole process. Sources, in order of value:
- Real user queries from logs or support tickets, with personal data removed.
- Known failures: every bug report becomes a test case.
- Edge cases written by domain experts: ambiguous questions, missing info, adversarial phrasing, other languages.
- Synthetic variations generated by an LLM to widen coverage. Review these; do not trust them blindly.
Practical tips:
- Start with 50–100 cases. A small set you actually run beats a large one you never finish.
- Tag each case by category (billing, returns, technical) so you can see where quality moves.
- Store expected behaviour, not just expected text: “must mention the 30-day return window”, “must not promise a refund”.
- Version the test set alongside your code.
For retrieval systems, also record which documents should be retrieved. That lets you measure retrieval separately from generation. See our RAG best practices for more.
Step 3: Choose metrics that match the task
Deterministic checks (cheap, run on everything)
- Schema validity, required fields, allowed enum values
- Exact or fuzzy match for classification and extraction
- Regex or keyword checks (“contains a link to the returns page”)
- Length limits, forbidden phrases, language detection
- Latency, token usage and cost per request
Retrieval metrics (for RAG)
- Recall@k: did the right document appear in the top k results?
- Precision / MRR: how high did it rank?
- Context relevance: how much of the retrieved text is actually useful?
Model-graded metrics (for open-ended output)
- Faithfulness / groundedness: are claims supported by the provided context?
- Answer relevance: does it answer the question asked?
- Rubric scores: tone, completeness, following policy
- Pairwise preference: is version B better than version A?
Agent metrics
For tool-using agents, grade the trajectory as well as the final answer. Did it call the right tools with valid arguments, avoid unnecessary steps and stop when done? See agentic AI design patterns for the architectures these tests cover.
Step 4: Use LLM-as-judge carefully
Using a strong model to grade outputs scales far better than human review. The research is encouraging but also comes with caveats. The paper Judging LLM-as-a-Judge found that strong judges agree well with human preferences, and it also documented position bias (favouring the first answer), verbosity bias (favouring longer answers) and self-enhancement bias (favouring outputs from the same model).
How we reduce those problems:
- Grade against a rubric, not “is this good?” Give the judge explicit criteria and ask for a short reason before the score.
- Prefer binary or small scales. “Pass/fail: does the answer mention the return window?” is more reliable than a 1–10 score.
- Swap positions in pairwise comparisons and count only consistent wins.
- Use a different model family from the one being tested, where practical.
- Provide the reference answer or source context so the judge checks facts instead of guessing.
- Calibrate against humans. Have a person label 30–50 cases and check that the judge agrees before you trust it at scale.
- Return structured output (JSON with
scoreandreason) so results are machine-readable.
Example judge prompt:
You are grading a customer-support answer.
<context>{retrieved_documents}</context>
<question>{user_question}</question>
<answer>{model_answer}</answer>
Criteria:
1. Every factual claim is supported by <context>.
2. The answer addresses the question directly.
3. No promises about refunds or delivery dates not stated in <context>.
For each criterion, give a one-sentence reason, then PASS or FAIL.
Return JSON: {"criteria": [{"id": 1, "reason": "...", "result": "PASS"}], "overall": "PASS"}
Step 5: Pick an evaluation tool
You can start with a spreadsheet and a script. Once you have more than a few dozen cases, a framework saves time. As of September 2026, these are the options we see most often:
| Tool | Type | Licence / hosting | Strong at |
|---|---|---|---|
| promptfoo | CLI + YAML configs | Open source (MIT); now part of OpenAI | Prompt/model comparisons, red-teaming, CI |
| DeepEval | Python, pytest-style | Apache 2.0 | Unit-test style evals, G-Eval, RAG and agent metrics |
| Ragas | Python toolkit | Apache 2.0 | RAG metrics, test data generation |
| Inspect | Python framework (UK AI Security Institute) | MIT | Rigorous, reproducible evals; agent and tool tasks |
| Langfuse | Tracing + evals platform | Open source core (MIT), self-hostable | Production traces, datasets, LLM-as-judge on live data |
| Arize Phoenix | Observability + evals | Elastic License 2.0, self-hostable | OpenTelemetry tracing, experiments |
How to choose:
- Developers who want tests next to code: DeepEval or promptfoo.
- RAG-heavy systems: Ragas metrics, or the RAG metrics in DeepEval.
- Research-grade reproducibility: Inspect. See also benchmarking LLMs for research.
- Production monitoring plus evals in one place: Langfuse or Phoenix, especially if data must stay on your own servers.
Step 6: Run regression tests in CI
Evaluation pays off when it runs automatically. The pattern we use:
- On every pull request that touches prompts, model config, retrieval or tools, run a fast subset of 50–100 cases, mostly deterministic checks plus a few judged cases.
- Compare against the main branch baseline, not an absolute number. Fail the build if the pass rate drops by more than a threshold you choose, or if any “critical” tagged case fails.
- Nightly or before release, run the full suite, including slower judged metrics and multiple samples per case to measure variance.
- Post a summary on the PR: pass rate by category, newly failing cases, cost and latency change.
- Pin model versions in config so that a provider update does not silently change results. Re-baseline deliberately when you upgrade.
Keep costs under control by caching model responses for unchanged inputs and using a smaller model as judge for simple criteria. Our LLM cost optimization guide has more.
Step 7: Close the loop with production data
Offline evals catch regressions. Production tells you what you missed.
- Log inputs, outputs, retrieved context and tool calls, with appropriate privacy controls.
- Collect lightweight feedback (thumbs up/down, “did this solve your problem?”).
- Sample a small percentage of live traffic for judged scoring each week.
- Turn every confirmed failure into a new test case.
Key takeaways checklist
- Written, measurable success criteria for each quality dimension
- A versioned test set built from real queries and known failures, tagged by category
- Deterministic checks first; LLM-as-judge only where needed, with rubrics and human calibration
- Retrieval measured separately from generation (for RAG)
- Evals running in CI on every prompt/model change, compared against a baseline
- Pinned model versions and deliberate re-baselining
- Production traces and feedback feeding new test cases
An evaluation suite turns LLM development from trial and error into ordinary engineering. If you want help building evals for a chatbot, RAG system or agent, or integrating them into your CI pipeline, see our AI solutions or talk to us.