What Are LLM Evals and How Do You Evaluate LLM Outputs?

LLM evals are structured tests that score model or agent outputs against defined criteria so you can compare versions and catch quality regressions. Use them when you need to choose between prompts, models, or agent configurations, or when you want an automated signal that quality changed after a deployment. They are not the same as observability: observability tells you what happened on live traffic, while evals tell you how good the output was on a controlled set of cases.

Evals vs. observability vs. monitoring

These three are often conflated, but they answer different questions and run at different times.

Practice Question it answers When it runs Typical input
LLM evals How good is this version on known cases? Before/after a change, on demand Curated test set with criteria
LLM observability What did the system actually do? Continuously, on live traffic Production traces and spans
LLM monitoring Did something change or break? Continuously, with alerts Metrics, logs, error rates

A practical split: use observability to see real agent behavior and collect interesting production cases, then promote those cases into your eval set. Respan's positioning reflects this pairing — it describes itself as an AI router with built-in observability and automated evals, and its page headings cover routing across models, shipping the best-scoring version, knowing when things change, and seeing what agents did. That is the loop: route, observe, evaluate, ship.

Core evaluation methods

There is no single metric that covers everything. Most teams combine two or three of these.

Automated metrics

Deterministic checks that need no model judgment. Good for tasks with a verifiable answer.

  • Exact match or normalized match for classification, extraction, or short answers.
  • Contains/regex checks for required fields, formats, or forbidden strings.
  • JSON/schema validity for structured output.
  • Retrieval metrics (precision, recall, hit rate) when the task depends on fetched context.

These are cheap, fast, and reproducible. They fail on open-ended generation, where many correct answers exist.

LLM-as-judge

A model scores outputs against a rubric. Useful for summarization, tone, helpfulness, and other qualities that resist exact matching.

  • Write the rubric explicitly: what counts as pass, what counts as fail, and how ties are handled.
  • Give the judge the input, the output, and the reference answer if you have one.
  • Ask for a score plus a short justification so you can audit disagreements.
  • Watch for known biases: judges tend to prefer longer answers and their own model family's style. Randomize or swap answer order when comparing two outputs.

Human review

People label a sample of outputs, usually for calibration rather than full coverage. Use it to validate that your automated metric actually tracks what you care about. If the judge and humans disagree often, fix the rubric before trusting the judge at scale.

Building a test set

The eval set is the part most teams underinvest in. A small, well-chosen set beats a large, noisy one.

  1. Start from real usage. Pull representative inputs from production traces rather than inventing them.
  2. Cover the edges. Include ambiguous requests, multi-step tasks, and cases where the correct behavior is to refuse or ask a clarifying question.
  3. Write expected outputs or criteria. For open-ended tasks, criteria ("must cite the source document") work better than a single gold answer.
  4. Label each case with the failure mode it targets, so a score drop tells you what broke, not just that something broke.
  5. Keep a held-out slice. Do not tune prompts against every case you own, or you will overfit.

Running an eval to compare versions

The workflow is the same whether you compare prompts, models, or agent versions.

  1. Freeze the variables. Change one thing at a time — prompt, model, or agent config — and hold the rest constant.
  2. Run every version on the same test set. Same inputs, same order, same judge settings.
  3. Score each run with your chosen metrics and record per-case results, not just the average.
  4. Compare. Look at the aggregate score and the per-case diff. A version that wins on average but breaks a critical case is usually not the one to ship.
  5. Ship the best-scoring version, then keep the eval set in CI so future changes are checked against it.

Respan frames the outcome as "ship the version that scores best," which is the decision this workflow is meant to support.

Catching regressions

A regression is a quality drop caused by a change that looked unrelated. Common triggers: a model version bump, a prompt edit, a retrieval change, or a new tool in an agent.

  • Re-run the eval set on every prompt, model, or agent change, ideally automatically.
  • Alert on score drops beyond a threshold you set, not on every fluctuation.
  • Track per-case pass/fail over time so you can see which categories degrade.
  • Keep a small "canary" subset of critical cases that must never fail.

Common pitfalls

  • Overfitting to a small eval set. If you iterate against the same 20 cases, you optimize for those cases, not the task. Hold out a slice and refresh cases from production.
  • Relying on a single metric. One number hides tradeoffs. Track at least one correctness metric and one quality metric.
  • Trusting an unvalidated judge. Check the judge against human labels before scaling it.
  • Eval sets that drift from reality. Production inputs change; a set built six months ago may no longer represent your users.
  • No baseline. Without a recorded baseline score, you cannot tell whether a change helped or hurt.

Where to start

If you are setting this up for the first time: pick one task, collect 30–50 real inputs, define pass/fail criteria, and run a single automated metric plus an LLM judge. Compare two prompt versions and record per-case results. Once that loop works, add it to CI and grow the set from production traces.

respan.ai
The AI router with built-in observability & automated evals.