Generative AI · Evals · AI Agents

Pick the grader before you write the question: eval design for a grounded assistant

Published · Updated

An eval set is not a bug log. If you cannot name the grading method before writing the question, the question ends up on someone's desk, and manual review does not survive a daily pipeline.

An assistant eval usually starts as a list of the questions that broke it last week. That list is cheap to make and misleading to act on: fix the ten failures and the suite stops telling you anything new. Anthropic's guidance on building tests puts the constraint first, which is to design evals that mirror your real-world task distribution. The set is a sample of what users actually ask, not an incident archive.

Sample the traffic, not the incidents

If a grounded assistant answers product questions in seven languages across four country domains, the eval set has to look like that mix. Our multilingual platform for a biomass boiler manufacturer serves seven languages across four separately indexed country domains, so a suite that only tests the home language tests the easy half. Weighting questions by where requests actually land is less satisfying than collecting dramatic failures, and far more predictive.

Pick the grading method before writing the question

The guidance is explicit about format: structure questions to allow for automated grading, such as multiple-choice, string match, code-graded or LLM-graded. Read that backwards and it becomes a design rule. If you cannot name the grader before you write the question, the question will end up on someone's desk for manual review, and manual review is the one step that does not scale alongside a pipeline that ships every day.

In practice this means rewriting open prompts into checkable ones. 'Is this answer good?' has no grader. 'Does the answer cite a retrieved passage, stay in the requested language, and avoid a price claim?' has three, each of them a string match, a classifier or a code check. The rewrite costs you some realism and buys a suite that can run on every deploy with nobody in the loop.

More questions, slightly lower signal

There is a trade the guidance states plainly: more questions with slightly lower-signal automated grading are better than fewer questions with high-quality human hand-graded evals. That feels wrong if you have ever read a hand-graded transcript and seen how much nuance the grader caught. The nuance is real. It does not survive contact with a weekly release cadence, and a suite that runs once a month is documentation rather than a guardrail.

One number hides the failure that matters

Most use cases need multidimensional evaluation along several success criteria. A grounded assistant fails in distinct ways: it answers from model memory instead of retrieval, it answers in the wrong language, or it answers confidently about something absent from the corpus. Collapsing those into a single pass rate lets a grounding regression hide behind a formatting improvement. Score the dimensions separately, and read them separately, even when a summary number is wanted.

The failure modes you are actually grading

OWASP's misinformation entry names the pair worth grading. Misinformation occurs when models produce false or misleading information that appears credible, and overreliance occurs when users place excessive trust in that output and fail to verify its accuracy. The mitigations listed are retrieval of relevant and verified information from trusted external databases during generation, plus tools and processes that automatically validate key outputs in high-stakes environments. The eval suite is where you find out whether retrieval changed behaviour or merely sat in the request path.

Why this matters more when nobody is watching

Our personal publishing platform writes and ships articles daily with three model fallback levels and roughly zero hours of weekly content operations. The multilingual platform publishes eight SEO articles every day with no manual intervention. Neither arrangement is comfortable without automated output checks, because fallback means the model producing an answer is not always the model you tested. The validation has to live beside the output rather than beside one chosen model.

Treat the grading method as part of the question rather than a later step. Sample the distribution you actually serve, convert every open prompt into something a grader can decide, accept a slightly noisier per-question signal in exchange for a suite that runs on every change, and report the dimensions apart instead of averaged. That is the version still worth trusting on a Friday evening, when the pipeline is publishing and nobody is reading the output.

Where to start if you have no suite

Sources

Claude Platform Docs (Anthropic) — Define success criteria and build evaluations — https://platform.claude.com/docs/en/test-and-evaluate/develop-tests

OWASP Gen AI Security Project — LLM09:2025 Misinformation — https://genai.owasp.org/llmrisk/llm092025-misinformation/

Neurolinks case study — Four markets, one codebase — https://neurolinks.be/work/tatano-energy

Neurolinks case study — A personal brand that publishes itself — https://neurolinks.be/work/matthieu-pesesse-media

Working on a project where these methods apply?