4.2 · Lesson 2 of 6
Design evaluation datasets and test frameworks using mixed methodologies
What You Need to Know
Evaluation datasets and test frameworks prove whether a Claude system is ready — they are not a box of friendly examples that always pass. Professional architects design versioned sets that cover happy path, edge cases, and adversarial cases, and they mix automatic scorers with human labeling where judgment is required. Tiny never-fail suites are exam traps: they create false confidence and miss regressions.
Frameworks must include more than offline prompt–response pairs. Scenario and tool-path tests belong in the harness for RAG and agents: empty retrieval, conflicting chunks, tool failures, and multi-turn policy flows. Version datasets with the system under test so every release can be compared apples-to-apples.
What a mixed-method suite includes
- Happy, edge, and adversarial cases with expected behaviors
- Automatic checks for structure, citations, latency, and retrieval hits
- Human (or expert) labels for faithfulness, nuance, and tone
- Versioned cases pinned to config snapshots
- Scenario and tool-path tests, not only static offline prompts
Decision rules
- Cover happy path, edge cases, and adversarial cases — not only friendly examples.
- Mix automatic scorers with human labeling where judgment is required.
- Version datasets with the system under test for regression comparability.
- Tiny never-fail sets are exam traps; they do not prove readiness.
- Scenario and tool-path tests belong in the framework, not only offline prompts.
Why friendly-only sets fail
Systems fail on the hard cases: conflicting policy, missing evidence, injected instructions in retrieved text, and brittle tool routes. A suite of twelve friendly examples that never fail cannot surface those failures or detect silent regressions when prompts or indexes change. Architects therefore treat coverage design and versioning as first-class design work, not a last-minute checklist.
Review checklist
- Does the set include edge and adversarial cases, not only happy paths?
- Are auto scores mixed with human judgment where needed?
- Is the dataset versioned and pinned to each eval run?
- Do scenario/tool-path tests cover retrieval and tool failures?
- Would a never-fail tiny set be rejected in design review?
Exam stems that boast perfect scores on a handful of friendly examples are inviting you to call the suite incomplete. The professional move is mixed methodologies, hard cases, and versioned regression harnesses.
Exam application
Tiny friendly datasets that never fail are traps; mixed methodologies and hard cases are required. Distractors: ship on never-fail demos, auto-only when judgment is needed, or prompt tweaks instead of fixing coverage. Prefer versioned sets plus scenario/tool-path tests.
Exam traps
Friendly-only datasets that never fail
Happy-path-only suites create false confidence. Exam items reward edge and adversarial coverage.
No dataset versioning
Without versions pinned to the system under test, you cannot compare regressions across releases.
Auto-only scoring when judgment is required
Policy nuance, tone, and citation faithfulness often need human labels mixed with automatic checks.
Offline prompts only — no scenario or tool-path tests
RAG and tool agents fail on retrieval and tool routes that never appear in static prompt-response pairs.
Practice scenario
A team ships a RAG policy assistant after running twelve friendly examples that always pass. They have no edge cases, no adversarial prompts, and no versioned dataset. What should the architect conclude?
Build exercise
Design a mixed-method eval suite for a RAG policy assistant with versioned cases
40 minutes
What you'll learn
- Plan happy, edge, and adversarial coverage
- Mix automatic scorers with human labels
- Version datasets against config snapshots
- Add scenario and tool-path tests to the harness
Step 1
Map coverage: happy, edge, and adversarial
For the RAG policy assistant, list must-pass happy paths, known edge policies (conflicts, outdated docs, multi-hop), and adversarial cases (jailbreaks, prompt injection via retrieved text, out-of-policy asks).
Why: Coverage design comes before harvesting examples. Friendly-only lists are incomplete by construction.
You should see: A coverage matrix with at least three rows per bucket and expected behaviors.
Step 2
Mix automatic scorers with human labeling
Use automatic checks for schema, citation presence, refusal keywords, and retrieval hit rates. Route faithfulness, policy nuance, and tone to human (or expert) labels on a sampled slice.
Why: Mixed methodologies catch both mechanical regressions and judgment failures auto scores miss.
You should see: Harness config showing auto checks plus a human-label queue with rubric.
Step 3
Version the dataset with the system under test
Store cases with ids, expected outcomes, source doc versions, and a dataset semver. Pin each eval run to dataset version + model/prompt/retrieval config hashes.
Why: Regression comparability requires versioning. Unversioned spreadsheets break auditability.
You should see: A versioned case file (or eval registry entry) linked to a config snapshot.
Step 4
Add scenario and tool-path tests to the framework
Beyond offline Q&A, script multi-turn scenarios and tool/retrieval paths: wrong corpus, empty retrieval, tool failure, and citation UI contracts. Run the suite on every candidate release.
Why: Offline prompts alone miss integration failures. Architects include scenario/tool-path tests in the framework.
You should see: Scenario scripts in CI plus a report that separates offline vs scenario results.
Sources
- Develop test cases — docs.anthropic.com
- Evaluation tooling — docs.anthropic.com
- Retrieval augmented generation — docs.anthropic.com — RAG systems need retrieval-aware evals