CCAR-P · Study Guide

← Domain 4: Evaluation, Testing & Optimization

4.2 · Lesson 2 of 6

Design evaluation datasets and test frameworks using mixed methodologies

What You Need to Know

Evaluation datasets and test frameworks prove whether a Claude system is ready — they are not a box of friendly examples that always pass. Professional architects design versioned sets that cover happy path, edge cases, and adversarial cases, and they mix automatic scorers with human labeling where judgment is required. Tiny never-fail suites are exam traps: they create false confidence and miss regressions.

Frameworks must include more than offline prompt–response pairs. Scenario and tool-path tests belong in the harness for RAG and agents: empty retrieval, conflicting chunks, tool failures, and multi-turn policy flows. Version datasets with the system under test so every release can be compared apples-to-apples.

What a mixed-method suite includes

  • Happy, edge, and adversarial cases with expected behaviors
  • Automatic checks for structure, citations, latency, and retrieval hits
  • Human (or expert) labels for faithfulness, nuance, and tone
  • Versioned cases pinned to config snapshots
  • Scenario and tool-path tests, not only static offline prompts

Decision rules

  • Cover happy path, edge cases, and adversarial cases — not only friendly examples.
  • Mix automatic scorers with human labeling where judgment is required.
  • Version datasets with the system under test for regression comparability.
  • Tiny never-fail sets are exam traps; they do not prove readiness.
  • Scenario and tool-path tests belong in the framework, not only offline prompts.

Why friendly-only sets fail

Systems fail on the hard cases: conflicting policy, missing evidence, injected instructions in retrieved text, and brittle tool routes. A suite of twelve friendly examples that never fail cannot surface those failures or detect silent regressions when prompts or indexes change. Architects therefore treat coverage design and versioning as first-class design work, not a last-minute checklist.

Review checklist

  • Does the set include edge and adversarial cases, not only happy paths?
  • Are auto scores mixed with human judgment where needed?
  • Is the dataset versioned and pinned to each eval run?
  • Do scenario/tool-path tests cover retrieval and tool failures?
  • Would a never-fail tiny set be rejected in design review?

Exam stems that boast perfect scores on a handful of friendly examples are inviting you to call the suite incomplete. The professional move is mixed methodologies, hard cases, and versioned regression harnesses.

Exam application

Tiny friendly datasets that never fail are traps; mixed methodologies and hard cases are required. Distractors: ship on never-fail demos, auto-only when judgment is needed, or prompt tweaks instead of fixing coverage. Prefer versioned sets plus scenario/tool-path tests.

Exam traps

  • Friendly-only datasets that never fail

    Happy-path-only suites create false confidence. Exam items reward edge and adversarial coverage.

  • No dataset versioning

    Without versions pinned to the system under test, you cannot compare regressions across releases.

  • Auto-only scoring when judgment is required

    Policy nuance, tone, and citation faithfulness often need human labels mixed with automatic checks.

  • Offline prompts only — no scenario or tool-path tests

    RAG and tool agents fail on retrieval and tool routes that never appear in static prompt-response pairs.

Practice scenario

A team ships a RAG policy assistant after running twelve friendly examples that always pass. They have no edge cases, no adversarial prompts, and no versioned dataset. What should the architect conclude?

Choose one answer

Build exercise

Design a mixed-method eval suite for a RAG policy assistant with versioned cases

40 minutes

What you'll learn

  • Plan happy, edge, and adversarial coverage
  • Mix automatic scorers with human labels
  • Version datasets against config snapshots
  • Add scenario and tool-path tests to the harness
  1. Step 1

    Map coverage: happy, edge, and adversarial

    For the RAG policy assistant, list must-pass happy paths, known edge policies (conflicts, outdated docs, multi-hop), and adversarial cases (jailbreaks, prompt injection via retrieved text, out-of-policy asks).

    Why: Coverage design comes before harvesting examples. Friendly-only lists are incomplete by construction.

    You should see: A coverage matrix with at least three rows per bucket and expected behaviors.

  2. Step 2

    Mix automatic scorers with human labeling

    Use automatic checks for schema, citation presence, refusal keywords, and retrieval hit rates. Route faithfulness, policy nuance, and tone to human (or expert) labels on a sampled slice.

    Why: Mixed methodologies catch both mechanical regressions and judgment failures auto scores miss.

    You should see: Harness config showing auto checks plus a human-label queue with rubric.

  3. Step 3

    Version the dataset with the system under test

    Store cases with ids, expected outcomes, source doc versions, and a dataset semver. Pin each eval run to dataset version + model/prompt/retrieval config hashes.

    Why: Regression comparability requires versioning. Unversioned spreadsheets break auditability.

    You should see: A versioned case file (or eval registry entry) linked to a config snapshot.

  4. Step 4

    Add scenario and tool-path tests to the framework

    Beyond offline Q&A, script multi-turn scenarios and tool/retrieval paths: wrong corpus, empty retrieval, tool failure, and citation UI contracts. Run the suite on every candidate release.

    Why: Offline prompts alone miss integration failures. Architects include scenario/tool-path tests in the framework.

    You should see: Scenario scripts in CI plus a report that separates offline vs scenario results.

Sources