CCAR-P · Study Guide

← Quick reference

Domain 4: Evaluation, Testing & Optimization

Cheat sheet for Evaluation: metrics, datasets, A/B, diagnosis, optimization, monitoring.

Evaluation metrics (4.1)

Balanced scorecard: accuracy/quality, latency, cost, safety, security — tied to outcomes. Set floors/ceilings before optimizing one axis.

  • Accuracy-only dashboards hide risk
  • Tool agents need security metrics (unauthorized attempts blocked)
  • Report Δ quality, Δ latency, Δ cost together

Eval datasets & frameworks (4.2)

  • Happy path + edge + adversarial
  • Mix automated scorers and human review
  • Version datasets with the system under test
  • Include scenario/tool-path tests — not offline prompts only
  • Tiny never-fail sets are traps

A/B testing (4.3)

  • Isolate one primary factor (or design multi-factor properly)
  • Power and traffic matter — anecdotes are not enough
  • Ship only when quality + cost + SLAs clear thresholds
  • Confounded model+prompt+retrieval releases cannot attribute wins

Diagnose issues (4.4)

Reproduce with fixed inputs/traces. Separate prompt failure, grounding/hallucination, retrieval miss, and model mismatch before fixing.

  • Citations with empty provenance → check retrieval/grounding first
  • Consistent capability gaps after good retrieval → model mismatch
  • Do not jump to a larger model without diagnosis

Optimize tokens / latency / cost (4.5)

  • Profile → trim → cache stable prefixes → reconsider tier
  • Re-check quality after every cost cut
  • Batch non-interactive work; keep chat hot path lean
  • Never delete safety to save cost

Monitor with observability (4.6)

  • Quality and cost alerts — not HTTP uptime alone
  • Redacted prompt/tool context for debug
  • Join request_id to outcomes
  • Alert → runbook → eval fixture capture