← Quick reference
Domain 4: Evaluation, Testing & Optimization
Cheat sheet for Evaluation: metrics, datasets, A/B, diagnosis, optimization, monitoring.
Evaluation metrics (4.1)
Balanced scorecard: accuracy/quality, latency, cost, safety, security — tied to outcomes. Set floors/ceilings before optimizing one axis.
- Accuracy-only dashboards hide risk
- Tool agents need security metrics (unauthorized attempts blocked)
- Report Δ quality, Δ latency, Δ cost together
Eval datasets & frameworks (4.2)
- Happy path + edge + adversarial
- Mix automated scorers and human review
- Version datasets with the system under test
- Include scenario/tool-path tests — not offline prompts only
- Tiny never-fail sets are traps
A/B testing (4.3)
- Isolate one primary factor (or design multi-factor properly)
- Power and traffic matter — anecdotes are not enough
- Ship only when quality + cost + SLAs clear thresholds
- Confounded model+prompt+retrieval releases cannot attribute wins
Diagnose issues (4.4)
Reproduce with fixed inputs/traces. Separate prompt failure, grounding/hallucination, retrieval miss, and model mismatch before fixing.
- Citations with empty provenance → check retrieval/grounding first
- Consistent capability gaps after good retrieval → model mismatch
- Do not jump to a larger model without diagnosis
Optimize tokens / latency / cost (4.5)
- Profile → trim → cache stable prefixes → reconsider tier
- Re-check quality after every cost cut
- Batch non-interactive work; keep chat hot path lean
- Never delete safety to save cost
Monitor with observability (4.6)
- Quality and cost alerts — not HTTP uptime alone
- Redacted prompt/tool context for debug
- Join request_id to outcomes
- Alert → runbook → eval fixture capture