Domain 4 · 16% of the exam
Evaluation, Testing & Optimization
Define metrics, build evals, run A/B tests, diagnose failures, and optimize cost and performance.
Objectives
- 4.1
Define evaluation metrics (accuracy, latency, cost, safety, security)
Define a balanced metric set covering accuracy/quality, latency, cost, safety, and security — tied to user and business outcomes, not a single vanity score.
- 4.2
Design evaluation datasets and test frameworks using mixed methodologies
Build versioned eval datasets and harnesses that mix automated checks, human review, and scenario tests — covering happy path, edge, and adversarial cases.
- 4.3
Conduct A/B testing and iterative improvements
Run controlled A/B tests on prompts, retrieval, and model settings — isolate factors, power the experiment, and ship only when quality and cost clear agreed bars.
- 4.4
Diagnose system issues (prompt failure, hallucinations, model mismatch)
Separate prompt failure, hallucination/grounding failure, retrieval miss, and model mismatch before fixing — reproduce with fixed inputs and traces.
- 4.5
Optimize token usage, latency, and cost-performance trade-offs
Cut tokens and latency without silently dropping required quality — cache and trim before buying capacity; re-measure quality after each cut; batch when interactivity is not required.
- 4.6
Monitor system performance using logging and observability tools
Operate with logs and observability that catch quality and cost regressions — alert beyond HTTP uptime, keep debug context, and close the loop into eval suites.