CCAR-P · Study Guide

← Glossary

Domain 4: Evaluation, Testing & Optimization

Eval terms: metrics, datasets, A/B tests, diagnosis, optimization, and monitoring loops.

  • Balanced metric set

    An evaluation scorecard covering accuracy/quality, latency, cost, safety, and security — not accuracy alone.

    Exam Accuracy-only dashboards are incomplete.

    See also: 4.1 Evaluation metrics
  • Safety / security metrics

    Explicit measures such as policy-block rate, unauthorized tool attempts blocked, and PII leakage incidents — separate from task accuracy.

    Exam Tool agents need security metrics, not only quality scores.

    See also: 4.1 Evaluation metrics
  • Mixed-methodology evals

    Combining automated scorers, human review, and scenario/tool-path tests so judgment-heavy and integration failures are both covered.

    Exam Auto-only or friendly-only sets miss real failure modes.

    See also: 4.2 Eval datasets & frameworks
  • Versioned eval dataset

    An evaluation set pinned to a version ID alongside the system under test so regressions stay comparable over time.

    Exam Version datasets with the system — floating unversioned sets hide drift.

    See also: 4.2 Eval datasets & frameworks
  • Confounded A/B

    An experiment that changes multiple factors (model, prompt, retrieval) at once so a win cannot be attributed to one change.

    Exam Isolate the primary factor or design a proper multi-factor test.

    See also: 4.3 A/B testing
  • Ship criteria

    Pre-agreed thresholds for quality, cost, and other SLAs that a winning variant must clear before release.

    Exam Ship on thresholds — not vibes or "more tokens."

    See also: 4.3 A/B testing
  • Prompt failure

    The model is capable but misses format, instructions, or constraints — fixed by clearer criteria/examples, not necessarily a larger model.

    Exam Reproduce first; do not upgrade model for format drift alone.

    See also: 4.4 Diagnose system issues
  • Hallucination / grounding failure

    Outputs invent facts or cite sources that were never retrieved. Diagnose with retrieval provenance before swapping models.

    Exam Citations with empty retrieval provenance → grounding check first.

    See also: 4.4 Diagnose system issues
  • Model mismatch

    Consistent capability gaps on a task class after retrieval and prompts look correct — the model tier or capability does not fit the work.

    Exam Suspect mismatch when failures are consistent and provenance is fine.

    See also: 4.4 Diagnose system issues
  • Token / cost optimization order

    Profile → trim context/tool output → cache stable prefixes → then reconsider model tier — re-check quality at each step.

    Exam Do not cut safety or ignore the quality floor to save cost.

    See also: 4.5 Optimize token/latency/cost
  • Interactive vs batch paths

    Keep chat on a low-latency path; batch overnight or non-interactive work where waiting is acceptable.

    Exam Do not batch the interactive chat to save cost.

    See also: 4.5 Optimize token/latency/cost
  • Quality monitoring loop

    Always-on quality and cost signals with alerts that route to runbooks and new eval fixtures — not HTTP uptime alone.

    Exam Green uptime with silent quality drop means the monitoring set is incomplete.

    See also: 4.6 Monitor with observability
  • Alert → eval fixture

    Closing the ops loop by capturing failing cases from alerts into the regression suite so the same failure is caught offline next time.

    Exam Alerts without fixture capture leave the same hole open.

    See also: 4.6 Monitor with observability