4.1 · Lesson 1 of 6
Define evaluation metrics (accuracy, latency, cost, safety, security)
What You Need to Know
Evaluation metrics for Claude systems are a balanced scorecard — not a single accuracy chart. Professional architects measure accuracy/quality, latency, cost, safety, and security together, and they tie those measures to user and business outcomes such as task success, deflection with quality, and CSAT. A rising accuracy number that hides slower replies, higher spend, or more policy violations is not a ship signal.
Safety and security need explicit measures: block rates, harmful-output rates, unauthorized tool attempts, and authZ failures on tool-using agents. Before optimizing any one axis, stakeholders must agree floors (quality/safety) and ceilings (latency/cost). Report metric deltas together when proposing config changes so regressions cannot hide behind a vanity win.
Axes of a complete metric pack
- Accuracy/quality tied to graded task success — not demos alone
- Latency (p50/p95) against user-facing SLAs
- Cost per successful task or conversation
- Safety: policy blocks, harmful outputs, escalation correctness
- Security: unauthorized tool attempts, data leakage / authZ failures
Decision rules
- Single-metric dashboards hide latency, cost, safety, and security risk.
- Safety and security need explicit measures (blocks, violations, unauthorized tool attempts).
- Tie metrics to user and business outcomes (task success, CSAT, deflection).
- State floors and ceilings stakeholders agreed before optimizing any one axis.
- Report metric deltas together when proposing config changes.
Why accuracy-only views fail
Accuracy is necessary but insufficient. A model that answers correctly after eight seconds, at triple the unit cost, or by calling tools it should not have, fails the product even if the quality chart looks green. Architects therefore treat floors and ceilings as ship gates: optimize under constraints, never by sacrificing an unbounded axis for a single score.
Review checklist
- Are accuracy, latency, cost, safety, and security all defined with owners?
- Do safety/security metrics cover tool misuse and policy violations?
- Are floors and ceilings written and stakeholder-agreed?
- Are metrics linked to user/business outcomes, not only model scores?
- Would a proposed change show all five deltas before ship?
Exam stems often show a glowing accuracy chart and ask what is missing. The professional move is to reject incomplete dashboards and require a balanced pack with floors, ceilings, and outcome linkage before shipping.
Exam application
Accuracy alone without latency/cost/safety/security is incomplete — prefer balanced scorecards. Distractors: ship on accuracy only, drop technical SLAs for surveys alone, or optimize later under live traffic without gates. Prefer answers that set floors/ceilings and report multi-axis deltas.
Exam traps
Optimizing accuracy alone
Accuracy-only dashboards hide latency, cost, safety, and security regressions. Exam items reward balanced scorecards tied to outcomes.
Ignoring security metrics on tool-using agents
Unauthorized tool attempts, policy violations, and authZ failures are first-class measures — not optional afterthoughts.
Optimizing one axis without floors or ceilings
Without stakeholder-agreed floors (quality/safety) and ceilings (latency/cost), gains on one axis can breach another silently.
Vanity metrics disconnected from user outcomes
Token savings or raw accuracy without task success, deflection, or CSAT do not prove the system works for users.
Practice scenario
A support-copilot dashboard shows only answer accuracy climbing week over week. Leadership wants to ship the latest config. What is incomplete about this decision basis?
Build exercise
Define a metric pack for a support copilot with floors, ceilings, and owners
35 minutes
What you'll learn
- Tie metrics to outcomes and named owners
- Cover accuracy, latency, cost, safety, and security
- Set stakeholder floors and ceilings before optimizing
- Report five-axis deltas for every proposed change
Step 1
Name outcomes and owners for the support copilot
Write the user and business outcomes (task success, deflection with quality, CSAT) and name who owns each metric. Refuse to define targets without owners.
Why: Metrics without owners become vanity charts. Exam stems reward tying measures to outcomes stakeholders accept.
You should see: A short outcome table: metric, owner, user/business link, reporting cadence.
Step 2
Draft the five-axis metric pack
Define measures for accuracy/quality, latency, cost, safety, and security. Include explicit safety blocks and unauthorized tool-attempt rates for the agent.
Why: A balanced pack prevents single-axis blindness. Safety and security need explicit measures on tool agents.
You should see: Five labeled metrics with definitions, units, and data sources.
Step 3
Set floors and ceilings before optimizing
Agree minimum quality/safety floors and maximum latency/cost ceilings with stakeholders. Document that no config ships if it breaches a floor or ceiling — even if accuracy rises.
Why: Optimization without bounds is how teams ship faster, cheaper, or "more accurate" systems that fail users or policy.
You should see: Written floors/ceilings signed off in the metric pack; ship gate checklist referencing them.
Step 4
Report deltas together when proposing a change
For any prompt, model, or tool change, show accuracy, latency, cost, safety, and security side by side versus baseline. Call out which floors/ceilings moved.
Why: Architects justify config with multi-axis evidence. Exam items punish proposing a win on one chart while ignoring regressions elsewhere.
You should see: A change proposal table with five deltas and a pass/fail against floors and ceilings.
Sources
- Test and evaluate overview — docs.anthropic.com
- Reduce hallucinations — docs.anthropic.com — quality and safety measures
- Pricing — docs.anthropic.com — cost as a first-class metric