3.3 · Lesson 3 of 8
Evaluate accuracy-latency trade-offs and justify configuration decisions
What You Need to Know
Accuracy–latency trade-offs sit at the center of Domain 3 configuration decisions. Model tier, retrieval depth, tool fan-out, and prompting style all move both quality and speed. Architects must state the accuracy floor and the latency (and often cost) SLA, quantify what each change buys, and refuse one-sided optimizations that violate a hard constraint without stakeholder waiver.
A useful pattern is to meet the accuracy floor with the smallest, fastest configuration that clears evals, then spend latency budget only where measured quality still fails. Cache stable prefixes, limit retrieval hops, and reserve heavier models for escalated or asynchronous paths when interactive SLAs are tight.
Decision rules
- Always name both bars: quality floor and latency/cost SLA.
- Quantify Δ accuracy, Δ latency, and Δ cost for every config proposal.
- If the floor is already met, prefer the faster/cheaper option unless a pillar demands otherwise.
- Reject or redesign changes that break a hard SLA without an explicit waiver.
- Profile before upgrading model tier — many wins come from caching and retrieval caps.
Levers that move both metrics
- Model tier and max output tokens
- Retrieval depth, hop count, and top-k
- Tool fan-out and sequential versus parallel calls
- Prompt length and whether stable prefixes are cached
- Whether rare hard cases stay on the interactive path
Decision frame
Write the accuracy floor and the latency/cost SLA before debating options. If the floor is unmet, spend latency budget up to the SLA. If the floor is met, prefer the faster and cheaper configuration unless a named business pillar explicitly pays for marginal quality. Escalate a waiver only with alternatives: reject, redesign (cache/index), or accept a slower premium path for a subset of queries.
Official-style items punish answers that celebrate accuracy while ignoring a stated p95. Ship criteria are dual — quality and contract — not vibes.
Exam application
Read the SLA first. If a candidate breaks p95 without a waiver, reject or redesign even when accuracy rises. Prefer caching and retrieval caps when the accuracy floor is already met. Distractors ignore cost/latency or hide monitoring.
Exam traps
Optimizing accuracy in isolation
Exam stems often give a latency or cost SLA. Ignoring it is a wrong answer even when accuracy rises.
Assuming larger models are always justified
If the accuracy floor is already met, extra capability that blows the SLA needs a business waiver or a redesign (cache, smaller model on the hot path).
Confusing perceived polish with measured latency
Streaming or UI tricks do not fix a p95 budget exceeded by multi-hop retrieval.
Skipping quantification
"It should be faster" is not a decision. State deltas for accuracy, latency, and cost.
Practice scenario
A live chat assistant has a contractual p95 latency SLA of 800ms and already clears the accuracy floor on the golden set. A proposed change adds a second retrieval hop and a larger model. Offline evals show +4% accuracy and +2.1s p95 latency. Stakeholders have not waived the SLA. What should the architect decide?
Build exercise
Decide a hot-path model and retrieval config
35 minutes
What you'll learn
- Define dual accuracy and latency bars
- Profile latency contributors
- Compare configs with quantified deltas
- Escalate waivers with real options
Step 1
Write the dual bar: accuracy floor + latency/cost SLA
Document the minimum acceptable quality metric and the hard p95 (or budget) constraint before changing model, retrieval depth, or tool fan-out.
Why: Without both bars, teams ship one-sided optimizations that fail production contracts.
You should see: A one-page decision sheet with numeric thresholds and owners.
Step 2
Profile where time goes before buying more model
Break latency into TTFT, retrieval, tool calls, and generation. Identify whether caching a stable prefix or trimming retrieval depth recovers budget.
Why: Many "need a bigger model" proposals are actually retrieval or cold-prefix problems.
You should see: A latency waterfall for the median and p95 request.
Step 3
Run a controlled config comparison on fixed evals
Compare baseline vs candidate on the same golden set and a latency harness. Record accuracy Δ, p95 Δ, and unit cost Δ.
Why: Exam answers reward quantified trade-offs, not anecdotal preference.
You should see: A table with baseline and candidate metrics and a ship/no-ship recommendation.
Step 4
Escalate for a waiver only with options
If stakeholders want the accuracy gain, present options: waive SLA, accept a slower tier for a subset of queries, or invest in caching/indexing to reclaim latency.
Why: Architects communicate trade-offs; they do not silently break contracts.
You should see: A short ADR or decision note with options and recommendation.
Sources
- Models overview — docs.anthropic.com
- Prompt caching — docs.anthropic.com — stable prefix latency/cost
- Context windows — docs.anthropic.com