4.5 · Lesson 5 of 6
Optimize token usage, latency, and cost-performance trade-offs
What You Need to Know
Optimizing token usage, latency, and cost means cutting waste without silently dropping required quality. Cache stable prefixes and trim tool or context bulk before you buy more capacity or downgrade the model. After every cut, re-measure against the quality floor. Batch when interactivity is not required; keep hot paths lean.
Cost spikes with flat quality usually mean redundant tokens or missed cache — not an automatic model swap. Unbounded max_tokens, stripping safety text, and batching interactive chat are common distractors. The quality SLA binds every cost decision.
Levers in preferred order
- Profile tokens and latency stages before spending
- Cache stable prefixes; trim redundant history and tool payloads
- Set task-appropriate max_tokens and context budgets
- Batch offline work; keep interactive paths synchronous and lean
- Reconsider model tier only after trim/cache and a quality re-check
Decision rules
- Cache stable prefixes and trim tool/context before escalating model tier.
- Measure quality after each cost or latency cut against the floor.
- Batch where interactivity is not required; keep hot paths lean.
- Cost cuts that ignore the quality SLA fail the item.
- Profile where tokens and time go before spending more.
Why trim/cache precedes tier changes
Model tier changes move the entire quality and latency curve. Trim and cache often recover the same cost headroom with less risk when quality is already meeting the floor. Architects treat tier as a controlled experiment after waste is removed — not as the first knob.
Review checklist
- Do you have a hotspot breakdown for tokens and p95 stages?
- Are stable policy and tool schemas eligible for prompt caching?
- Would this cut remove safety or quality-critical instructions?
- Is the workload interactive (no batch) or offline (batch OK)?
- What is the written quality floor you must not breach?
Prefer answers that profile, trim/cache, re-check quality, then reconsider tier. Reject safety deletion, unbounded max tokens, ignored floors, and batching interactive chat.
Exam application
Cost spike + flat quality → trim and cache first, then maybe tier — always re-measure. Quality SLA outranks raw unit-cost reduction.
Exam traps
Delete safety content to save cost
Guardrails and policy packs are not free tokens to strip. Cost cuts that ignore the quality and safety floor fail exam items.
Leave max_tokens unbounded forever
Unbounded completion budgets hide waste and inflate cost. Architects set task-appropriate caps after profiling.
Ignore the quality floor while cutting spend
Cheaper that fails the SLA is not an optimization. Re-measure golden and online quality after every cut.
Batch the interactive chat path
Batching helps non-interactive workloads. Interactive assistants need lean hot paths, not overnight queues.
Practice scenario
A hot-path assistant's weekly cost spiked while quality on the golden set stayed flat. The team proposes jumping to a cheaper smaller model immediately. What optimization sequence best aligns to cost-performance trade-offs?
Build exercise
Optimize a hot-path assistant for tokens and latency
40 minutes
What you'll learn
- Profile token and latency hotspots before spending
- Apply cache and trim without cutting safety
- Re-measure quality after each cut, then reconsider tier
- Batch only non-interactive workloads
Step 1
Profile where tokens and latency go on the hot path
Break down input vs output tokens, cache hit rate, tool-result size, and p95 stages (retrieve, model, tools). Identify the top waste sources before buying capacity.
Why: Spend without a profile optimizes the wrong layer. Exam items reward measure-first.
You should see: A hotspot table: stage → tokens or ms → % of request cost or latency.
Step 2
Cache stable prefixes and trim redundant context
Enable prompt caching for stable policy and tool schemas. Trim unused history, duplicate tool dumps, and low-value chunks. Keep safety and required quality content.
Why: Cache and trim beat premature model downgrades when quality is already flat.
You should see: Config diff showing cache breakpoints, trimmed context budget, and unchanged safety sections.
Step 3
Re-measure quality after each cut, then reconsider tier
Run the golden set and key online metrics after trim/cache. Only if cost still misses budget and quality stays above floor, evaluate a smaller/faster tier with the same fixtures.
Why: Quality SLA binds cost cuts. Tier changes without a re-check are silent regressions waiting to happen.
You should see: Before/after: $/request, p95, golden pass rate — with a go/no-go on tier change.
Step 4
Batch only non-interactive work; keep the hot path lean
Move summaries, backfills, and offline scoring to batched jobs. Keep the customer chat path synchronous and profiled. Document rejected "batch the UI" proposals.
Why: Interactivity constraints are architectural. Batching the wrong surface saves money and breaks UX.
You should see: Two workload classes in the design: interactive SLA vs batch cost targets.
Sources
- Prompt caching — docs.anthropic.com — stable prefix cost/latency lever
- Pricing — docs.anthropic.com — token economics for trade-offs
- Models overview — docs.anthropic.com — tier choice after trim/cache