Domain 5 · 15% of the exam
Context Management & Reliability
Manage context windows effectively, implement caching strategies, handle long conversations, and build reliable production systems.
Task statements
- 5.1
Context Window Management
Progressive summarisation drops numerical values, dates, percentages, and customer-stated expectations, so a $247.83 refund for order #8891 becomes a recent order. A persistent case facts block, included in every prompt and never summarised, holds those facts. Multi-issue sessions keep a separate structured issue layer. Key findings go at the start of aggregated inputs. Tool results are trimmed before they enter history. The Messages API is stateless, so each request carries the full conversation. Upstream agents return structured findings. Prompt caching exists; implementation details beyond knowing it exists are outside the current exam.
- 5.2
Escalation & Ambiguity Resolution
A support agent escalates for exactly three reasons: the customer explicitly asks for a human, the request is a policy gap or exception, or the agent has tried and cannot make progress. Honour an explicit human request on that turn. A policy violation has a documented refusal; a gap is silent and needs a human. Frustration and a self-reported confidence score are unreliable triggers. A frustrated customer with a straightforward issue gets the resolution; if they then insist on a human, escalate. Multiple customer matches require another identifier. The first fix is explicit escalation criteria with few-shot examples in the system prompt.
- 5.3
Error Propagation in Multi-Agent Systems
A failed subagent returns structured error context: failure type, the action it attempted, partial results, and alternative approaches. The four failure types are transient, validation, business, and permission. Silent suppression returns an empty success so the coordinator never recovers. Workflow termination throws away the subagents that finished. An access failure never executed and may be retried. A valid empty result executed and found nothing, so it is the answer. Synthesis output annotates coverage gaps. Transient failures are retried locally before they reach the coordinator.
- 5.4
Codebase Exploration & Context Degradation
Context degradation is an attention-quality problem: after extended codebase exploration the model cites typical patterns instead of the specific classes, methods, and paths it already found. Verbose discovery output buries earlier findings, and a larger context window does not fix that. Scratchpad files persist those findings outside the conversation and should be maintained from the start. Subagents isolate verbose exploration so the coordinator keeps structured summaries. Phase 1 summaries are injected into Phase 2 prompts to prevent a cold start. /compact is used proactively to protect context quality. A structured state manifest lets a crashed session resume without repeating the exploration.
- 5.5
Human Review & Confidence Calibration
A 97% aggregate accuracy figure can hide 45–60% accuracy on handwritten receipts, scanned PDFs, and international formats, because standard invoices dominate the volume. Validate accuracy by document type AND field segment before automating. Sample every stratum, including high-confidence extractions that are already automated. Raw confidence is relative: 0.90 on dates can mean 94% actual accuracy and 0.90 on amounts only 82%, so calibrate against labelled validation sets. Route fields above the calibrated threshold to automation with stratified sampling, and put the highest-uncertainty items first in a dynamic reviewer queue. Reduce human review only after that sequence, on segments that stay accurate.
- 5.6
Information Provenance & Multi-Source Synthesis
Every finding carries a structured mapping: claim, source URL, document name, relevant excerpt, and publication date. Attribution dies when a synthesis agent paraphrases those mappings away, so each step of the pipeline — research, analysis, synthesis, report — has to merge them forward. When two credible sources disagree, annotate both values with attribution and dates and let the reader decide. Different dates often describe a trend, such as growth moving from 8% to 12%. Render financial data as tables, news as prose, and technical findings as lists. An analysis that hits a conflict finishes with the conflict annotated for the coordinator.