2.4 · Lesson 4 of 5
Optimize context windows and manage token usage
What You Need to Know
Context windows are a budget, not a landfill. Architects decide what must stay verbatim (IDs, amounts, error codes, binding policy quotes) versus what can be summarized or dropped. Dumping entire corpora — logs, wikis, ticket histories — into every turn burns tokens, latency, and attention. Prefer retrieval and structured tool results over paste-all.
Position matters: critical constraints and live state can suffer when buried in the middle of huge contexts (lost-in-the-middle risk). Trim tool outputs with explicit retention rules so compression does not delete the one line that matters. Persist transactional facts in a durable structured block outside narrative summaries — progressive summarization alone often drops IDs.
Context engineering levers
- Per-layer token budgets with headroom for the answer
- Verbatim vs summarize vs drop decisions
- Live-state / facts blocks for critical IDs
- Positioning of must-follow constraints
- Structured retention when trimming tool results
Decision rules
- Budget tokens by layer — do not dump corpora.
- Keep transactional IDs outside narrative summaries.
- Trim tool results with structured retention rules.
- Watch position / lost-in-the-middle for critical facts.
- Prefer retrieval and excerpts over paste-all history.
Summaries vs facts
Summaries are for narrative continuity. Facts blocks are for machine-usable truth. When exams show IDs disappearing after summarization, the fix is persistence of structured fields — not “summarize harder” or larger windows alone.
Review checklist
- Is there an explicit context budget?
- Are verbatim fields listed and protected?
- Do tool adapters bound and structure output?
- Is live state separated from long history prose?
- Would a smaller excerpt still support the task?
Prefer answers that trim with retention and protect IDs. Distractors celebrate dumping more context or raising unrelated limits.
Exam application
Choose token budgeting, structured trim, and durable facts blocks. Reject paste-all corpora and summary-only designs that lose transactional truth.
Exam traps
Dumping entire corpora into context
Raw logs, wikis, or ticket histories waste budget and bury signal. Exam stems reward retrieval/trim over paste-all.
Ignoring lost-in-the-middle / position effects
Critical facts buried mid-context are easier to miss. Place must-use constraints where the design expects attention — and keep them short.
Narrative summarization that drops transactional IDs
Summaries lose case_id, account_id, and amounts. Persist critical facts outside the prose summary.
No retention rules when trimming tool results
Blind truncation can delete the one error line that matters. Structure what must stay verbatim vs what can compress.
Practice scenario
Each agent turn dumps 50k tokens of raw application logs into context. Critical case IDs also vanish after narrative summarization of older turns. What is the best immediate redesign?
Build exercise
Budget and trim context for a tool-using support agent
40 minutes
What you'll learn
- Allocate a per-layer context budget
- Separate verbatim facts from summarizable noise
- Place live state to reduce lost-in-the-middle risk
- Trim tool results with structured retention
Step 1
Set a context budget
Allocate tokens across system instructions, retrieved docs, tool results, and history for a support turn — with a reserve for the model response.
Why: Without a budget, every layer expands until quality and latency collapse.
You should see: A budget table with per-layer caps and owners.
Step 2
Decide verbatim vs summarize
List fields that must stay exact (IDs, amounts, error codes, policy quotes) versus content that can be compressed (long logs, chat fluff).
Why: Truth-critical tokens deserve retention; narrative noise does not.
You should see: A retention checklist marked verbatim / summarize / drop.
Step 3
Place critical constraints carefully
Keep must-follow rules and live IDs out of the buried middle of a giant paste. Prefer a short “live state” block near where the model acts.
Why: Position and lost-in-the-middle effects punish giant undifferentiated dumps.
You should see: A live-state block template separate from history prose.
Step 4
Trim tool results with structured retention
Implement adapters that return structured excerpts (signature, key fields, bounded tail) instead of full log blobs each turn.
Why: Tool-path trimming is the durable fix for dump-driven context bloat.
You should see: Before/after token counts for one tool call.
Sources
- Context windows — docs.anthropic.com — window behavior and budgeting
- Prompting best practices — docs.anthropic.com — clear, well-structured context
- Working with messages — docs.anthropic.com — message structure for history and tools