CCAR-P · Study Guide

← Domain 2: Claude Models, Prompting & Context Engineering

2.4 · Lesson 4 of 5

Optimize context windows and manage token usage

What You Need to Know

Context windows are a budget, not a landfill. Architects decide what must stay verbatim (IDs, amounts, error codes, binding policy quotes) versus what can be summarized or dropped. Dumping entire corpora — logs, wikis, ticket histories — into every turn burns tokens, latency, and attention. Prefer retrieval and structured tool results over paste-all.

Position matters: critical constraints and live state can suffer when buried in the middle of huge contexts (lost-in-the-middle risk). Trim tool outputs with explicit retention rules so compression does not delete the one line that matters. Persist transactional facts in a durable structured block outside narrative summaries — progressive summarization alone often drops IDs.

Context engineering levers

  • Per-layer token budgets with headroom for the answer
  • Verbatim vs summarize vs drop decisions
  • Live-state / facts blocks for critical IDs
  • Positioning of must-follow constraints
  • Structured retention when trimming tool results

Decision rules

  • Budget tokens by layer — do not dump corpora.
  • Keep transactional IDs outside narrative summaries.
  • Trim tool results with structured retention rules.
  • Watch position / lost-in-the-middle for critical facts.
  • Prefer retrieval and excerpts over paste-all history.

Summaries vs facts

Summaries are for narrative continuity. Facts blocks are for machine-usable truth. When exams show IDs disappearing after summarization, the fix is persistence of structured fields — not “summarize harder” or larger windows alone.

Review checklist

  • Is there an explicit context budget?
  • Are verbatim fields listed and protected?
  • Do tool adapters bound and structure output?
  • Is live state separated from long history prose?
  • Would a smaller excerpt still support the task?

Prefer answers that trim with retention and protect IDs. Distractors celebrate dumping more context or raising unrelated limits.

Exam application

Choose token budgeting, structured trim, and durable facts blocks. Reject paste-all corpora and summary-only designs that lose transactional truth.

Exam traps

  • Dumping entire corpora into context

    Raw logs, wikis, or ticket histories waste budget and bury signal. Exam stems reward retrieval/trim over paste-all.

  • Ignoring lost-in-the-middle / position effects

    Critical facts buried mid-context are easier to miss. Place must-use constraints where the design expects attention — and keep them short.

  • Narrative summarization that drops transactional IDs

    Summaries lose case_id, account_id, and amounts. Persist critical facts outside the prose summary.

  • No retention rules when trimming tool results

    Blind truncation can delete the one error line that matters. Structure what must stay verbatim vs what can compress.

Practice scenario

Each agent turn dumps 50k tokens of raw application logs into context. Critical case IDs also vanish after narrative summarization of older turns. What is the best immediate redesign?

Choose one answer

Build exercise

Budget and trim context for a tool-using support agent

40 minutes

What you'll learn

  • Allocate a per-layer context budget
  • Separate verbatim facts from summarizable noise
  • Place live state to reduce lost-in-the-middle risk
  • Trim tool results with structured retention
  1. Step 1

    Set a context budget

    Allocate tokens across system instructions, retrieved docs, tool results, and history for a support turn — with a reserve for the model response.

    Why: Without a budget, every layer expands until quality and latency collapse.

    You should see: A budget table with per-layer caps and owners.

  2. Step 2

    Decide verbatim vs summarize

    List fields that must stay exact (IDs, amounts, error codes, policy quotes) versus content that can be compressed (long logs, chat fluff).

    Why: Truth-critical tokens deserve retention; narrative noise does not.

    You should see: A retention checklist marked verbatim / summarize / drop.

  3. Step 3

    Place critical constraints carefully

    Keep must-follow rules and live IDs out of the buried middle of a giant paste. Prefer a short “live state” block near where the model acts.

    Why: Position and lost-in-the-middle effects punish giant undifferentiated dumps.

    You should see: A live-state block template separate from history prose.

  4. Step 4

    Trim tool results with structured retention

    Implement adapters that return structured excerpts (signature, key fields, bounded tail) instead of full log blobs each turn.

    Why: Tool-path trimming is the durable fix for dump-driven context bloat.

    You should see: Before/after token counts for one tool call.

Sources