6.1 · 3.8% of the exam · Topic 1 of 3
Context Engineering
The context window is the only memory a call has. Drift and bloat happen when old tool dumps crowd out the instruction. Prune tool output, compact a long transcript into a summary, and isolate a subtask so its scratch work never enters the parent.
Learning objectives
- Treat the window as the sum of instructions, history, tool results, thinking, and the answer.
- Recognize context drift and bloat from the transcript, not from the model name.
- Prune tool results once the parent has the fact it needed.
- Compact older turns into a summary and keep the recent tail intact.
- Isolate a multi-step subtask in its own context and return only the result.
Detailed theory
What this skill covers
Domain 6 is Prompt and Context Engineering, 11.0% of the CCDV-F exam. This topic is context engineering, 3.8%: window management, drift and bloat, tool-output pruning, compaction, and isolation through subagents and multi-step workflows.
Prompt wording is the next topic. This one is what you allow into the window at all. A precise instruction loses to a transcript that no longer contains it.
What the window is holding
Every Messages call resends the context you choose: system text, tools, the message array, and room for the new output. Tool results and thinking blocks are part of that array. They are not free side notes.
The window is a budget. A 1 million token model still charges you to resend what you kept, and the recent tokens compete with the original task. Managing context means deciding which bytes are still evidence and which are a log.
Drift and bloat
Bloat is size. A search tool that returns ten full pages, a file read of a generated log, and the same policy pasted on every turn will fill the window with text the next decision does not need. The bill rises because you resend it. The answer gets worse because the task is a smaller fraction of the prefix.
Drift is the behavior that follows. Later turns start answering the latest tool dump, a side question, or a tone from an earlier draft. The original constraint is still in the system prompt and still loses, because the conversation has taught a different job. Drift is not fixed by raising the model. It is fixed by shortening what the next call sees.
Tool-output pruning
A tool result should be the fact the next step needs: an id, a total, an error, a short excerpt. The page, the log, and the intermediate rows can stay in your system. Once Claude has used them, keeping the full result in the parent transcript is bloat.
You can prune in your own history before the next POST: replace an old tool_result body with a one-line remainder. The API can also clear old tool results server-side when the conversation crosses a threshold you set. The clear-tool-uses strategy removes the oldest results in order and leaves a placeholder so the model knows the body was removed. By default the tool call itself stays. Clearing the call input is a separate switch.
Clearing edits the prefix, so a prompt cache that covered those results misses and must be written again. Clear enough tokens that the miss is worth it. A one-line trim on every turn rewrites the cache for almost no saving.
Compaction
Compaction replaces a long stretch of the transcript with a summary block. On the next request, content before that block is dropped. The conversation continues from the summary. You have to send the block back. Dropping it and sending only the new user message throws away the summary too.
Threshold compaction runs when input tokens cross a trigger. The documented default trigger is 150,000 input tokens, and the trigger value has to be at least 50,000. Custom summary instructions replace the default summary prompt. They do not append to it. Pause after compaction when you need to attach something, such as a recent turn, before the model continues.
Keep the tail if the last few turns must stay word for word. Summarize only the older messages, then place the compaction block in front of the turns you kept, including their thinking blocks. The API summarizes every message you include in that request, so a turn you still need verbatim cannot be inside the summarized span.
messages = [
compactionBlock,
...recentTurnsKeptVerbatim,
{ role: "user", content: nextQuestion },
];Isolation and multi-step workflows
A subagent starts with its own window. Its tool chatter stays there. The parent receives the result you asked the subagent to return. That is how a file-by-file review avoids pouring every file into the parent. The subagent does not already know the parent transcript. The task you send has to carry the facts it needs.
A multi-step workflow whose steps are known can be separate calls. Each call gets the artifact from the previous step, not the whole search log. Isolation is the same idea without a subagent: the parent stores the intermediate dump, and the next prompt sees the decision, the ids, and the error.
Do not isolate by deleting the constraint. A summary that drops the refund policy, the current tool error, or the user's latest correction is a new kind of drift. Compact the exploration. Keep the rule and the latest evidence.
Core concepts
Context bloat
- What
- Tokens in the window that the next decision does not need, usually old tool results and repeated text.
- Why
- They are resent on every call and they crowd the instruction.
- When
- A transcript grows by pages of tool output between decisions.
- When not
- The long text is the document the user asked you to use. That document is the task.
Context drift
- What
- Later turns follow the recent transcript instead of the original task.
- Why
- The window is what the model conditions on. Recency can outrank an early rule that is still present.
- When
- The answer matches the last tool dump and misses the system instruction.
- When not
- The instruction itself was ambiguous. That is a prompt problem.
Tool-output pruning
- What
- Replacing a used tool result with the short fact, or clearing old results past a threshold.
- Why
- The parent keeps the conclusion and loses the dump.
- When
- The tool already returned and the next step needs an id, a total, or an error string.
- When not
- The model has not read the result yet, or the full body is still the evidence.
Compaction block
- What
- A summary the API uses as the new start of the conversation. Content before it is ignored on later calls.
- Why
- A long agent loop cannot keep every search page.
- When
- Input tokens are near the trigger and the early turns are exploration.
- When not
- You still need those turns verbatim. Keep them after the block, or do not include them in the summary request.
Context isolation
- What
- A subagent or a separate call whose scratch work is not copied into the parent window.
- Why
- The parent stays on the decision. The child spends its own window on the search.
- When
- A subtask would dump logs, files, or many tool calls into the main transcript.
- When not
- The next step needs the parent's intermediate reasoning. Then pass that reasoning in the task.
Practical examples
The log that replaced the policy
A refund agent reads a policy in the system prompt, then a tool returns a 40,000-token event log. The next answer quotes a line from the log and ignores the policy cap. The policy is still in the request.
That is drift from bloat. Store the log in your system. Put a short finding in the tool result: the event id and whether it matches the policy. The next call can follow the cap again.
A summary that kept the tail
An investigation is at 160,000 input tokens. The last two turns contain the user's correction. A compaction request that includes those turns will summarize the correction instead of preserving it.
Summarize the older search. Place the compaction block, then the last two turns unchanged, then the new question. The correction stays literal. The search pages do not.
Claude-specific considerations
- The Messages API does not drop old tool results unless you prune them or enable a clear strategy.
- Server-side tool-result clearing removes the oldest results past a threshold and leaves a placeholder. Clearing tool inputs is optional.
- A clear rewrites the cached prefix. Clear enough that the cache miss earns its write.
- A compaction block is the new start. Content before it is ignored once you send the block back.
- The default compaction trigger is 150,000 input tokens. The trigger value must be at least 50,000.
- Custom compaction instructions replace the default summary prompt.
- A subagent's window is fresh. The parent sees the returned result, not the child's tool log.
Architecture decisions
Tradeoffs
Keeping everything makes the next turn informed and expensive, until the window is mostly logs. Pruning and compaction are lossy. Isolation spends another call so the parent stays small.
Quick reference
- The window holds instructions, tools, history, tool results, thinking, and the new output.
- Bloat is unused text you keep resending. Drift is the model following that text instead of the task.
- Prune a tool result to the fact once the next step no longer needs the dump.
- Server-side clearing drops the oldest tool results past a threshold and leaves a placeholder.
- Clearing changes the cached prefix. Clear a meaningful number of tokens.
- A compaction block replaces everything before it. Send the block back on the next call.
- Default compaction trigger: 150,000 input tokens. Minimum trigger: 50,000.
- Custom compaction instructions replace the default summary prompt.
- Keep a verbatim tail by leaving those turns out of the summary and placing them after the block.
- A subagent isolates the search. The parent gets the result you asked for, not the child's log.
Decision rules for the exam
Common exam traps
Exam tips
- If the stem says the instruction is still present and the answer follows a tool dump, prune the dump.
- If the stem says a summary forgot a recent correction, that correction was inside the summarized span.
- Subagents are for isolation. They are not a second copy of the same full transcript.
Common mistakes
Leaving every search page in the parent after the page has been used.
Store the page outside the transcript. Keep the id and the finding.
Compacting the turn that contains the user's correction.
Summarize older turns only. Place the correction after the block.
Spawning a subagent with no task and expecting the parent discussion to be there.
Write the goal, the constraints, and the facts into the child task.
Rewriting the system prompt to shout over a huge log.
Remove the log from the next request. The rule can stay as it was.
Practice questions
Original questions for this topic. They are study items, not questions from the live exam.
The system prompt still states a $50 refund cap. After a tool returns a long event log, Claude approves $400 and quotes the log. What should change first?
An agent loop crosses the compaction trigger. The summary is generated. The next user question is sent as the only message. Claude has no memory of the investigation. What was left out?
You want a summary of a long search, and you need the user's last correction to remain word for word. How do you compact?
A parent agent must review 40 files. Putting each file read into the parent window crowds out the review standard. What keeps the standard in the parent?
Scenario questions
The cache-aware clear
A coding agent caches a large system prompt and then accumulates tool results until the window is mostly file bodies. Someone proposes clearing a 20-token result every turn so the cache prefix barely changes. Quality is already drifting toward the latest file.
Which clear matches both the drift and the cache?
Build exercise
Cut a transcript down to the decision
Intermediate · 35 minutes
What you will learn
- Which lines are bloat.
- What a compaction tail must preserve.
- What a subagent should return.
Step 1
Write a 12-line transcript
Include a system rule, three tool dumps, a user correction, and a final question. Mark the lines the next answer still needs.
Why: The skill is the cut, not a longer prompt.
You should see: A short list of keepers and a longer list of dumps.
Step 2
Prune the tools
Rewrite each used tool result as one line: an id and a finding. Leave a result that has not been read yet.
Why: Pruning too early deletes evidence.
You should see: Short results for the finished tools, one full result for the unread one.
Step 3
Place a compaction block
Say which lines go into the summary and which two lines sit after the block.
Why: The tail is verbatim only if it is outside the summary request.
You should see: The correction and the final question after the block.
Step 4
Isolate one dump
Take the largest tool dump and make it a subagent task. Write the one paragraph the parent should receive.
Why: Isolation is a returned result, not a copied log.
You should see: A task sentence and a result sentence.
Review checklist
Checks are saved in this browser.
Key takeaways
- Drift is the model following leftover context. Remove the leftover text before you rewrite the instruction.
- Tool results are pruned to the finding. A large clear beats a one-token clear if you care about the cache.
- Compaction continues from the summary block. Turns that must stay literal sit after that block.
- A subagent spends its own window. The parent keeps the rule and the result.
Sources
- Compaction — Summary blocks, the input-token trigger, and continuing from the block.
- Context editing — Clearing old tool results and the cache cost of a clear.
- Build with Claude — The message array that is the context.
- CCDV-F blueprint notes — Domain 6 context skill and weight. Study notes, not exam items.