1.3 · 4.9% of the exam · Topic 3 of 3
Agent Patterns and Frameworks
Patterns are how an agent uses tools, memory, and context: a tool loop, isolated subagents, and a window you prune on purpose. Frameworks such as Strands, LangGraph, and PydanticAI standardise loop control, state, and branching when that orchestration is actually complex. They do not make the model more capable.
Learning objectives
- Describe a tool-use loop as the basic agent pattern, including what the application appends between turns.
- Use subagents for context isolation, and state what crosses the boundary.
- Separate working context from durable memory, and name when memory does not fix a full window.
- Choose compaction, pruning, or isolation when a transcript fills with tool output.
- Match Strands, LangGraph, and PydanticAI to control, portability, and the pattern each one favours, and skip a framework when a plain loop is enough.
Detailed theory
The tool-use loop is the pattern underneath the others
Every richer pattern still reduces to the same cycle. The application sends messages and tool definitions. If stop_reason is tool_use, it runs the calls and appends tool results. If stop_reason is end_turn, it stops. Prompt chaining, routing, and an evaluator that checks a draft are workflows or agents wrapped around that cycle, depending on who chooses the next node.
Anthropic's workflow patterns stay relevant here even when you later wrap them in a framework. Chaining is fixed stages. Routing sends the input to a specialist. Parallelisation fans out independent calls. An orchestrator-worker pair is the supervisor from 1.1. An evaluator-optimizer loop drafts, checks, and revises. Use them when they match the dependency. A framework is how you express the graph, not a reason to add a graph.
Subagents and context isolation
A subagent has its own message list. It does not inherit the parent transcript. You pass the goal, the inputs, and any findings it must not rediscover. It returns a final message. The search queries, file bodies, and dead ends stay in the worker. That is context isolation: the parent's window is spent on the decision, not on the search.
Isolation fails in two familiar ways. The prompt is too thin, so the worker repeats work the parent already did, or it answers without the constraint that mattered. Or the parent pastes the worker's entire trace "so nothing is lost," which deletes the benefit. Pass structured findings. If the parent must audit a particular tool call, ask for that call in the return contract instead of merging histories.
Independent subagents run together when the parent emits several delegations in one turn. A worker that needs a sibling's output waits. Nested workers are a depth limit you should set, because a tree of investigators will fill a budget before the supervisor writes the answer.
Memory is not the context window
The Messages API does not remember a session for you. Whatever you resend is the working context. That window holds instructions, the dialogue, and tool results, and it bounds both the conversation and the response. Durable memory is different: a file, a database row, or an agent-memory note that survives the process and is loaded back in later on purpose.
Memory helps when a fact should outlive the chat: the customer's preferred language, a decision already made, a summary of yesterday's investigation. It does not enlarge the window. If you load every note on every turn, you have rebuilt the full transcript under another name. Retrieve the slice that this turn needs. Treat memory writes as a tool with a schema, not as a diary the model appends without a reason.
Managing the window
Tool output is usually what fills the window. A directory listing, a retrieved page, or a test log can dwarf the user message. Keep the raw artifact outside the transcript when you can: store it, and put a short, cited excerpt in the tool result. When the history itself is long, compact the older turns into a structured summary and drop the verbatim middle. Keep the system prompt, the latest user goal, and the unresolved facts.
Subagents are the isolation valve for exploration. Compaction is the valve for a single thread that must continue. Pruning is the valve for one oversized tool result. Using all three at once without a rule produces a summary of a summary and a worker that lacks the source. Pick the valve that matches the object that grew.
Drift looks like contradictions and forgotten constraints near the end of a long run. The remedy is the same hygiene. Raising temperature, repeating the system prompt after every turn, or cutting max_tokens does not remove the stale tool output.
When a framework earns its weight
An agentic framework provides a graph or a state object for multi-step, tool-using work. It standardises three things teams otherwise rewrite: loop control, how state moves between steps, and how the run branches. The model's abilities stay the model's. A framework that "makes the agent smarter" is a distractor.
Adopt one when the orchestration is the maintenance burden: many branches, checkpoints, human interrupts, retries with state, or a graph you need to draw in review. Skip it when the job is one tool in a loop until end_turn. The Agent SDK already runs that loop. A graph library around a single node is overhead, and it is the same mistake as reaching for an agent when a workflow would do.
Strands, LangGraph, and PydanticAI
Strands is a model-driven agent library. You register tools and let the model run the loop, with a small amount of orchestration code. It fits when you want that loop, across model providers, and you do not want to draw every edge. The trade is less explicit control of the graph. Portability is relatively high because the agent is not tied to one vendor's state machine.
LangGraph represents the run as a stateful graph: nodes, edges, and conditional edges, with checkpoints. It fits when you need durable steps, a human interrupt, or a replay of a failed branch. The trade is ceremony and a closer tie to the LangChain ecosystem. You take it for control and auditability, not because the model will score higher.
PydanticAI is a typed Python agent layer. Dependencies and results are Pydantic models, so a bad structure fails in your code and can be retried with the validation error. It fits when the pain is typed outputs and injected services, not a large graph. It is the wrong purchase if you wanted checkpointed branches and bought a validator.
There is no universal winner. Choose on control, portability, and whether the built-in pattern matches the task. You can outgrow a choice. Starting on the heavy framework "to be future-proof" pays the cost now for a graph you may never need.
Core concepts
Tool-use loop
- What
- Repeat model calls until end_turn, executing tools when stop_reason is tool_use.
- Why
- It is the smallest agent, and the piece every framework still runs.
- When
- The model must choose among a few tools and the path is short.
- When not
- You already know the sequence. Then the loop is a workflow with model calls inside stages.
Context isolation
- What
- A subagent's transcript is separate. Only the designated return value joins the parent.
- Why
- Intermediate tool output would otherwise evict the parent's instructions and goal.
- When
- Exploration is verbose and the parent needs a conclusion, not the search log.
- When not
- The parent must reason over every intermediate observation. Then do that work in the parent, or return those observations explicitly and accept the cost.
Working context
- What
- The tokens you send on this call: instructions, history, and tool results.
- Why
- The API is stateless. Unsent facts do not exist for that turn.
- When
- You are deciding what this call needs.
- When not
- You are deciding what the business should remember next month. That is durable memory, loaded later on purpose.
Durable memory
- What
- State stored outside the transcript and retrieved into a later turn.
- Why
- Some facts should survive compaction and new sessions.
- When
- A preference, a decision, or a case summary will be needed again and is small enough to load selectively.
- When not
- You are trying to fix a window that is already full of today's tool logs. Write a summary, but prune the logs. Do not also paste the summary plus the logs.
Agentic framework
- What
- A library that expresses multi-step agents as a graph or typed state. Strands, LangGraph, and PydanticAI are the named examples.
- Why
- Loop control, state passing, and branching stop being copy-pasted per project.
- When
- The orchestration is complex enough that the framework's pattern removes real code.
- When not
- A single tool loop, or when the goal is a smarter model. Frameworks do not change model capability.
Practical examples
One tool, and a proposal to adopt LangGraph
The agent calls a product API in a loop until it can answer a stock question. A lead wants LangGraph "so the agent becomes more capable and future-proof." Capability will not change. The graph would contain one node. Keep the tool loop, or the Agent SDK. Revisit a graph when there are real branches, a checkpoint, or a human interrupt worth drawing.
A research parent that contradicts itself
After forty minutes the assistant denies a constraint from the first user message. The trace is mostly full-page fetches. The fix is to move each fetch into a subagent that returns a cited paragraph, and to compact the parent's older turns. A larger window would hold a few more pages and then fail the same way, at a higher cost.
Picking the library for a claims graph
A claims process has a typed decision record, a human interrupt on amounts over a threshold, and the ability to resume after a crash. PydanticAI fits the decision record. LangGraph fits the interrupt and the checkpoint. Using both is reasonable: typed results inside graph nodes. Strands fits a simpler investigator that only needs tools and a loop. The exam wants the matching pattern, not the brand you like.
Claude-specific considerations
- Compaction in Managed Agents is part of the hosted harness. In a custom loop or the Agent SDK you still decide what to drop, unless you call a compaction step yourself.
- Subagent return values are the isolation boundary in Claude's agent tooling. Design that payload. Do not rely on the worker "remembering" the parent.
- Claude Code agent memory (user, project, local scopes) is durable memory for that product. It still has to be pulled into context to matter on a turn. It is not a second hidden window.
- Strands, LangGraph, and PydanticAI can all call Claude. None of them is an Anthropic runtime. The Agent SDK and Managed Agents are Anthropic's construction options from 1.2. A framework is an orchestration library on top of a model API.
- Evaluator-optimizer is a pattern: generate, check, revise. The checker can be a schema validator in code. If the check must always hold, the validator is code, and the model revises from the error.
Architecture decisions
Tradeoffs
Framework choice is control against portability against the pattern the library already likes. The model stays constant across the row.
Quick reference
- Tool loop: tool_use continues, end_turn stops, results are appended.
- Subagent context is isolated. Pass the slice. Accept the final message.
- Parallel subagents are multiple delegations in one parent turn.
- Working context is what you send. Durable memory is what you store and load later.
- A bigger window is not the first fix for tool-log bloat.
- Prune one huge result. Compact a long thread. Isolate an exploration.
- Frameworks standardise loop control, state, and branching.
- They do not raise model capability.
- LangGraph: explicit graph and checkpoints. PydanticAI: typed results. Strands: model-driven loop.
- Skip the framework when a single tool loop is the whole design.
Decision rules for the exam
Common exam traps
Exam tips
- If the stem claims a library increases intelligence, reject it. Look for loop control, state, or branching.
- Match the library to the pain: checkpoint, type, or a simple loop. Brand loyalty is not a rationale.
- Context questions want a structural cut: prune, compact, or isolate. Parameter tweaks are the distractors.
- A subagent that "shares context automatically" is wrong. Explicit prompt in, final message out.
- Multi-step does not mean multi-agent, and multi-agent does not mean a third-party framework.
Common mistakes
Storing every tool result verbatim for audit inside the prompt.
Store the artifact outside the window. Keep a citation and a short excerpt in the transcript.
Loading all durable memory on every turn.
Retrieve by the current task. Memory that always loads is just a longer prompt.
Adopting a graph for a one-node agent.
Stay on the tool loop until there is state and branching to draw.
Expecting the worker to know the parent's constraints.
Write the constraints into the worker prompt. Isolation includes ignorance.
Compacting away the open question.
The summary has to keep the goal, the constraints, and the unresolved facts. Drop the raw middle, not the task.
Practice questions
Original questions for this topic. They are study items, not questions from the live exam.
What do Strands, LangGraph, and PydanticAI provide for agent work?
An assistant's context is almost full of verbose tool output, and it has started ignoring earlier constraints. What is the right remedy?
Which memory design matches the Messages API?
A process must checkpoint after each case stage and pause for a person above a dollar threshold. The team also wants the final decision as a typed object. Which pairing fits?
Scenario questions
Future-proofing a stock checker
The production agent answers stock questions by calling one inventory tool until stop_reason is end_turn. Latency and cost are acceptable. An architect proposes adopting LangGraph before the next launch so the system is "ready for multi-step planning" and "more capable." There is no second tool and no human interrupt on the roadmap for this quarter.
What should the team do for this launch?
The parent that reread the web
A supervisor asks a subagent to compare two vendors. The subagent prompt is "research both and be thorough." The parent context already contains a 30-page fetch of vendor A from an earlier turn, but that text is not copied into the prompt. The subagent fetches vendor A again, then the parent pastes both full transcripts into the synthesis call. The synthesis call fails because the window is full.
Which fix addresses the failure?
Build exercise
Isolate a search and pick the orchestration
Intermediate · 45 minutes
What you will learn
- How a subagent return contract protects the parent window.
- How compaction differs from dumping memory back in.
- How to reject a framework when the graph has one node.
- How LangGraph-style and PydanticAI-style pains differ.
Step 1
Write a parent and one worker
The parent answers "which of these three docs mentions the refund cap?" The worker may read files. The worker's return schema is doc_name, quote, and line. The parent must not receive file bodies.
Why: The schema is the isolation boundary. If the body fits in the return, you have not isolated anything.
You should see: The parent transcript grows by three short findings, not by three documents.
Step 2
Compact the parent on purpose
After the findings, replace older tool chatter with a summary that keeps the user question and the three quotes. Call the model once for the answer.
Why: Compaction is a cut you can show. The answer should still cite a document name.
You should see: The synthesis request is visibly smaller than the sum of the file reads.
Step 3
Score three frameworks against a one-node graph
In a short note, say why Strands, LangGraph, and PydanticAI are each a mismatch or a match for this search. Then add one requirement that would make LangGraph the right extra: a human must approve before any quote is sent outside the team, and the run must resume after a crash.
Why: The exam pairs the library with a concrete pain. The approval-plus-resume line is that pain.
You should see: The first note rejects a heavy graph. The second names checkpoints and an interrupt.
Review checklist
Checks are saved in this browser.
Key takeaways
- The tool-use loop is the base pattern. Frameworks express it when the graph is worth expressing.
- Subagents isolate context. You pass the inputs and receive the final message.
- Memory stores facts. The window is only what you send. Loading everything defeats both.
- Prune a huge result, compact a long thread, isolate an exploration.
- Strands favours a model-driven loop, LangGraph an explicit checkpointed graph, PydanticAI typed results. None of them upgrades the model.
Sources
- Building effective agents — Anthropic engineering
- Context windows — Claude docs
- Agent SDK subagents — Claude docs