1.1 · 4.5% of the exam · Topic 1 of 3
Agent Architecture
Choose a workflow when the path is known, and an agent when the path is discovered at runtime. A supervisor delegates to isolated subagents. Sequential steps wait on each other. Parallel steps do not. The simplest structure that meets the requirement wins on reliability, predictability, cost, latency, and context.
Learning objectives
- State the difference between a workflow, which owns control flow in code, and an agent, which lets the model choose the next step.
- Apply the decision test: known and repeatable steps stay a workflow; a route that depends on discoveries becomes an agent.
- Describe a manager or supervisor hierarchy, including what the supervisor owns and what a subagent does not inherit.
- Choose sequential execution when a later step needs an earlier result, and parallel execution when the slices are independent.
- Explain the trade on reliability, predictability, cost, latency, and context, and reject an agent when a workflow already meets the requirement.
Detailed theory
Two ways to arrange the same parts
A Claude application is a model, some tools, and the code that decides what happens next. A workflow keeps that decision in your code. You name the stages, the branches, and the stop condition before the run starts. Claude fills a language-shaped gap inside a stage: extract these fields, draft this note, classify this ticket. The path itself does not depend on the model inventing the next stage.
An agent inverts the ownership. You give the model a goal and a tool list. The model chooses which tool to call, in which order, and when it is finished. Your code still runs the tools and still decides whether a call is allowed, but the route through the task is chosen at runtime. One model in a tool-use loop is already an agent. Agent does not mean a fleet of agents.
The agent loop
An agent loop is control flow, not a sentence in the prompt. Each turn sends the current messages, reads the response, and either stops or continues. The field that decides is stop_reason. tool_use means the application runs the requested tools, appends the results, and calls again. end_turn means the model is finished and the loop returns the content.
The application executes the tools. Claude does not. The tool result goes back as a user turn, which is why the history grows on every pass. A text block that says the work is done is not a stop signal. The same response can contain prose and a tool call. An iteration cap is a ceiling so a runaway loop cannot spend without bound. It is not how you detect completion.
while (turns < MAX_TURNS) {
const response = await client.messages.create({ model, tools, messages });
if (response.stop_reason === "end_turn") return textFrom(response);
if (response.stop_reason !== "tool_use") throw new Error(response.stop_reason);
messages.push({ role: "assistant", content: response.content });
messages.push({ role: "user", content: runTools(response) });
turns += 1;
}Decision criteria
Prefer the simplest structure that meets the requirement. Start at a single model call. Move to a workflow when you need several fixed stages. Move to an agent only when you cannot enumerate the route in advance and that non-determinism is acceptable.
A validation failure you already know about is not a reason to become an agent. "If the vendor is missing, reject the row" is a branch you can write. An agent earns its place when the follow-up itself is unknown: the lookup may reveal a fraud hold, a partial shipment, or a contract exception, and each of those opens a different investigation you cannot list at design time.
Two further signals keep a borderline task as a workflow. If a wrong step is expensive, you want a guard between named stages. If operators need to see the path in ordinary traces, a fixed pipeline is easier to follow than a transcript of model-chosen tool calls.
- Known, identical steps on every run: workflow. Cheaper, faster, and testable stage by stage.
- Path depends on discoveries: agent. You pay in extra turns, variable cost, and harder tests.
- Either structure would work: workflow. Autonomy that the task does not need is a defect.
- A single tool-use loop is an agent. Do not demand a multi-agent design to justify the word.
Supervisors, subagents, and decomposition
A manager or supervisor hierarchy is for a task that is too wide for one context, or that splits into jobs with different tools. The supervisor owns decomposition, delegation, aggregation, and the decision to stop or ask a human. Each subagent receives a bounded goal and the slice of context it needs. It does not inherit the supervisor's transcript, and it does not message its peers. Results return to the supervisor, which is the only place a cross-cutting decision is made.
Task decomposition is the supervisor's job. A thin report that misses whole categories is a decomposition failure: the workers did the slice they were given. Widen the split, or send the gap back out. Do not blame the specialist that was never assigned that slice.
Isolation is the point of the subagent, not an inconvenience. A research worker can read dozens of files. The supervisor should receive a short finding with sources, not the file contents. That is how a hierarchy stays inside a context budget. The cost is explicit handoff: anything the worker must know has to be in its prompt.
Sequential and parallel execution
Sequential execution is the right shape when step N consumes the output of step N−1. Invoice extraction, then schema validation, then a ledger write is sequential even when each stage calls Claude. Running them at the same time would write a row from an extraction that has not been checked.
Parallel execution is for independent slices that share no intermediate state. Classifying four unrelated tickets, or researching solar and geothermal as separate questions, can run together. Wall-clock time drops. Token cost does not: you still pay for every call. Fan out only the work that does not need the other results. The supervisor fans in, then does the synthesis that actually depends on all of them.
A hybrid is normal. Independent investigations run in parallel. The synthesis that combines them is sequential and waits. Putting the synthesis in the parallel phase produces a report that ignores a still-running worker.
What you are trading
Reliability is about whether the same input produces an acceptable result, and whether a failure is visible. A workflow fails at a named stage you can unit-test. An agent can take a path you never exercised. You buy reliability back with hooks, evals, and a bounded tool list, not by adding more agents.
Predictability is whether operators can say what will happen before the run. Workflows are predictable because the graph is yours. Agents are predictable only at the boundaries you coded: which tools exist, which calls a hook will block, and when the loop must stop.
Cost scales with model turns and with how much history each turn resends. An agent that explores pays for wrong branches. A workflow pays a fixed number of calls. Parallelism lowers latency and raises concurrent spend. Context pressure is the hidden cost: long tool logs make later turns worse and more expensive. Keep raw exploration in a subagent and return a summary.
Core concepts
Workflow
- What
- Code owns the sequence, the branches, and the stop condition. Claude is called inside a stage.
- Why
- A fixed path is cheaper, faster, and easier to test than a model choosing the path.
- When
- You can write the steps before the run, and a known branch is an if in code.
- When not
- The next action depends on facts the system has not seen yet and cannot list in advance.
Agent
- What
- The model chooses the next tool and when to stop. A single tool-use loop qualifies.
- Why
- Open-ended investigation cannot be encoded as a stable procedure.
- When
- The route is data-dependent and that variability is acceptable.
- When not
- The steps are the same on every run, or a wrong autonomous step is unacceptable and you have not put a code gate in front of it.
Supervisor
- What
- The agent that decomposes work, delegates, and aggregates. It is the only component that sees every result.
- Why
- Someone has to decide coverage, order, and the final answer. Peer-to-peer chatter duplicates work and splits accountability.
- When
- The job splits into specialists, or one transcript would be flooded by intermediate tool output.
- When not
- The task is one tool loop with one concern. A supervisor with a single worker is overhead.
Subagent
- What
- A worker with its own context, a narrow tool list, and a prompt that contains only the context it was given.
- Why
- Isolation keeps search logs and file bodies out of the parent window.
- When
- A slice is self-contained and its intermediate trace is not something the parent must quote.
- When not
- The worker needs the full parent history, or the parent must inspect every intermediate tool call. Pass a summary, or do the work in the parent.
Sequential execution
- What
- Stage N starts after stage N−1 returns, and it consumes that output.
- Why
- A later decision made on missing input is a wrong decision.
- When
- There is a real data dependency, including a check that must pass before a side effect.
- When not
- The slices do not read each other's output. Waiting then adds latency for no correctness gain.
Parallel execution
- What
- Independent slices run at the same time. A later fan-in combines them.
- Why
- Wall-clock time is the sum of the slowest slice, not the sum of all slices.
- When
- No slice needs another's result, and you can tolerate the concurrent token spend.
- When not
- One slice writes state that another reads, or the combination itself is the task and cannot start early.
Practical examples
Receipts into a ledger
Every receipt takes the same path: OCR, extract amount and vendor, validate against the chart of accounts, write the row. A validation failure is a known branch: reject or queue for a clerk. That is a workflow. An agent would spend turns re-deciding an order that never changes, and the cost per receipt would vary for no benefit.
The tell that flips the design is a document that might be a receipt, a contract, or a claim, where the follow-up lookup depends on which kind it is and cannot be listed up front. That path is an agent, still with a code gate in front of the ledger write.
A support case with an unknown cause
A customer writes "the charge looks wrong." The next lookup might be the order, the invoice, a promo, or a duplicate charge, and that choice depends on what the first lookup returns. The investigation is an agent. The refund itself is not. A hook or a prerequisite check blocks the refund tool until a verified customer and an amount inside policy exist. The investigation can be autonomous. The money movement cannot.
Six energy types, one report
A supervisor splits "renewable generation" into solar, wind, geothermal, hydro, biomass, and storage. The six research slices share no state, so they run in parallel. Each subagent returns findings with sources. The supervisor then writes the comparison, which is sequential because it needs every slice. A report that covers only solar and wind is a decomposition miss on the supervisor, even if those two subagents were thorough.
Claude-specific considerations
- On the Messages API the conversation is stateless. The loop resends history. That is why a subagent with its own message list protects the supervisor's context, and why you pass findings explicitly.
- stop_reason tool_use continues the loop. stop_reason end_turn ends it. pause_turn, max_tokens, refusal, and model_context_window_exceeded are not end_turn. A production loop must not treat them as success.
- A coordinator spawns workers with the Agent tool (Task on the exam's older wording). Workers do not see the parent transcript unless you put it in the prompt. Only the worker's final message comes back.
- Parallel workers are several Agent tool calls in one supervisor response. One call, then another on a later turn, is sequential and spends the latency you meant to save.
- Claude Managed Agents and the Agent SDK are construction choices, covered in 1.2. They do not change this decision. A known path is still a workflow if you host it yourself.
Architecture decisions
Tradeoffs
Choosing an agent buys flexibility. You pay for it on every axis below. A workflow is the left column. An agent is the right column. Parallelism is a separate dial: it cuts latency for independent work and does not cut token cost.
Quick reference
- Workflow: code chooses the next step. Agent: the model chooses the next step.
- Known path → workflow. Discovered path → agent. Tie → workflow.
- One tool-use loop is an agent. Multi-agent is optional.
- Supervisor decomposes, delegates, and aggregates. Subagents do not share memory.
- Pass context explicitly. Only the subagent's final message returns.
- Sequential when there is a data dependency. Parallel when there is not.
- stop_reason is the loop control. An iteration cap is a safety ceiling.
- Side effects stay behind code, even when the investigation is an agent.
Decision rules for the exam
Common exam traps
Exam tips
- Find the constraint in the first sentence: cost, testability, latency, or a side effect. The credited option changes the structure to match that constraint.
- If the stem says the steps never change, eliminate every option whose advantage is autonomy.
- If the stem says the parent context is full of search logs, the fix is isolation or compaction, not a larger window as the first move.
- "Must always" and "never" on a tool call point at code: a workflow stage or a hook. A stronger prompt stays probabilistic.
- A missing section in an aggregated report is a planning failure on the supervisor when the workers were scoped narrowly on purpose.
Common mistakes
Calling every Claude integration an agent.
A single call, or a fixed pipeline of calls, is not an agent. The model has to direct its own steps.
Building a supervisor for a one-tool job.
Add a hierarchy when context or tool scope demands it. Otherwise the extra hop adds latency and cost.
Letting subagents talk to each other.
Route every handoff through the supervisor so coverage and conflicts have one owner.
Running dependent stages in parallel to save time.
Check the data dependency first. Parallelism is only valid when no slice reads another slice.
Treating context overflow as a model problem.
The architecture put raw tool output in the parent. Move exploration to a subagent and return a summary.
Practice questions
Original questions for this topic. They are study items, not questions from the live exam.
A nightly job reformats a fixed CSV, validates it against one schema, and emails a summary. The stages never change. Which design fits?
A support agent looks up an account, and the next action depends on whether it finds a fraud hold, a billing dispute, or a shipping exception. Those cases are not a fixed list you can branch on today. What is the structure?
Which statement about a subagent is accurate?
Four product reviews must be tagged, and none of the tags depends on the others. A fifth step writes one summary from the four tags. How should this run?
Scenario questions
The ledger agent proposal
A finance team uploads vendor invoices. Every invoice is OCR'd, parsed into amount, vendor, and date, checked against the chart of accounts, and posted. Failed checks go to a clerk queue. A lead wants an autonomous agent "so the system can handle surprises." In six months of samples, the only surprise is a missing vendor id, which already routes to the queue.
What should the team build?
A report that skipped geothermal
A supervisor researches renewable generation. It assigns solar to one subagent and wind to another. Both return sourced, detailed notes. The published report has no geothermal, hydro, or storage sections. Someone suggests the search tool is weak and someone else suggests a larger context window.
Where is the defect?
Build exercise
Classify three jobs and draw the control flow
Intermediate · 40 minutes
What you will learn
- How to tell a known path from a discovered path.
- Where a supervisor belongs, and where it is extra.
- Which edges are sequential and which can run together.
- Where a side effect needs a code gate even if the rest is an agent.
Step 1
Write the three jobs on one page
Job A posts invoices with the stages in the ledger example. Job B investigates "this charge looks wrong" and may issue a refund under policy. Job C tags a batch of independent reviews and then writes one digest.
Why: The exam hands you mixed jobs. Sorting them before you pick tools keeps you from using one pattern everywhere.
You should see: Three short paragraphs, each ending with the words workflow or agent.
Step 2
Mark data dependencies
For each job, draw boxes for stages and arrows only where a box reads another's output. Circle any arrow that is a side effect: post, refund, send.
Why: An arrow is a reason to stay sequential. No arrow is permission to run in parallel. A circle is a reason for a gate.
You should see: A is a line of four boxes. B has a branchy lookup and a circled refund. C has four boxes into one.
Step 3
Place the supervisor only where isolation pays
Add a supervisor box only if some worker would otherwise dump raw tool output into the parent, or if two specialists have different tools. Write the exact fields that cross the boundary.
Why: Handoff fields are the exam's context-passing point. If a fact matters, it is in the prompt.
You should see: B's research worker returns claim, source, and amount. C needs no supervisor if one process can tag and then summarise.
Step 4
Price one run in turns
Count model calls for the workflow version and for an agent version of job A. Note which count is stable.
Why: Cost and latency arguments need a number, even a rough one. The fixed path has a fixed count.
You should see: Job A as a workflow has a small fixed count. The agent version has a range, with an iteration ceiling written beside it.
Review checklist
Checks are saved in this browser.
Key takeaways
- The simplest structure that meets the requirement is the right one. A known path is a workflow.
- An agent lets the model choose the next step. One tool-use loop is enough to qualify.
- Supervisors decompose and aggregate. Subagents are isolated and receive only the context you pass.
- Sequential means a data dependency. Parallel means there is none. Fan-in waits.
- Autonomy spends reliability, predictability, money, time, and context. Pay for it only when the path is discovered at runtime.
Sources
- Building effective agents — Anthropic engineering
- Agent SDK overview — Claude docs
- CCDV-F exam guide, Domain 1 skill weights — Public blueprint summary