CCDV-F · Study Guide

← Domain 1: Agents and Workflows

1.2 · 5.3% of the exam · Topic 2 of 3

Agent Construction with Claude

Once an agent is justified, pick who owns the loop. A custom harness gives you the mechanics. The Claude Agent SDK runs the loop, tool dispatch, and sessions in your process. Claude Managed Agents runs the harness for you, on an Anthropic sandbox or a self-hosted one. Hooks and human approval make the steps that must not fail into code.

Learning objectives

  • Place the Messages API, a custom harness, the Claude Agent SDK, and Managed Agents on one spectrum of control.
  • List what an agent definition declares: instructions, tools, and permissions.
  • Implement the custom loop contract: send, branch on stop_reason, run tools, append results, repeat.
  • Choose a deterministic hook or a human checkpoint for an irreversible action, instead of a stronger prompt.
  • Match Anthropic-hosted and self-hosted deployment to data-residency and operational constraints.

Detailed theory

Four ways to own the loop

Construction starts after the 1.1 decision. If the path is a workflow, you do not need an agent runtime. If it is an agent, you still choose how much of the harness you write.

The Messages API is the request. It does not store a session and it does not run tools. A custom harness is your loop on top of that API. The Claude Agent SDK is that loop, plus tool dispatch and session handling, running in your process while you still execute tools. Claude Managed Agents is the furthest step: Anthropic runs the harness, and the session executes in a sandbox you configure as Anthropic-hosted or self-hosted.

  • Messages API: one call. You hold history. You run tools.
  • Custom harness: you own stopping, retries, telemetry, and where hooks sit.
  • Agent SDK: you own instructions, tools, and permissions. The SDK owns the loop in your process.
  • Managed Agents: you own the agent definition and the environment choice. Anthropic owns the loop and the sandbox lifecycle.

What you declare

An agent definition answers three questions. The system prompt says what the agent is for and how it should behave. The tools say what it can call, each with a name, a description, and an input schema. Permissions, often an allowed-tool list, say which of those tools this run may actually use. The SDK and Managed Agents both need that definition. Neither one invents it for you.

A vague prompt or a blurry tool description still produces a bad agent on a perfect loop. Adopting the SDK to "fix tool choice" misses the split: the SDK automates mechanics, not judgment. Whether filesystem configuration such as CLAUDE.md and Skills loads under the SDK is a setting you set explicitly (settingSources), not something to assume.

The custom loop

A custom loop is the same cycle the SDK runs, written where you can see it. Send the messages. Inspect stop_reason. On tool_use, run every tool_use block, append the assistant message, append a user message of tool_result blocks, and continue. On end_turn, return. Any other stop_reason is a specific failure: max_tokens means the output was cut, refusal means the model declined, and context-window exhaustion means the history does not fit. None of those are end_turn.

You build this when the product needs control the SDK does not expose: a custom retry budget, a trace shaped like the rest of your platform, a stop rule that is not "the model ended," or a side effect that must pause for a person before the tool binary runs. The cost is that you now own every bug in the harness.

if (response.stop_reason === "tool_use") {
  messages.push({ role: "assistant", content: response.content });
  const results = await Promise.all(
    toolUses(response).map(async (block) => ({
      type: "tool_result" as const,
      tool_use_id: block.id,
      content: await gatedRun(block),
    })),
  );
  messages.push({ role: "user", content: results });
}

Tool execution inside the harness

The model requests a tool. Your process runs it. The harness is responsible for matching tool_use ids, timeouts, and what goes back into the transcript. A tool error should return as a tool_result the model can read, with enough structure to distinguish a transient timeout from a validation reject, so the next turn can repair or stop. Swallowing the error and continuing with a stale result is how the loop lies.

Least privilege is a construction choice. Do not register a refund tool on an agent whose job is summarising URLs. A capability the process never holds cannot be prompted into existence. When the tool must exist, a hook is the gate, not a sentence in the prompt.

Managed deployment

Claude Managed Agents is the option when you want the loop, tool runtime, and sandbox without operating them. A session is a running agent in an environment. Events are the user turns, tool results, and status updates your application exchanges with that session. The harness also applies built-in caching and compaction, which matters on long investigations.

The environment is the deployment choice the exam cares about. An Anthropic-hosted cloud sandbox minimises infrastructure: packages, network, and the execution host are Anthropic's. Session state and outputs are stored server-side, so this option is the wrong fit for a contract that requires Zero Data Retention or a HIPAA BAA. A self-hosted sandbox keeps execution on infrastructure you control for residency or compliance, while you still use the managed harness rather than writing the loop. Pick the sandbox from the data constraint, not from a preference for "cloud."

Hooks, deterministic actions, and people

A hook is code at a lifecycle boundary. PreToolUse runs before the tool. It can allow, deny, or rewrite the input. The side effect has not happened yet, so this is where a refund ceiling, a path allowlist, or a production delete ban belongs. PostToolUse runs after the tool and can normalise what the model will see. It cannot undo the delete. A prompt that says "never refund above $200" remains probabilistic.

Some decisions are not a rule. A human checkpoint pauses the loop and routes to a person: before an irreversible send, after a plan and before its side effects, or when a tool result is outside the range a retry would responsibly handle. The harness must actually wait. Asking the model to "check with the user" inside the prompt is not a checkpoint, because the model can skip the sentence.

Pair stop_reason with an explicit budget. A cap on turns, tokens, or money ends a loop that keeps requesting tools. The cap is the safety net. Completion is still end_turn.

Lifecycle and errors

A practical lifecycle is define, start, loop, gate, stop, and record. Define the prompt, tools, and permissions. Start a session or a message list. Loop on the model. Gate side effects with hooks or approvals. Stop on end_turn or on the budget. Record the trace so the next failure is diagnosable.

Handle errors at the layer that owns them. Transient tool failures retry in the harness with a limit. Validation failures go back to the model as a tool result so it can correct the arguments. Permission failures stop that tool and surface a reason; they are not a cue to try a more privileged tool. An unhandled stop_reason fails closed: return the error to the caller instead of appending empty success.

Core concepts

Claude Agent SDK

What
A library that runs the agent loop, dispatches tools, and handles session state on top of the Messages API, inside your process.
Why
Most agents need that machinery and do not need a novel harness.
When
You want supported loop behaviour and you still execute tools in your environment.
When not
You need a stop rule, a trace, or a retry policy the SDK does not expose. Build the loop, or use hooks if the gap is only enforcement.

Custom harness

What
Your loop: history, stop_reason, tool execution, hooks, budgets, and traces.
Why
You can see and test every transition.
When
Integration with existing telemetry, unusual stopping conditions, or a gate the managed runtime cannot express.
When not
You are about to reimplement ordinary tool dispatch. That is the SDK's job.

Managed Agents

What
Anthropic runs the harness. You choose an Anthropic-hosted sandbox or a self-hosted sandbox.
Why
Long-running sessions, scheduled runs, and a sandbox you do not want to build.
When
The product value is the task, not the loop, and the data-handling terms fit the sandbox you pick.
When not
You must keep every token and file inside a process you already operate, or you need ZDR/BAA terms the hosted service does not offer.

Hook

What
Deterministic code around a tool call. PreToolUse can block the call. PostToolUse can only shape the result.
Why
A prompt cannot guarantee an irreversible action.
When
The rule is "never" or "only when this precondition holds."
When not
The rule is a style preference. Put that in the prompt.

Human approval

What
The harness pauses and waits for a person before continuing.
Why
Some choices are policy judgments, not predicates you can code completely.
When
The worst case of an unattended step is unacceptable, and a fixed deny would block legitimate work.
When not
A code rule already decides the case. An approval dialog for every read is latency, not safety.

Practical examples

Wrong tool, SDK blamed

An agent calls search_customer when the user asked for an invoice. The team plans to "switch to the Agent SDK so it handles the logic." The SDK will not repair a tool description that says "looks up customer things" on two different tools. Keep the SDK if the loop is the pain. Rewrite the descriptions and the allowed-tool list if the pain is selection.

Shell access and a prompt

A coding agent can run a shell. The safeguard is a system line: "Never delete outside the workspace." Tests look clean. The design is still wrong, because a prompt is not a guarantee. Intercept the shell tool, reject any delete whose path is outside the workspace, and return that refusal as the tool result. The model can continue. The filesystem does not.

Residency decides the sandbox

A bank wants a long research agent and must keep execution in its own VPC. Managed Agents with a self-hosted sandbox matches that constraint. The Anthropic-hosted sandbox would move the session onto infrastructure the bank cannot accept. Writing a custom loop is only necessary if the bank also rejects the managed harness, not merely the hosted sandbox.

Claude-specific considerations

  • The Agent SDK sits on the Messages API. It is not a replacement transport and it does not remove the need for tool schemas.
  • Agent definitions carry the system prompt, the tools, and the permissions. The SDK does not infer a safe tool list from the prompt.
  • settingSources controls whether CLAUDE.md, skills, and similar filesystem configuration load. Set it. Do not assume the SDK mirrors an interactive Claude Code session.
  • Managed Agents stores session state for the hosted path and documents that this is outside Zero Data Retention and HIPAA BAA. Deleting a session is an API action you can take; it is not the same as the data never having been stored.
  • PreToolUse is the block. PostToolUse sees a tool that already ran. A refund ceiling implemented only as PostToolUse is too late.
  • Human approval in the harness is a pause. A sentence that tells the model to ask is still a prompt.

Architecture decisions

SituationChooseBecause
Ordinary tool loop, tools run in your service, no exotic stop rule.Claude Agent SDKYou keep the environment and skip reimplementing dispatch and sessions.
The trace, retry budget, and stop condition must match an existing platform.Custom harness on the Messages APIThose controls are the thing you would fight a framework to express.
Long sessions, and you do not want to operate a sandbox.Managed Agents, Anthropic-hostedThe harness and the cloud sandbox are the product you are buying.
The same managed harness, but execution must stay in your network.Managed Agents, self-hosted sandboxResidency is an environment setting, not a reason to rewrite the loop.
A refund above a limit must be impossible.PreToolUse deny, or do not register the toolThe guarantee has to be code on the call path.
A payment is sometimes legitimate and sometimes not, and policy is not a pure function.Human approval pauseA person holds the remaining judgment. The loop waits in the harness.

Tradeoffs

Moving right gives away control and operational load. Moving left gives both back, including the bugs. Hosting is a separate choice from who writes the loop.

AxisYou own moreAnthropic owns more
ControlCustom loop: every transition is yours.Managed Agents: you configure the agent and the environment.
Where tools runSDK and custom harness: in your process.Managed Agents: in the sandbox, cloud or self-hosted.
DataYour logs and your store.Hosted sessions are stored server-side and are outside ZDR and HIPAA BAA.
Failure ownershipYou debug dispatch, retries, and stop handling.You debug prompts, tools, and the session events you receive.
EnforcementHooks in your loop, before the side effect.Hooks and permissions still apply. A prompt is still not a guarantee.

Quick reference

  • SDK provides the loop, tool dispatch, and sessions. You provide prompt, tools, and permissions.
  • The SDK is a layer on the Messages API, not a different model.
  • Custom loop: send, read stop_reason, run tools, append results, repeat until end_turn.
  • Do not stop on prose. Do not treat max_tokens as success.
  • PreToolUse can block. PostToolUse cannot undo.
  • Human approval is a harness pause, not a prompt sentence.
  • Anthropic-hosted sandbox: least infrastructure. Self-hosted sandbox: residency.
  • Hosted Managed Agents is not the ZDR or HIPAA BAA path.
  • Iteration and spend caps are safety nets beside end_turn.

Decision rules for the exam

If the question says…The answer is likely…
"SDK will fix wrong tool choice"No. Fix descriptions and permissions.
"need custom retries and traces"Custom harness
"standard loop in our process"Agent SDK
"do not operate a sandbox"Managed Agents, Anthropic-hosted
"data must stay in our VPC"Self-hosted sandbox or your own process
"must never refund above N"PreToolUse or omit the tool
"stop when the text looks done"Wrong. Branch on stop_reason.
"PostToolUse to prevent the delete"Wrong. The delete already ran.

Common exam traps

TrapCorrect answer
The SDK replaces the Messages API.It calls the API and adds the loop.
A firm prompt guarantees a destructive action will not happen.Use a hook or remove the tool.
Self-hosted means you must write the loop.Self-hosted can be the sandbox under Managed Agents.
Hosted agents are fine for a ZDR contract.Hosted session state is stored server-side.
A tool error should be an empty success so the loop stays simple.Return the error as a tool result or fail the turn.
Ask the model to request approval.The harness pauses until a person responds.

Open the Domain 1 sheet

Exam tips

  • If the stem is about behaviour quality, the SDK is usually a distractor. Look at the prompt and the tool descriptions.
  • If the stem says never, always, or irreversible, the answer is a hook or a missing tool. Delete the options that only edit the prompt.
  • If the stem is about residency, answer with where the sandbox runs. If it is about loop mechanics, answer with who owns the harness. Those are different questions.
  • PostToolUse offered as a way to stop a payment is the trap. The payment has happened.
  • end_turn with no tool call means return the content. Do not invent another tool round.

Common mistakes

  • Rewriting the loop when the defect is a tool description.

    Change the description, the schema, or the allowed list. Keep the SDK.

  • Stopping when the assistant text sounds finished.

    Read stop_reason. Continue while it is tool_use.

  • Putting the refund ceiling only in the system prompt.

    Deny in PreToolUse, or do not give the agent the tool.

  • Choosing the Anthropic sandbox because the feature list is longer.

    Read the data contract first. ZDR and BAA push you off the hosted session store.

  • Using a human for every deterministic check.

    Code the rule you can state exactly. Save the person for the judgment you cannot.

Practice questions

Original questions for this topic. They are study items, not questions from the live exam.

What does an agent definition in the Claude Agent SDK declare?

Choose one answer

A custom loop must stop reliably and must never issue a refund above a configured limit. Which design is right?

Choose one answer

A team wants Managed Agents, and the contract requires execution inside their own network. Which configuration matches?

Choose one answer

The Agent SDK is already in use, and the agent still picks the wrong tool about half the time. What is the productive change?

Choose one answer

Scenario questions

Deletes that tests did not catch

A repository agent may run shell commands. The system prompt says "Never delete files outside this repository." A week of manual tests shows no violations. A reviewer still blocks the launch. The tool result path currently appends whatever the shell prints, with no inspection.

What change makes the boundary real?

Choose one answer

Choosing a runtime for a claims investigator

The claims team needs an agent that runs for many minutes, calls internal APIs, and writes a case file. They already run services in a private cluster. Counsel says session contents cannot be stored on a vendor's multi-tenant control plane. They do not want to maintain their own tool-dispatch loop if they can avoid it.

Which construction fits all three constraints?

Choose one answer

Build exercise

Build a refund loop with a real gate

Advanced · 50 minutes

What you will learn

  • How a custom loop treats stop_reason.
  • How a tool error returns to the model.
  • How a PreToolUse gate differs from a prompt line.
  • How a human pause is different from both.
  1. Step 1

    Write the loop skeleton

    Call the Messages API with two tools: get_customer and refund. Continue only on tool_use. Return only on end_turn. On any other stop_reason, throw with that value. Cap turns at 8.

    Why: The cap is visible, and completion is still the stop reason. Mixing them is the bug the exam describes.

    You should see: A short prompt that needs both tools ends with end_turn after two tool rounds. A prompt that stops early is not cut off by the cap.

  2. Step 2

    Return a structured tool failure

    When get_customer is unknown, return a tool_result that says the id was not found. Do not throw out of the loop.

    Why: The model can apologise or ask for a better id. A thrown exception hides the failure from the only component that can change course.

    You should see: The next assistant turn mentions the missing customer instead of calling refund.

  3. Step 3

    Block refunds over the limit in code

    Before the refund function runs, deny amounts over 200. Return the denial as the tool result. Leave a prompt line that says the same thing, and then try a user message that tells the agent to ignore the limit.

    Why: You want one run that proves the prompt can be talked past and one run that proves the hook cannot.

    You should see: The refund function is never entered for 201, including on the jailbreak-style follow-up.

  4. Step 4

    Add one human pause

    For amounts from 50 to 200, pause and require a boolean approve from the caller before the refund function runs. Amounts under 50 proceed. Amounts over 200 still deny with no prompt.

    Why: The middle band is judgment. The top band is a rule. Asking a human about the top band adds delay and still risks a yes.

    You should see: Three cases: 40 runs, 80 waits, 500 never calls the function.

Review checklist

Checks are saved in this browser.

Key takeaways

  • The SDK owns the loop, dispatch, and sessions. You own the prompt, the tools, and the permissions.
  • A custom harness is for control, not for reproducing the default loop.
  • Managed Agents can run in an Anthropic sandbox or a self-hosted one. The hosted path stores session state.
  • High-stakes actions are hooks or human pauses in the harness. Prompts are guidance.
  • Fail closed on an unexpected stop_reason, and feed real tool errors back into the transcript.

Sources

Domain 1 overview · Quick reference

View progress