CCDV-F · Study Guide

← Domain 4: Eval, Testing, and Debugging

4.1 · 2.6% of the exam · Topic 1 of 1

Debugging and Error Handling

Name the failure before you change the prompt. An HTTP error never reached a finished answer. A stop_reason tells you whether a successful response is complete, truncated, or waiting on a tool. The trace shows which of those happened, and which layer owns the fix.

Learning objectives

  • Separate an HTTP error from a successful Messages response that still is not usable.
  • Match a recovery to the failure: retry, fix the request, continue the loop, or treat the text as model output.
  • Read a trace for stop_reason, request_id, token use, and tool_result blocks.
  • Decide whether the break is in the integration layer or in the model output.

Detailed theory

What this skill covers

Domain 4 is Eval, Testing, and Debugging, and it is 2.6% of the CCDV-F exam. The task in this lesson is debugging and error handling: identify the error type, choose a recovery, read the trace, and isolate an integration-layer problem from a model-output problem.

The exam is asking where you look first. A prompt rewrite does not fix a 401. A retry does not finish a tool call you never returned. A schema failure on a complete end_turn response is not an API outage.

Two different signals

An error means the Messages request was rejected or failed while the service handled it. The body is JSON with type error, an error object that has type and message, and a request_id. There is no assistant answer to parse.

stop_reason is on a successful response. It says why generation stopped. end_turn means Claude finished. max_tokens means your output cap was hit. stop_sequence means one of your stop sequences appeared. tool_use means you must run the client tools and send the results. pause_turn means a server-tool loop hit its iteration limit and you continue by sending that assistant message back. refusal means Claude declined, and stop_details says more. model_context_window_exceeded means the response filled the model window and the text is truncated.

Streaming can return HTTP 200 and then an error event later in the stream. That event is still an error. The partial tokens before it are not a finished answer.

Error identification

Read the HTTP status and error.type together. Catch the SDK's typed exception. Do not branch on the human message string. The type can gain new values over time. The request-id header, and request_id in the error body, is what you keep for support. On Claude Platform on AWS, keep the AWS request id for CloudTrail and the Anthropic request id for Anthropic.

A 400 invalid_request_error is the request shape: missing model, max_tokens, or messages, roles that do not alternate, a tool_use id with no tool_result immediately after, a tool_result that is not the first content in that user message, or a thinking block that was edited or dropped before you sent the assistant turn back. A bad model id is a 404 not_found_error. A malformed or revoked key is 401 authentication_error. A key that cannot use the resource is 403 permission_error. A billing problem is 402 billing_error. A body over the size limit is 413 request_too_large. A conflict with current state is 409 conflict_error.

429 rate_limit_error can be a rate limit, a usage-tier spend cap, a workspace spend limit, or a short acceleration limit. A tier spend-cap 429 has no retry-after header and keeps failing until access resumes. 500 api_error, 529 overloaded_error, and connection failures are service-side. 504 timeout_error means the request ran long. Long generations belong on the streaming Messages API.

  • 400, 401, 402, 403, 404, 413: fix the request, the key, billing, or the payload. A blind retry repeats the same failure.
  • 409: resolve the conflict, then retry.
  • 429 with retry-after: back off. 429 with no retry-after on a spend cap: wait until the cap lifts.
  • 500, 529, timeouts, and dropped connections: retry with backoff. The official SDKs already do that twice by default.

Recovery strategies

Recovery is the next action that matches the signal you already identified. Retry is for transient failures. The official SDKs retry connection errors, rate limits, and 5xx responses with exponential backoff, twice by default, and they honor retry-after when it is present. A second retry loop around the SDK retries the same call again. Turn the SDK retries up or down on the client when you need a different budget. Do not also string-match the message.

A tool_use stop is not a failure and not a retry. Execute every client tool, then append the assistant message exactly as you received it and a user message whose tool_result blocks come first, one result per tool_use id. If the tool itself failed, set is_error to true and put an actionable reason in the result. Claude can choose another step. An empty string with is_error false tells the model the tool succeeded and found nothing.

pause_turn is the other continuation. Send the assistant content back unchanged so the server-tool loop can proceed. Do not invent client tool_result blocks for that stop. A response that is waiting on your tools uses tool_use, not pause_turn.

max_tokens means the text you have is a prefix. Raise the cap or continue from that assistant message. Do not validate the prefix as if it were a finished object. model_context_window_exceeded is also truncation, and a larger max_tokens does not create window space the input already used. Shorten the history or the tool output, then call again. refusal is a decline: read stop_details and send the task to a fallback model when the product allows it. stop_sequence is the stop you configured. Read which sequence fired before you call it a bug.

if (response.stop_reason === "tool_use") {
  const results = await runClientTools(response.content);
  messages.push({ role: "assistant", content: response.content });
  messages.push({ role: "user", content: results });
} else if (response.stop_reason === "pause_turn") {
  messages.push({ role: "assistant", content: response.content });
} else if (response.stop_reason === "max_tokens") {
  // prefix only: raise the cap or continue, then parse
}

Trace analysis

A trace is one turn and the next request that used it. For a rejected call, store the status, error.type, the message, and request_id. For a 200, store the model, stop_reason, stop_sequence, output token count against max_tokens, the content block types, and each tool_use id paired with the tool_result you sent, including is_error. That row is enough to tell a truncation from a bad answer and a missing tool result from a model mistake.

Walk the row in order. No 200 means you never received model output. tool_use or pause_turn means the harness stopped early. output tokens equal to max_tokens, or stop_reason max_tokens, means the text is cut off. end_turn with output tokens under the cap means generation finished, and the remaining question is whether the content passes your check.

Failure modes you should be able to name from that row: the request was rejected, the stream died after 200, the output was truncated, the tool loop was not closed, the tool failed and the app hid it, the model finished and the content is wrong, or the same prompt varied because sampling is not a contract. Non-determinism is not an HTTP error. Pin the behavior with checks, not with a retry that you treat as success when any sample passes.

Integration layer versus model output

The integration layer is everything your process does around the model: the key, the URL, the JSON you assemble, the tool runner, the parser, and the retry policy. Model output is the assistant content on a request the API accepted and finished.

Use the trace as the split. Status 4xx or 5xx, a mid-stream error event, a missing or reordered tool_result, a rewritten thinking block, or a tool exception turned into an empty success are integration. end_turn, under the token cap, with the tool results honest, and a total or a field that fails the spec, is model output. Fix that with the instructions, the examples, a schema the response must match, or a retry that includes the validation error. Filing the request_id as an outage does not change the words.

The common mix-up is a parser error on truncated JSON. The parser failed in your process, and the cause is max_tokens. Repair the stop, then parse. The other mix-up is a confident wrong answer with a clean trace. That is not a 500, and a larger timeout will not correct it.

Core concepts

HTTP error

What
A non-success response, or a streaming error event, with error.type, a message, and request_id. There is no finished assistant answer.
Why
The status decides whether you retry, fix the request, or stop.
When
The call fails before you have a usable stop_reason.
When not
The API returned 200 and an assistant message. That path uses stop_reason.

stop_reason

What
The field on a successful Messages response that says why generation stopped.
Why
end_turn, max_tokens, tool_use, and pause_turn need different next steps.
When
You have a message object and you are about to parse it or loop.
When not
The body is an error object. stop_reason is not how HTTP errors are classified.

request_id

What
The request-id header and the request_id field on an error body. SDKs expose it on the response or the raw response.
Why
Support and CloudTrail correlate a failure to one call.
When
An error, a timeout, or a streaming error needs a ticket.
When not
A finished end_turn answer is wrong. The id does not explain a bad total.

tool_result is_error

What
A flag on the tool_result block you send back when the tool failed.
Why
Claude recovers from a stated failure. An empty success removes that choice.
When
The tool threw, timed out, or returned a domain error.
When not
The tool succeeded. A real empty result stays is_error false and says it is empty.

Trace

What
The ordered record of status or stop_reason, token use, content block types, and tool ids for one turn.
Why
The row separates a rejected call, an unfinished loop, a truncation, and a bad answer.
When
A production failure has to be assigned to a layer.
When not
You are designing the happy path and do not yet have a call to inspect.

Integration layer

What
Your client, request assembly, tool runner, parser, and retry policy.
Why
Those bugs are deterministic and are yours to fix. The model never saw a valid finished turn.
When
The trace shows a 4xx or 5xx, a dropped tool_result, a mutated thinking block, or a hidden tool error.
When not
The API completed the turn and the content fails a business check.

Model output

What
The assistant content after a successful, complete turn.
Why
Quality problems are handled with instructions, checks, and a retry that includes the validation error.
When
stop_reason is end_turn, the token cap was not hit, and the tools reported honestly.
When not
The text is a prefix from max_tokens or the window. That prefix is not the model's finished claim.

Practical examples

A parser error that is really truncation

An invoice extractor returns HTTP 200. stop_reason is max_tokens. output tokens equal the cap you set, 256. The text ends in the middle of a JSON string. The JSON parser throws, and the alert says the model produced invalid JSON.

The integration received a prefix and treated it as a document. The recovery is to raise max_tokens or continue the assistant message, then parse. The prompt does not need a new instruction about braces until a complete end_turn response still fails the parser.

A tool that failed in silence

The trace shows tool_use for get_invoice, then a user message with that tool_use_id, is_error false, and content "". The database client had thrown a timeout. The catch block replaced the exception with an empty string. Claude answers that the invoice is missing.

The model did what the result said. The integration layer must set is_error true and include the timeout. A second call can then retry the tool or tell the user the lookup failed.

A wrong total with a clean trace

HTTP 200, stop_reason end_turn, 180 output tokens against a cap of 1024, and no tools. The total does not match the line items in the user message. request_id is present because every response has one.

This is model output. Retrying the identical request may sample a different wrong total. The check belongs in your code, and a correction retry should include the failed check. Opening a support ticket with the request id does not change the arithmetic.

Claude-specific considerations

  • stop_reason is not an error type. It exists on a successful message.
  • Official SDKs retry connection errors, 429s, and 5xx responses twice by default and honor retry-after. Other 4xx responses need a different request, not another attempt.
  • A spend-cap 429 has no retry-after. Backoff does not clear it.
  • Client tools continue on tool_use with tool_result blocks first in the next user message. pause_turn continues by echoing the assistant message.
  • When thinking is on, the next request must include every thinking and redacted_thinking block unchanged, including an empty thinking field.
  • A streaming error event can follow HTTP 200. Partial text from that stream is unfinished.
  • Claude Platform on AWS has an AWS request id and an Anthropic request id. Use each with the system that indexed it.

Architecture decisions

SituationChooseBecause
The body is an error and status is 401 or 403.Fix the key or the permission. Do not retry.The model was not asked a question it could answer.
Status is 429 and retry-after is present.Back off for that delay. Let the SDK retry policy do it once.The limit is temporary. A tight loop spends the quota again.
Status is 429, there is no retry-after, and the workspace is at its spend cap.Stop calling until the cap allows traffic.The same request will fail for the rest of the window.
Status is 200 and stop_reason is tool_use.Run the tools and return every tool_result before any text.The turn is waiting on your process.
Status is 200 and stop_reason is max_tokens.Raise the cap or continue. Parse only after a later end_turn.The text is a prefix.
Status is 200, end_turn, and the field fails the spec.Reject the output and retry with the validation error, or change the instructions.The API already completed the turn.

Tradeoffs

A retry hides a transient outage and also hides a bug if you retry the wrong class. A stricter client fails closed and costs you an answer when the service would have succeeded on the second attempt. The trace is the way you pick.

AxisRetry or continueChange the request or the output
Transport500, 529, timeout, or a dropped connection.400, 401, 402, 403, 404, or 413.
Finished textend_turn under the cap, then a schema check.max_tokens or model_context_window_exceeded. The text is not finished.
Toolstool_use or pause_turn. The loop is still open.A tool exception stored as an empty success. The loop is lying.
A wrong answerOne more sample of the same prompt.A check, then a retry that includes the failure.

Quick reference

  • An error body has error.type and request_id. A message body has stop_reason. They are different events.
  • 400, 401, 402, 403, 404, and 413 are fixed by changing the request. They are not prompt bugs.
  • 429 with retry-after backs off. A spend-cap 429 has no retry-after and will not succeed until the cap lifts.
  • 500, 529, timeouts, and connection drops are retried with backoff. The SDKs already retry those twice by default.
  • tool_use: one tool_result per id, results first in the user message. A failed tool sets is_error true.
  • pause_turn: send the assistant content back unchanged. Do not invent client tool results.
  • max_tokens and model_context_window_exceeded are truncation. Parse after the turn is actually complete.
  • end_turn under the cap, with honest tool results, and a failed check: model output.
  • A thinking block must be echoed unchanged. Editing it is a 400 from the integration.
  • Log status or stop_reason, request_id, output tokens versus max_tokens, and tool ids with is_error.

Decision rules for the exam

If the question says…The answer is likely…
"the JSON parser threw"Read stop_reason before you blame the model
"thinking blocks cannot be modified"The app edited the assistant turn. Send it back unchanged
"tool_use ids were found without tool_result blocks"Return every result, and put those blocks before any text
"429 and no retry-after"A spend cap. Retry will not help
"HTTP 200 then the stream errors"Unfinished. Discard the partial text
"stop_reason pause_turn"Echo the assistant message to continue server tools
"stop_reason tool_use"You run the client tools. The model does not
"end_turn and the total is wrong"Model output. A validation retry beats a support ticket
"output tokens equal max_tokens"Truncation, even if the prefix looks like JSON
"504 on a long generation"Stream the Messages request

Common exam traps

TrapCorrect answer
Any failure is fixed by a stronger prompt.A 4xx or a missing tool_result is the integration.
stop_reason max_tokens is an HTTP 400.The call succeeded and the answer is cut off.
An empty tool_result means the tool failed.is_error true is the failure. Empty content can look like success.
pause_turn and tool_use continue the same way.Client tools need tool_result. pause_turn echoes the assistant message.
A wrong end_turn answer should be filed with request_id as an outage.request_id is for API failures. A finished wrong answer is output.
Raising max_tokens fixes model_context_window_exceeded.The window is full. Remove input before you generate more.

Open the Domain 4 sheet

Exam tips

  • The domain is small. The point is usually which signal you read, not a catalog of status codes.
  • If the stem includes a status code, stay in the integration. If it includes stop_reason, stay on that response.
  • If a tool result is mentioned, ask whether is_error was set and whether every tool_use id has a block.
  • Do not treat a truncated prefix as evidence the model cannot follow a schema.

Common mistakes

  • Wrapping the SDK in a second retry that also retries 400s.

    Keep retries on connection failures, 429 with retry-after, and 5xx. Change the request for the other 4xx codes.

  • Catching a tool exception and returning an empty tool_result.

    Set is_error true and describe the failure so the next turn can recover.

  • Parsing assistant text before reading stop_reason.

    max_tokens and model_context_window_exceeded mean the text is incomplete.

  • Sending a support ticket for a wrong total on end_turn.

    Use the request id when the API failed. Use a check and a validation retry when the turn completed.

Practice questions

Original questions for this topic. They are study items, not questions from the live exam.

A Messages call returns HTTP 200. stop_reason is max_tokens. The body text stops inside a JSON object and JSON.parse throws. What failed?

Choose one answer

Production starts returning 429 rate_limit_error with no retry-after header. The console shows the workspace monthly spend cap is exhausted. What is the recovery?

Choose one answer

The previous assistant message contains two tool_use blocks. The next call returns 400 and says tool_use ids were found without tool_result blocks immediately after. What does the user message need?

Choose one answer

HTTP 200, stop_reason end_turn, 140 output tokens, max_tokens 1024, no tools. The refund amount does not match the policy in the prompt. Which layer owns the fix?

Choose one answer

The database client throws. The tool handler catches it and returns a tool_result with is_error false and an empty string. Claude tells the user the record does not exist. Where is the bug?

Choose one answer

A response that only used server tools stops with pause_turn. How do you continue?

Choose one answer

Scenario questions

The 400 after a tool call

Extended thinking is on. The assistant turn has a thinking block, a short text block, and a tool_use block. The app keeps the text and the tool_use, drops the thinking block to save tokens, and sends a tool_result. The next call is 400 invalid_request_error and the message says thinking blocks cannot be modified.

What is the recovery?

Choose one answer

Two alerts, one outage

The on-call channel has two rows from the same hour. Row 1 is HTTP 529 overloaded_error with a request_id and no assistant message. Row 2 is HTTP 200, stop_reason end_turn, output tokens far under max_tokens, and a city name that is not in the source paragraph.

How should the two rows be handled?

Choose one answer

Build exercise

Split one failing turn into a layer

Intermediate · 35 minutes

What you will learn

  • How an error body differs from stop_reason.
  • Which failures are safe to retry.
  • How a trace records tools and truncation.
  • When the fix belongs in the integration and when it belongs in the output.
  1. Step 1

    Write four log lines

    Invent four one-line traces: a 401, a 429 with no retry-after, a 200 with stop_reason max_tokens and output tokens equal to the cap, and a 200 with end_turn whose extracted date is impossible.

    Why: Identification starts from the row, not from a guess about the prompt.

    You should see: Four lines, each with status and either error.type or stop_reason.

  2. Step 2

    Name the recovery for each line

    For each line, write one next action: change the key, stop until the cap lifts, continue or raise the cap, or reject the field and retry with the validation error.

    Why: The same word retry is wrong for three of the four lines.

    You should see: Four actions, and no prompt rewrite on the 401 or the spend cap.

  3. Step 3

    Close a tool loop on paper

    The assistant message has tool ids toolu_a and toolu_b. toolu_a succeeds with a total. toolu_b throws a timeout. Write the next user message shape: both results, results before text, is_error true only on toolu_b.

    Why: A missing or dishonest result is the integration bug that looks like a model bug.

    You should see: Two tool_result blocks and a timeout string on the failed id.

  4. Step 4

    Mark the layer

    Add a column to the four lines and the tool turn: integration or model output. One sentence each, citing the field that decided it.

    Why: The exam item is the split, not the vocabulary list.

    You should see: The impossible date is the only model-output row among the four logs. The tool timeout is integration until is_error is set.

  5. Step 5

    Say what you would not file

    Write the request you would not send to API support, and the one you would, using the lines above.

    Why: request_id is for a failed call. A finished wrong field is your check.

    You should see: The 429 or other rejected call is the ticket. The bad date is not.

Review checklist

Checks are saved in this browser.

Key takeaways

  • Identify the signal first: HTTP error.type or stop_reason. They are not interchangeable.
  • Retry transient overload and rate limits that say when to come back. Fix the request for the other client errors.
  • A tool loop is closed with honest tool_result blocks. pause_turn is closed by sending the assistant turn back.
  • Truncation is a harness problem. A complete end_turn that fails your check is model output.
  • The trace is the isolation tool: status, request_id, stop_reason, tokens against the cap, and is_error.

Sources

  • Claude API errors — HTTP status, error.type, request_id, and which failures the SDKs retry.
  • Stop reasons — end_turn, max_tokens, stop_sequence, tool_use, pause_turn, refusal, and model_context_window_exceeded.
  • Build with Claude — Messages, tools, and the surrounding application loop.
  • Troubleshooting tool use — Missing tool_result blocks and thinking blocks that must not be edited.
  • CCDV-F blueprint notes — Domain 4 scope and weight. Study notes, not exam items.

Domain 4 overview · Quick reference

View progress