4.1 Debugging and Error Handling
- An error body has error.type and request_id. A message body has stop_reason. They are different events.
- 400, 401, 402, 403, 404, and 413 are fixed by changing the request. They are not prompt bugs.
- 429 with retry-after backs off. A spend-cap 429 has no retry-after and will not succeed until the cap lifts.
- 500, 529, timeouts, and connection drops are retried with backoff. The SDKs already retry those twice by default.
- tool_use: one tool_result per id, results first in the user message. A failed tool sets is_error true.
- pause_turn: send the assistant content back unchanged. Do not invent client tool results.
- max_tokens and model_context_window_exceeded are truncation. Parse after the turn is actually complete.
- end_turn under the cap, with honest tool results, and a failed check: model output.
- A thinking block must be echoed unchanged. Editing it is a 400 from the integration.
- Log status or stop_reason, request_id, output tokens versus max_tokens, and tool ids with is_error.
| If the question says... | The answer is likely... |
|---|---|
| "the JSON parser threw" | Read stop_reason before you blame the model |
| "thinking blocks cannot be modified" | The app edited the assistant turn. Send it back unchanged |
| "tool_use ids were found without tool_result blocks" | Return every result, and put those blocks before any text |
| "429 and no retry-after" | A spend cap. Retry will not help |
| "HTTP 200 then the stream errors" | Unfinished. Discard the partial text |
| "stop_reason pause_turn" | Echo the assistant message to continue server tools |
| "stop_reason tool_use" | You run the client tools. The model does not |
| "end_turn and the total is wrong" | Model output. A validation retry beats a support ticket |
| "output tokens equal max_tokens" | Truncation, even if the prefix looks like JSON |
| "504 on a long generation" | Stream the Messages request |
| Trap | Correct answer |
|---|---|
| Any failure is fixed by a stronger prompt. | A 4xx or a missing tool_result is the integration. |
| stop_reason max_tokens is an HTTP 400. | The call succeeded and the answer is cut off. |
| An empty tool_result means the tool failed. | is_error true is the failure. Empty content can look like success. |
| pause_turn and tool_use continue the same way. | Client tools need tool_result. pause_turn echoes the assistant message. |
| A wrong end_turn answer should be filed with request_id as an outage. | request_id is for API failures. A finished wrong answer is output. |
| Raising max_tokens fixes model_context_window_exceeded. | The window is full. Remove input before you generate more. |