5.1 LLM Fundamentals
- A response is next-token sampling until a stop_reason. It is not a lookup of a stored answer.
- The context window holds input and output. max_tokens caps output only. Thinking tokens spend that cap.
- Current windows: Opus 5.5 and Sonnet 5.5 are 1 million tokens. Haiku 4.5 is 200,000.
- Temperature and related sampling controls change variety where the model accepts them. They do not make a contract.
- Extended thinking is budget_tokens. Later models want adaptive thinking and effort.
- Effort lives in output_config.effort: low, medium, high, xhigh, max. adaptive is not one of those values.
- Opus 5.5 defaults to medium effort. Haiku 4.5 has no effort parameter.
- Disabling thinking at xhigh or max on Opus 5 and later is a 400.
- Fast mode is a premium speed setting on supported Opus models on the Claude API.
- Zero-shot is instructions. Single-shot is one example. Multi-shot is several, including the edge.
| If the question says... | The answer is likely... |
|---|---|
| "the same prompt, two different totals" | Sampling. Check the field |
| "stop inside a thinking block" | max_tokens was shared with thinking |
| "adaptive as an effort value" | adaptive is a thinking type |
| "Haiku, set effort to low" | Haiku 4.5 does not support effort |
| "disable thinking at max effort on Opus 5" | That pair is a 400. Lower the effort or leave thinking on |
| "the format is wrong and the instructions are already specific" | Add examples |
| "need tokens faster on Opus" | Fast mode, premium price, Claude API only |
| "1 million token window, so max_tokens can be anything" | max_tokens is still the output cap |
| Trap | Correct answer |
|---|---|
| Temperature 0 means the answer is repeatable. | Sampling can still diverge. A check is the control. |
| The context window and max_tokens are the same limit. | The window fits the call. max_tokens caps the output. |
| Omitted thinking is free. | The tokens are billed and they count toward max_tokens. |
| Fast mode is Haiku. | Fast mode speeds a supported Opus model and costs a premium. |
| More rules beat one example when the shape is wrong. | An example of the shape is the direct fix. |
5.2 Technical Fundamentals
- Messages is REST. One call is one HTTPS request with the full message array.
- An official SDK wraps that request. It does not add memory or a different model.
- stream true uses server-sent events. The client reads deltas and can assemble a Message.
- Tool results and the next user turn are a new POST. They are not frames on the SSE response.
- A long non-streaming call can die on a quiet connection. Stream it.
- A WebSocket is two-way and persistent. It belongs between your client and your server.
- The Messages API does not stream tokens over a WebSocket.
- Keep the API key on the server. The browser talks to you.
- An async client runs many REST calls. It does not merge them into one window.
- Platform model ids differ. The protocol stays a message.
| If the question says... | The answer is likely... |
|---|---|
| "the SDK will remember the chat" | You resend the messages array |
| "stream the tokens" | Server-sent events on the Messages response |
| "open a WebSocket to the model" | Use SSE for tokens. A socket, if any, ends at your server |
| "the HTTP call sits quiet for minutes" | Stream, or raise the timeout |
| "put the API key in the page" | The server holds the key and calls the SDK |
| "tool result mid-stream" | Finish the turn, then POST the tool_result |
| "same JSON, different cloud id" | Pin the id you evaluated on that platform |
| "async client means one shared context" | Each call is its own request and bill |
| Trap | Correct answer |
|---|---|
| Streaming is a WebSocket to Anthropic. | Streaming is server-sent events on HTTP. |
| The SDK stores the conversation on Anthropic's servers. | You store it and send it again. |
| SSE lets the client push tool results upstream. | The response is one-way. The next request carries the result. |
| A raw curl call uses a different model than the SDK. | Both send the Messages body. The model field decides. |
| Streaming is cheaper. | The final message is priced the same. Streaming changes delivery. |
5.3 Model Selection and Tradeoffs
- Haiku: fastest, cheapest, high volume, simpler tasks. Haiku 4.5 has a 200k window and no effort parameter.
- Sonnet: speed and intelligence in the middle. Sonnet 5.5 defaults to high effort and a 1M window.
- Opus: hard reasoning and long agentic work. Opus 5.5 defaults to medium effort. Adaptive thinking stays on.
- Start from a measured set. The smallest passing model wins.
- Effort is the lever inside a model. Fast mode is a premium speed on supported Opus models.
- A new id can change tool use, format, and which thinking fields are legal.
- Pin the snapshot you evaluated. Repeat the eval on a new id and on a different platform's id.
- Output tokens cost more than input tokens on every family.
- Do not promote a model from one flattering transcript.
| If the question says... | The answer is likely... |
|---|---|
| "high volume, already accurate on the small model" | Haiku |
| "the policy answer fails on Sonnet" | Opus, then remeasure |
| "Opus is correct and too slow" | Lower effort if the set holds, or fast mode |
| "use the alias, it will track the latest" | Pin the evaluated snapshot |
| "we tested the Claude API id and deployed the Bedrock id" | Test the id you serve |
| "budget_tokens on a current Opus" | That model expects adaptive thinking and effort |
| "one demo looked better" | Compare the eval set |
| "put Opus on the classifier to be safe" | Route the easy call down |
| Trap | Correct answer |
|---|---|
| The largest model is the safe default for every route. | Use the smallest model that passes the set. |
| Fast mode is how you turn Opus into Haiku prices. | Fast mode is a premium speed on Opus. |
| A new release honors the old request byte for byte. | Thinking fields and default effort can change and reject or shift. |
| Sonnet, Opus, and Haiku share one context window. | Haiku 4.5 is 200k. Sonnet 5.5 and Opus 5.5 are 1M. |
| Quality, latency, and cost can all be maxed. | The family is a tradeoff. Effort only tunes inside it. |
5.4 Cost and Token Management
- Cost uses four buckets: uncached input, cache write, cache read, and output.
- Output tokens, including thinking, use the output price. Do not blend them with input.
- Current sticker prices: Haiku 4.5 $1/$5, Sonnet 5.5 $2/$10, Opus 5.5 $4/$20 per million input and output tokens.
- Standard cache write is 1.25× input. Standard cache read is 0.1× input. Check the card for model-specific read rates.
- The checkpoint is the end of the shared prefix. Stable tools and system text first. The changing user text after.
- Automatic caching moves the breakpoint forward. Explicit breakpoints: up to four, on the blocks you name.
- Under the minimum length, the call succeeds and the cache fields stay zero.
- Five-minute TTL by default, refreshed on a hit. A one-hour TTL costs more to write.
- Changing model, fast mode, thinking config, or top-level effort misses the old prefix.
- The first call writes. Judge the hit on the second call's cache_read_input_tokens.
| If the question says... | The answer is likely... |
|---|---|
| "cache_read is zero on the first call" | Expected. Look at the second call |
| "reads died after an effort edit" | Top-level effort is part of the prefix |
| "the marker is set and creation is zero" | The prefix is under the minimum length |
| "the user id is in the system prompt" | That token changes the prefix. Move it after the checkpoint |
| "warm the cache with a dummy user message at the end" | Put the explicit breakpoint before the dummy |
| "hourly job, five-minute TTL" | The entry expires. Use the longer TTL or accept a rewrite |
| "fast mode for the demo, standard speed in prod" | The speed change misses the other prefix |
| "bill equals tokens times one rate" | Split input, output, write, and read |
| Trap | Correct answer |
|---|---|
| Caching makes output tokens cheaper. | Reads discount the prefix. Output is still output. |
| A cache_control field guarantees a write. | The prefix must clear the model's minimum length. |
| Automatic caching is right for a pre-warm that ends in a placeholder. | The breakpoint would sit on the placeholder. Use an explicit one. |
| Batch half-price applies to a user who is waiting. | Batch is for work that can wait. Interactive calls use the standard card. |
| Word count times a dollar figure is a cost model. | Use the usage buckets and the price card. |