CCDV-F · Study Guide

Domain 516.8%

Quick Reference: Domain 5 — Model Selection and Optimization

5.1 LLM Fundamentals

  • A response is next-token sampling until a stop_reason. It is not a lookup of a stored answer.
  • The context window holds input and output. max_tokens caps output only. Thinking tokens spend that cap.
  • Current windows: Opus 5.5 and Sonnet 5.5 are 1 million tokens. Haiku 4.5 is 200,000.
  • Temperature and related sampling controls change variety where the model accepts them. They do not make a contract.
  • Extended thinking is budget_tokens. Later models want adaptive thinking and effort.
  • Effort lives in output_config.effort: low, medium, high, xhigh, max. adaptive is not one of those values.
  • Opus 5.5 defaults to medium effort. Haiku 4.5 has no effort parameter.
  • Disabling thinking at xhigh or max on Opus 5 and later is a 400.
  • Fast mode is a premium speed setting on supported Opus models on the Claude API.
  • Zero-shot is instructions. Single-shot is one example. Multi-shot is several, including the edge.
If the question says...The answer is likely...
"the same prompt, two different totals"Sampling. Check the field
"stop inside a thinking block"max_tokens was shared with thinking
"adaptive as an effort value"adaptive is a thinking type
"Haiku, set effort to low"Haiku 4.5 does not support effort
"disable thinking at max effort on Opus 5"That pair is a 400. Lower the effort or leave thinking on
"the format is wrong and the instructions are already specific"Add examples
"need tokens faster on Opus"Fast mode, premium price, Claude API only
"1 million token window, so max_tokens can be anything"max_tokens is still the output cap
TrapCorrect answer
Temperature 0 means the answer is repeatable.Sampling can still diverge. A check is the control.
The context window and max_tokens are the same limit.The window fits the call. max_tokens caps the output.
Omitted thinking is free.The tokens are billed and they count toward max_tokens.
Fast mode is Haiku.Fast mode speeds a supported Opus model and costs a premium.
More rules beat one example when the shape is wrong.An example of the shape is the direct fix.

5.2 Technical Fundamentals

  • Messages is REST. One call is one HTTPS request with the full message array.
  • An official SDK wraps that request. It does not add memory or a different model.
  • stream true uses server-sent events. The client reads deltas and can assemble a Message.
  • Tool results and the next user turn are a new POST. They are not frames on the SSE response.
  • A long non-streaming call can die on a quiet connection. Stream it.
  • A WebSocket is two-way and persistent. It belongs between your client and your server.
  • The Messages API does not stream tokens over a WebSocket.
  • Keep the API key on the server. The browser talks to you.
  • An async client runs many REST calls. It does not merge them into one window.
  • Platform model ids differ. The protocol stays a message.
If the question says...The answer is likely...
"the SDK will remember the chat"You resend the messages array
"stream the tokens"Server-sent events on the Messages response
"open a WebSocket to the model"Use SSE for tokens. A socket, if any, ends at your server
"the HTTP call sits quiet for minutes"Stream, or raise the timeout
"put the API key in the page"The server holds the key and calls the SDK
"tool result mid-stream"Finish the turn, then POST the tool_result
"same JSON, different cloud id"Pin the id you evaluated on that platform
"async client means one shared context"Each call is its own request and bill
TrapCorrect answer
Streaming is a WebSocket to Anthropic.Streaming is server-sent events on HTTP.
The SDK stores the conversation on Anthropic's servers.You store it and send it again.
SSE lets the client push tool results upstream.The response is one-way. The next request carries the result.
A raw curl call uses a different model than the SDK.Both send the Messages body. The model field decides.
Streaming is cheaper.The final message is priced the same. Streaming changes delivery.

5.3 Model Selection and Tradeoffs

  • Haiku: fastest, cheapest, high volume, simpler tasks. Haiku 4.5 has a 200k window and no effort parameter.
  • Sonnet: speed and intelligence in the middle. Sonnet 5.5 defaults to high effort and a 1M window.
  • Opus: hard reasoning and long agentic work. Opus 5.5 defaults to medium effort. Adaptive thinking stays on.
  • Start from a measured set. The smallest passing model wins.
  • Effort is the lever inside a model. Fast mode is a premium speed on supported Opus models.
  • A new id can change tool use, format, and which thinking fields are legal.
  • Pin the snapshot you evaluated. Repeat the eval on a new id and on a different platform's id.
  • Output tokens cost more than input tokens on every family.
  • Do not promote a model from one flattering transcript.
If the question says...The answer is likely...
"high volume, already accurate on the small model"Haiku
"the policy answer fails on Sonnet"Opus, then remeasure
"Opus is correct and too slow"Lower effort if the set holds, or fast mode
"use the alias, it will track the latest"Pin the evaluated snapshot
"we tested the Claude API id and deployed the Bedrock id"Test the id you serve
"budget_tokens on a current Opus"That model expects adaptive thinking and effort
"one demo looked better"Compare the eval set
"put Opus on the classifier to be safe"Route the easy call down
TrapCorrect answer
The largest model is the safe default for every route.Use the smallest model that passes the set.
Fast mode is how you turn Opus into Haiku prices.Fast mode is a premium speed on Opus.
A new release honors the old request byte for byte.Thinking fields and default effort can change and reject or shift.
Sonnet, Opus, and Haiku share one context window.Haiku 4.5 is 200k. Sonnet 5.5 and Opus 5.5 are 1M.
Quality, latency, and cost can all be maxed.The family is a tradeoff. Effort only tunes inside it.

5.4 Cost and Token Management

  • Cost uses four buckets: uncached input, cache write, cache read, and output.
  • Output tokens, including thinking, use the output price. Do not blend them with input.
  • Current sticker prices: Haiku 4.5 $1/$5, Sonnet 5.5 $2/$10, Opus 5.5 $4/$20 per million input and output tokens.
  • Standard cache write is 1.25× input. Standard cache read is 0.1× input. Check the card for model-specific read rates.
  • The checkpoint is the end of the shared prefix. Stable tools and system text first. The changing user text after.
  • Automatic caching moves the breakpoint forward. Explicit breakpoints: up to four, on the blocks you name.
  • Under the minimum length, the call succeeds and the cache fields stay zero.
  • Five-minute TTL by default, refreshed on a hit. A one-hour TTL costs more to write.
  • Changing model, fast mode, thinking config, or top-level effort misses the old prefix.
  • The first call writes. Judge the hit on the second call's cache_read_input_tokens.
If the question says...The answer is likely...
"cache_read is zero on the first call"Expected. Look at the second call
"reads died after an effort edit"Top-level effort is part of the prefix
"the marker is set and creation is zero"The prefix is under the minimum length
"the user id is in the system prompt"That token changes the prefix. Move it after the checkpoint
"warm the cache with a dummy user message at the end"Put the explicit breakpoint before the dummy
"hourly job, five-minute TTL"The entry expires. Use the longer TTL or accept a rewrite
"fast mode for the demo, standard speed in prod"The speed change misses the other prefix
"bill equals tokens times one rate"Split input, output, write, and read
TrapCorrect answer
Caching makes output tokens cheaper.Reads discount the prefix. Output is still output.
A cache_control field guarantees a write.The prefix must clear the model's minimum length.
Automatic caching is right for a pre-warm that ends in a placeholder.The breakpoint would sit on the placeholder. Use an explicit one.
Batch half-price applies to a user who is waiting.Batch is for work that can wait. Interactive calls use the standard card.
Word count times a dollar figure is a cost model.Use the usage buckets and the price card.

Topics in this domain