Next-token generation
The response is a sequence of draws. Each token is sampled from the distribution over the prefix so far.
Exam context: A wrong sentence is a path of draws, which is why a second call can diverge. When: You are explaining a flaky completion or why a retry is not a fix. When not: You are debugging an HTTP error. No tokens were generated.
See also: 5.1 LLM Fundamentals
Context window
The maximum tokens for the request plus the response being generated.
Exam context: The window and max_tokens are different caps. One can have room while the other is exhausted. When: You are sizing a model or explaining a truncation. When not: You are setting the bill. Price is per token, not per window.
See also: 5.1 LLM Fundamentals
Sampling
The rule that turns next-token probabilities into one chosen token. Temperature, top_p, and top_k are sampling controls where the model accepts them.
Exam context: They change variety. They do not create a deterministic contract. When: The product needs more or less variation and the model allows the parameter. When not: You need the same answer every time. Use a check, not a lower temperature alone.
See also: 5.1 LLM Fundamentals
Adaptive thinking
thinking type adaptive. The model decides when and how deeply to think. Effort steers that depth.
Exam context: On current Opus and Sonnet models this replaces a fixed thinking budget. When: The task sometimes needs scratch work and sometimes does not. When not: The model only accepts extended thinking, such as Haiku 4.5, or you meant an effort level named adaptive. That name is not an effort.
See also: 5.1 LLM Fundamentals
Effort
output_config.effort, from low through max, controlling how much work the response spends.
Exam context: It is the first lever for cost and latency inside one model. When: A classification can be cheaper than a decision, on a model that supports the parameter. When not: Haiku 4.5, which does not support effort, or a request that disables thinking at xhigh or max on Opus 5 and later.
See also: 5.1 LLM Fundamentals
Multi-shot prompting
Several complete examples of the input and the output you want, placed before the live task.
Exam context: Examples fix a format or an edge that instructions keep missing. When: Zero-shot instructions are already clear and the misses are shape or boundary cases. When not: The request is invalid or the output is truncated. Examples will not close a tool loop.
See also: 5.1 LLM Fundamentals
Messages REST call
An HTTPS POST whose body is the model, the token cap, and the messages, and whose success is JSON or an event stream.
Exam context: Every SDK feature bottoms out in this request. When: You are tracing a call or writing a client the SDK does not cover. When not: You are choosing Opus versus Haiku. That choice is a field on the request, not a different protocol.
See also: 5.2 Technical Fundamentals
Official SDK
A typed client that builds the REST request, retries transient failures, and can accumulate a stream.
Exam context: It removes hand-rolled header and SSE bugs without changing the model. When: The language has a maintained Anthropic SDK. When not: You need a transport the Messages API does not speak. The SDK will not open a WebSocket to the model.
See also: 5.2 Technical Fundamentals
Server-sent events
A one-way stream of events on the HTTP response after stream is set true.
Exam context: This is how token deltas arrive, and how a long generation keeps the connection in use. When: A person is watching the answer, or the output cap implies a long run. When not: The client also needs to send tool results on that same response. Results go on the next POST.
See also: 5.2 Technical Fundamentals
WebSocket
A persistent two-way connection between two of your processes, or between a browser and your server.
Exam context: It is the wrong name for Messages streaming, and the right name for a live channel you operate. When: A client must push and receive without a new HTTP request each time, and your server is the other end. When not: You want Claude's tokens. Call the API and read SSE.
See also: 5.2 Technical Fundamentals
Client-side transcript
The message array your application stores and resends.
Exam context: The API and the SDK forget the call when the response ends. When: A conversation continues, or a second server resumes it. When not: You expected the SDK constructor to remember the previous user. It does not.
See also: 5.2 Technical Fundamentals
Haiku
The fastest family, for high volume and simpler tasks that already pass an eval.
Exam context: It is the low price and the low latency, with a smaller window on Haiku 4.5. When: The task is a label, a short extract, or a fan-out, and the set passes. When not: The set fails and the failure is reasoning depth. A lower price will not invent that depth.
See also: 5.3 Model Selection and Tradeoffs
Sonnet
The balance of speed and intelligence for product work between a label and the hardest reasoning.
Exam context: It is the middle of the price and latency ladder. When: Haiku misses the eval and Opus is more than the task needs. When not: You have not measured. The name is not a default for every route.
See also: 5.3 Model Selection and Tradeoffs
Opus
The family for hard reasoning and long agentic work.
Exam context: Quality on those tasks is why the higher price exists. When: The eval fails on smaller models, or the work is a long tool loop that needs the deeper model. When not: A cheap classification you attached out of habit. Route that call down.
See also: 5.3 Model Selection and Tradeoffs
Pinned model id
The exact id whose eval you are willing to serve.
Exam context: A floating alias or a platform's different id is a different behavior until you test it. When: Production, and any job that quotes a previous score. When not: A local sketch where you intend to throw the transcript away.
See also: 5.3 Model Selection and Tradeoffs
Release regression
A behavior change that arrives with a new id: prompts, tools, thinking settings, or defaults.
Exam context: The previous eval described the previous id. When: You move from one snapshot to the next, including a default effort that shifted. When not: You are comparing samples of the same pinned id. That spread is sampling.
See also: 5.3 Model Selection and Tradeoffs
Usage object
The token buckets on a response: uncached input, output, cache creation, and cache read.
Exam context: The invoice is these counts times the price card. When: You are explaining a bill or checking that a cache hit happened. When not: You are estimating a prompt you have not sent. Use the token counting endpoint, and still expect output to be unknown.
See also: 5.4 Cost and Token Management
Cache checkpoint
The cache_control breakpoint at the end of the prefix you want reused.
Exam context: Only the bytes up to that point can hit. A later edit of those bytes misses. When: Tools, system text, or history repeat across calls. When not: The shared text is below the model's minimum cache length. The marker will not write.
See also: 5.4 Cost and Token Management
Automatic caching
A top-level ephemeral cache_control. The breakpoint rides the last cacheable block and moves forward.
Exam context: A growing conversation can reuse the prefix without you moving a marker. When: The whole history so far is the shared prefix. When not: You are pre-warming a system prompt and the last block is a dummy user message. Use an explicit breakpoint before that dummy.
See also: 5.4 Cost and Token Management
Cache read
Tokens served from a matching prefix, billed at a fraction of the input price.
Exam context: This is the saving. A write on every call is not a saving. When: cache_read_input_tokens is non-zero on the second call. When not: The first call, which should show creation tokens instead.
See also: 5.4 Cost and Token Management
Price card
The input, output, cache write, cache read, batch, and fast-mode rates for one model id.
Exam context: Families differ, and fast mode does not use the standard card. When: You turn usage counts into money. When not: You compare quality. A cheaper card does not describe the eval.
See also: 5.4 Cost and Token Management