CCDV-F · Study Guide

← Domain 5: Model Selection and Optimization

5.1 · 5.2% of the exam · Topic 1 of 4

LLM Fundamentals

Claude generates one next token at a time inside a finite context window. Sampling makes that choice vary. Thinking, effort, fast mode, and examples change how many tokens that process spends and what shape the text takes.

Learning objectives

  • Explain next-token generation, tokens, and the context window as separate limits.
  • Treat sampling as the source of non-determinism, and a single completion as a sample.
  • Choose extended thinking, adaptive thinking, or an effort level for the work in front of you.
  • Place zero-shot, single-shot, and multi-shot prompting against the failure you are seeing.
  • Describe fast mode as a speed and price setting on models that offer it.

Detailed theory

What this skill covers

Domain 5 is Model Selection and Optimization, 16.8% of the CCDV-F exam. This topic is LLM fundamentals, 5.2%: tokens, context windows, sampling, non-determinism, next-token generation, fast mode, extended and adaptive thinking, effort, and zero-shot through multi-shot prompting.

The later topics pick a model, a client, and a bill. This topic is the unit they all share. A token is the thing being generated, cached, and priced. The window is how much of the conversation fits. The sample is one draw from that process.

Next-token generation

A Claude response is built one token at a time. Given the prefix so far, the model assigns a probability to each possible next token and draws one. That token is appended, and the draw repeats until the turn stops: a natural end, your max_tokens cap, a stop sequence, or another stop_reason.

There is no stored answer for the whole prompt that the API looks up. The words you read are the path of those draws. A fact that was likely at the first token can become unlikely after an earlier draw goes a different way. That is why a second call can diverge even when the messages you sent are identical.

Tokens and the context window

A token is the model's piece of text, smaller than a word on average and not a stable fraction of characters. Estimates on the current tokenizer put about 555,000 words in 1 million tokens. Models before the Opus 4.7 tokenizer fit more like 750,000 words in that million. A 200,000-token window is on the order of 150,000 words. Count the string you will send. A word count is a sketch.

The context window is the space for the request and the response being generated. Current Opus 5.5 and Sonnet 5.5 windows are 1 million tokens. Haiku 4.5 is 200,000. max_tokens is a different knob: the most output tokens you will allow on this call. Input plus that output has to fit the window. Filling max_tokens stops the turn with stop_reason max_tokens even when the window still has room.

Thinking tokens are output tokens. They count toward max_tokens and toward the bill, including when the thinking text is omitted and the block comes back with an empty thinking field. A tight max_tokens can end the turn inside the reasoning and leave you no answer text.

Sampling and non-determinism

Sampling parameters such as temperature, top_p, and top_k change which tokens are eligible and how sharply the draw prefers the likely ones. A low temperature concentrates the draw. It does not turn the model into a database. Two calls can still differ. Once they differ at one token, every later token is conditioned on a different prefix.

Whether a given model accepts a non-default temperature is part of that model's API. A rejection is a 400 on the request. The quality of a finished answer is a separate question, answered with an evaluation set rather than with another unchecked sample.

Thinking and effort

Thinking is scratch work before the answer, returned as thinking blocks ahead of the text. The text in the block is a summary, not a raw trace. You pass the blocks back unchanged on later turns, the same way you pass tool_use blocks back. Display summarized shows the summary. Display omitted, the default on many current models, leaves the thinking field empty. The tokens are billed either way.

Extended thinking is the older mode: thinking type enabled plus a budget_tokens cap on the scratch work. It is deprecated on Opus 4.6 and Sonnet 4.6, and later models reject it. Adaptive thinking sets type to adaptive and lets the model decide when and how deeply to think. On Opus 5.5 and several other current models, adaptive thinking is already on. adaptive is a thinking mode. It is not an effort level.

Effort is output_config.effort. The ladder is low, medium, high, xhigh, and max. It steers how much work the response spends, and in adaptive mode that includes how often and how deeply the model thinks. Defaults differ: Opus 5.5 defaults to medium, and several other current models default to high. Haiku 4.5 does not support effort. Not every model accepts every rung. On Opus 5 and later, disabling thinking at xhigh or max returns a 400. Keep thinking on, or drop the effort to high or below.

Use a lower effort on a short classification or a narrow subagent. Use a higher effort when the decision is the product and your evals improve. Changing the top-level effort rewrites the cached prefix, so a conversation that is caching should keep one effort unless the model supports a per-message effort change that leaves the prefix in place.

{
  "model": "claude-opus-5-5",
  "max_tokens": 16000,
  "thinking": { "type": "adaptive" },
  "output_config": { "effort": "medium" },
  "messages": [{ "role": "user", "content": "Check this refund against the policy." }]
}

Fast mode

Fast mode is a research preview that raises output speed, up to about 2.5 times, on supported Opus models on the Claude API. You set speed to fast and send the fast-mode beta. The tokens are the same kind of tokens. The price is a premium schedule, not the model's standard input and output rates. It is not a smaller model and it is not a lower effort.

It is not offered on Amazon Bedrock, Google Cloud, Microsoft Foundry, or Claude Platform on AWS for the Opus models that document the preview. Switching between standard and fast speed starts a new cache prefix.

Zero-shot, single-shot, and multi-shot

Zero-shot is the task and the criteria, with no worked example. Use it when the instruction already names the format and the boundary. A second paragraph of rules is still zero-shot.

Single-shot adds one complete example of input and the output you want. Use it when there is one canonical shape and the model is missing that shape. Multi-shot adds several examples, including the edge you keep seeing in reviews. Examples teach a format and a boundary case more directly than another rule about that case.

Each example sits in the context and on the bill. Put stable examples before the variable user turn so a cache prefix can cover them. An example that contradicts the instruction trains the contradiction. An example is still a sample of the pattern, not a validator. Check the result.

Core concepts

Next-token generation

What
The response is a sequence of draws. Each token is sampled from the distribution over the prefix so far.
Why
A wrong sentence is a path of draws, which is why a second call can diverge.
When
You are explaining a flaky completion or why a retry is not a fix.
When not
You are debugging an HTTP error. No tokens were generated.

Context window

What
The maximum tokens for the request plus the response being generated.
Why
The window and max_tokens are different caps. One can have room while the other is exhausted.
When
You are sizing a model or explaining a truncation.
When not
You are setting the bill. Price is per token, not per window.

Sampling

What
The rule that turns next-token probabilities into one chosen token. Temperature, top_p, and top_k are sampling controls where the model accepts them.
Why
They change variety. They do not create a deterministic contract.
When
The product needs more or less variation and the model allows the parameter.
When not
You need the same answer every time. Use a check, not a lower temperature alone.

Adaptive thinking

What
thinking type adaptive. The model decides when and how deeply to think. Effort steers that depth.
Why
On current Opus and Sonnet models this replaces a fixed thinking budget.
When
The task sometimes needs scratch work and sometimes does not.
When not
The model only accepts extended thinking, such as Haiku 4.5, or you meant an effort level named adaptive. That name is not an effort.

Effort

What
output_config.effort, from low through max, controlling how much work the response spends.
Why
It is the first lever for cost and latency inside one model.
When
A classification can be cheaper than a decision, on a model that supports the parameter.
When not
Haiku 4.5, which does not support effort, or a request that disables thinking at xhigh or max on Opus 5 and later.

Multi-shot prompting

What
Several complete examples of the input and the output you want, placed before the live task.
Why
Examples fix a format or an edge that instructions keep missing.
When
Zero-shot instructions are already clear and the misses are shape or boundary cases.
When not
The request is invalid or the output is truncated. Examples will not close a tool loop.

Practical examples

The cap that ate the answer

A proof task uses adaptive thinking and max_tokens 512. The response stops with stop_reason max_tokens. The content is a thinking block and no text. The context window is 1 million tokens and the input was short.

The window had room. The output cap did not, and thinking tokens spend that cap. Raise max_tokens so the answer has room after the scratch work, or lower effort if the evals still pass.

Two routers, one effort

A pipeline classifies a ticket, then decides a refund. Both calls use Opus at max effort. The classifier is a few labels. The decision quotes a policy.

Leave the decision at a high effort if the evals need it. Drop the classifier to low or medium, or move it to a smaller model. The two steps do not need the same rung.

Claude-specific considerations

  • Thinking tokens are output tokens. They count toward max_tokens even when display is omitted and the thinking field is empty.
  • adaptive is a thinking type. low, medium, high, xhigh, and max are effort levels.
  • Opus 5.5 defaults effort to medium. Several other current models default to high. Haiku 4.5 does not support effort.
  • On Opus 5 and later, thinking disabled together with effort xhigh or max is a 400.
  • Extended thinking with budget_tokens is rejected on models after the 4.6 generation. Those models use adaptive thinking.
  • Fast mode is speed fast plus the fast-mode beta, on supported Opus models on the Claude API, at a premium price.
  • A cache prefix changes when top-level effort or the thinking configuration changes.

Architecture decisions

SituationChooseBecause
The instruction is clear and the format still comes back wrong.Add one or more complete examples.The failure is the shape, and examples sit in the prefix the model continues.
A short classifier shares a model with a hard decision.A lower effort, or a smaller model, on the classifier only.Effort and model choice are per call.
Users are waiting and the Opus eval already passes.Fast mode on a supported Opus model, priced on the fast schedule.It raises output speed. It does not change which model you evaluated.
The same prompt produces two refunds.A check on the field, then a retry that includes the failure.Sampling can diverge. Another unchecked sample is the same process.
max_tokens is hit inside a thinking block.A larger output cap, or less effort if quality holds.Thinking and the answer share max_tokens.

Tradeoffs

More thinking, more examples, and a higher effort buy a better chance at a hard task and spend more of the window and the bill. Fast mode buys speed at a premium price and leaves the model's quality story to your evals.

AxisFewer tokens, less workMore work on this call
ReasoningA low effort, or thinking off where the model allows it.Adaptive thinking at a higher effort.
FormatZero-shot instructions.Single-shot or multi-shot examples.
LatencyFast mode on a supported Opus model.Standard speed, standard price.
CertaintyOne sample, accepted as-is.A check that rejects a bad sample.

Quick reference

  • A response is next-token sampling until a stop_reason. It is not a lookup of a stored answer.
  • The context window holds input and output. max_tokens caps output only. Thinking tokens spend that cap.
  • Current windows: Opus 5.5 and Sonnet 5.5 are 1 million tokens. Haiku 4.5 is 200,000.
  • Temperature and related sampling controls change variety where the model accepts them. They do not make a contract.
  • Extended thinking is budget_tokens. Later models want adaptive thinking and effort.
  • Effort lives in output_config.effort: low, medium, high, xhigh, max. adaptive is not one of those values.
  • Opus 5.5 defaults to medium effort. Haiku 4.5 has no effort parameter.
  • Disabling thinking at xhigh or max on Opus 5 and later is a 400.
  • Fast mode is a premium speed setting on supported Opus models on the Claude API.
  • Zero-shot is instructions. Single-shot is one example. Multi-shot is several, including the edge.

Decision rules for the exam

If the question says…The answer is likely…
"the same prompt, two different totals"Sampling. Check the field
"stop inside a thinking block"max_tokens was shared with thinking
"adaptive as an effort value"adaptive is a thinking type
"Haiku, set effort to low"Haiku 4.5 does not support effort
"disable thinking at max effort on Opus 5"That pair is a 400. Lower the effort or leave thinking on
"the format is wrong and the instructions are already specific"Add examples
"need tokens faster on Opus"Fast mode, premium price, Claude API only
"1 million token window, so max_tokens can be anything"max_tokens is still the output cap

Common exam traps

TrapCorrect answer
Temperature 0 means the answer is repeatable.Sampling can still diverge. A check is the control.
The context window and max_tokens are the same limit.The window fits the call. max_tokens caps the output.
Omitted thinking is free.The tokens are billed and they count toward max_tokens.
Fast mode is Haiku.Fast mode speeds a supported Opus model and costs a premium.
More rules beat one example when the shape is wrong.An example of the shape is the direct fix.

Open the Domain 5 sheet

Exam tips

  • If the stem says the text changed between identical calls, stay on sampling. Do not reach for a new model first.
  • If the stem names effort or thinking, check that the model supports that control.
  • Separate the window, the output cap, and the thinking budget. A truncation question is usually one of those three.

Common mistakes

  • Setting effort to adaptive.

    Use thinking type adaptive, and set effort to low, medium, high, xhigh, or max.

  • Judging a model from one pretty completion.

    The completion is one sample. Compare models on the evaluation set.

  • Putting a huge example set after the user text.

    Stable examples belong in the prefix, before the turn that changes every call.

  • Turning on fast mode and expecting the standard Opus price.

    Fast mode uses the premium schedule published for that model.

Practice questions

Original questions for this topic. They are study items, not questions from the live exam.

Two Messages calls send the same refund policy and the same receipt. The totals differ. The status is 200 and stop_reason is end_turn both times. What explains the difference?

Choose one answer

Adaptive thinking is on, the input is a short question, and stop_reason is max_tokens. The only content block is a thinking block. The model window is 1 million tokens. What is exhausted?

Choose one answer

A team sets output_config.effort to adaptive on Claude Opus 5.5 so the model will think only on hard tickets. What is wrong with that request?

Choose one answer

Support replies must quote a ticket id in one exact line. Longer instructions have not fixed the line. Which prompting change matches that miss?

Choose one answer

Interactive coding on Claude Opus 5.5 is correct on the eval set and too slow for the product. The team wants higher output speed on the Claude API. What fits?

Choose one answer

Scenario questions

The classifier and the decision

One service runs two calls. The first maps a support note to a category. The second decides whether a refund matches a written policy. Both calls use Claude Opus 5.5 at max effort. The category call is most of the bill. The refund call is the one customers dispute.

What change matches the two jobs?

Choose one answer

Build exercise

Label one call's controls

Beginner · 30 minutes

What you will learn

  • How a window, an output cap, and thinking share a request.
  • Where effort ends and adaptive thinking begins.
  • When an example belongs in the prompt.
  1. Step 1

    Write the job in one line

    Pick a real task you would send to Claude, such as classifying a note or checking a refund. Write the input size in a sentence and whether a person is waiting.

    Why: The controls depend on the job, not on a favorite parameter.

    You should see: One sentence with a rough size and a latency need.

  2. Step 2

    Separate the caps

    Name the model family you would start with, its window from the current lineup, and an output cap large enough for an answer after thinking.

    Why: max_tokens is not the window.

    You should see: A model, a window, and a max_tokens value greater than a short answer.

  3. Step 3

    Set thinking and effort

    Write the thinking type and the effort you would send. If the model has no effort control, say so.

    Why: The names are easy to swap.

    You should see: thinking type adaptive or enabled, and an effort rung or an explicit none.

  4. Step 4

    Choose the shot

    Decide zero-shot, single-shot, or multi-shot. If you add examples, sketch one that includes the edge you care about, and put it before the live user text.

    Why: Examples are a prefix, not a trailing hint.

    You should see: A choice and, if needed, one short example.

Review checklist

Checks are saved in this browser.

Key takeaways

  • A Claude answer is a chain of sampled tokens. Identical inputs can diverge. A check beats another untested sample.
  • The window holds the call. max_tokens holds the output. Thinking spends the output cap.
  • Adaptive thinking and effort are different fields. Fast mode is a premium speed setting, not a model family.
  • Examples earn their tokens when the missing piece is a shape or an edge the instructions already name.

Sources

  • Models overview — Windows, default effort, thinking mode, and prices for the current lineup.
  • Thinking — Adaptive and extended thinking, and how thinking tokens spend max_tokens.
  • Effort — output_config.effort and how a top-level change affects the cache.
  • Build with Claude — Messages, prompting, and the application around the model.
  • CCDV-F blueprint notes — Domain 5 skill list and weight. Study notes, not exam items.

Domain 5 overview · Quick reference

View progress