5.3 · 2.7% of the exam · Topic 3 of 4
Model Selection and Tradeoffs
Haiku, Sonnet, and Opus are a speed, quality, and price ladder. Effort and fast mode move a call along that ladder without a new name. A new model id is a behavior change: pin the snapshot you evaluated, and rerun the eval before you serve it.
Learning objectives
- Place Haiku, Sonnet, and Opus on quality, latency, and cost.
- Prefer effort, routing, and fast mode before a jump to a larger model.
- Treat a model release as a change that can break prompts, tools, and thinking settings.
- Pin the model id that the evaluation actually ran.
Detailed theory
What this skill covers
Model selection and tradeoffs are 2.7% of the exam. The skill is the Opus, Sonnet, and Haiku use cases, the quality, latency, and cost tradeoff, and behavior that changes between releases.
The current catalog also includes models outside those three names. The decision the exam is practicing is still which rung the task needs, and what you re-check when the id changes.
The three families
Haiku is the fastest current family member and the one you point at high volume and simpler work. Haiku 4.5 is priced at $1 per million input tokens and $5 per million output tokens, with a 200,000-token window. Its thinking mode is extended thinking. It does not take an effort level. Reach for it when the eval set already passes and the bill or the latency is dominated by easy calls.
Sonnet is the balance of speed and intelligence. Sonnet 5.5 is $2 and $10 per million input and output tokens, with a 1 million token window, adaptive thinking, and a default effort of high. It is the usual home for a product task that is harder than a label and cheaper than your hardest reasoning.
Opus is the deep reasoning and long agentic work. Opus 5.5 is $4 and $20 per million input and output tokens, with a 1 million token window. Adaptive thinking is always on, and the default effort is medium. The models overview says to start with Opus 5.5 when you are unsure. That is a starting point for quality, not a reason to put it on every route. A model above Opus exists in the current lineup for work whose evals still fail on Opus 5.5. You still justify the step with those evals.
Quality, latency, and cost
The three move together. Opus costs more per token and spends more time, especially when it thinks. Haiku returns sooner and costs less, and it will lose on tasks that need the longer reasoning. Sonnet sits between them. Output tokens cost more than input tokens on every rung, so a chatty answer dominates a short prompt.
Inside one model, effort is the finer knob. Drop it where the eval still passes. Raise it before you rename the model. Fast mode, on supported Opus models on the Claude API, buys output speed at a premium price. It is how you keep Opus when the eval needs Opus and the product needs tokens sooner. It is not a substitute for Haiku on a task Haiku already passes.
Route the easy step and the hard step apart. A classifier in front of a policy decision should not inherit the decision's model. Subagents that search or label can be Haiku while the parent stays on Opus. Measure the split on the eval set. A demo transcript is one sample.
Releases change behavior
A new model id can follow the same prompt differently: tool choice, format, tone, how much it thinks, and which parameters it accepts. Extended thinking with budget_tokens is rejected on models after the 4.6 generation. Opus 5 and later return a 400 if you disable thinking at effort xhigh or max. Opus 5.5's default effort is medium, where Opus 5's default was high. Carrying a configuration across that line without an eval is how a silent quality shift ships.
Current Claude API model ids are pinned snapshots, including the dateless ids from the 4.6 generation onward. Pin the id you tested. An alias that floats to a future release is a different deployment. The same string is not the same snapshot on every platform. Bedrock, Google Cloud, Foundry, and the Claude API publish their own ids. Pin the one on the platform that served the eval.
Retirement dates are part of the choice. Haiku 4.5's commitment runs no sooner than October 2026 on Anthropic-operated platforms. The Opus 5.5 and Sonnet 5.5 commitments run no sooner than September 2027. A pinned id still ends. Plan the next eval before the date, not after the traffic has moved.
A selection order
Write the behaviors the product promises. Run them on the candidate. Start from the docs' suggestion when you have no data, then move down a rung wherever the set still passes. Adjust effort on the rung you kept. Turn on fast mode only when that rung is required and latency is still the complaint, and only where the preview exists.
When a release notes page says the default effort, the thinking mode, or the sampling parameters changed, rerun the set. Update the request the new model rejects. Ship the new id when the set, not a single chat, says the behaviors still hold.
Core concepts
Haiku
- What
- The fastest family, for high volume and simpler tasks that already pass an eval.
- Why
- It is the low price and the low latency, with a smaller window on Haiku 4.5.
- When
- The task is a label, a short extract, or a fan-out, and the set passes.
- When not
- The set fails and the failure is reasoning depth. A lower price will not invent that depth.
Sonnet
- What
- The balance of speed and intelligence for product work between a label and the hardest reasoning.
- Why
- It is the middle of the price and latency ladder.
- When
- Haiku misses the eval and Opus is more than the task needs.
- When not
- You have not measured. The name is not a default for every route.
Opus
- What
- The family for hard reasoning and long agentic work.
- Why
- Quality on those tasks is why the higher price exists.
- When
- The eval fails on smaller models, or the work is a long tool loop that needs the deeper model.
- When not
- A cheap classification you attached out of habit. Route that call down.
Pinned model id
- What
- The exact id whose eval you are willing to serve.
- Why
- A floating alias or a platform's different id is a different behavior until you test it.
- When
- Production, and any job that quotes a previous score.
- When not
- A local sketch where you intend to throw the transcript away.
Release regression
- What
- A behavior change that arrives with a new id: prompts, tools, thinking settings, or defaults.
- Why
- The previous eval described the previous id.
- When
- You move from one snapshot to the next, including a default effort that shifted.
- When not
- You are comparing samples of the same pinned id. That spread is sampling.
Practical examples
One model on every route
A support product sends classification, retrieval ranking, and the customer-facing answer to Opus 5.5 at high effort. The answer is a small share of the calls and most of the complaints. The classifier is most of the tokens.
Keep Opus on the answer if that is what passes. Move the classifier to Haiku, or to a lower effort on a model that has effort, after the category eval passes. The bill follows the tokens you stop sending to the top rung.
The weekend upgrade
Production pins an older Opus id. A deploy swaps in the next id because the alias was convenient. Refund tone and tool choice shift on Monday. No eval ran.
Put the tested id back. Run the behavior set on the new snapshot, including the thinking and effort fields the new model accepts. Ship the new id with the diff of that set.
Claude-specific considerations
- Haiku 4.5: $1 / $5 per million tokens, 200k window, extended thinking, no effort parameter.
- Sonnet 5.5: $2 / $10, 1M window, adaptive thinking, default effort high.
- Opus 5.5: $4 / $20, 1M window, adaptive thinking always on, default effort medium.
- The models overview suggests starting with Opus 5.5 when you are unsure, then proving a smaller model with evals.
- Fast mode is an Opus speed upgrade on the Claude API, at a premium price, not a Haiku rename.
- Dateless ids from the 4.6 generation on are pinned snapshots. Older aliases can move. Pin what you tested.
- Disabling thinking at xhigh or max fails on Opus 5 and later. budget_tokens fails on models that only accept adaptive thinking.
Architecture decisions
Tradeoffs
Moving up the family buys a better chance at hard work and a higher price on every token. Moving down buys speed and a risk the eval will catch. Effort and fast mode sit inside that choice.
Quick reference
- Haiku: fastest, cheapest, high volume, simpler tasks. Haiku 4.5 has a 200k window and no effort parameter.
- Sonnet: speed and intelligence in the middle. Sonnet 5.5 defaults to high effort and a 1M window.
- Opus: hard reasoning and long agentic work. Opus 5.5 defaults to medium effort. Adaptive thinking stays on.
- Start from a measured set. The smallest passing model wins.
- Effort is the lever inside a model. Fast mode is a premium speed on supported Opus models.
- A new id can change tool use, format, and which thinking fields are legal.
- Pin the snapshot you evaluated. Repeat the eval on a new id and on a different platform's id.
- Output tokens cost more than input tokens on every family.
- Do not promote a model from one flattering transcript.
Decision rules for the exam
Common exam traps
Exam tips
- The stem usually names a constraint: volume, a failing eval, or a release. Answer that constraint.
- A request for the newest model without an eval is the wrong change.
- Keep Haiku's missing effort parameter in mind when the stem sets a rung on every model.
Common mistakes
One Opus route for labels and decisions.
Split the calls. Spend Opus where the set requires it.
Shipping the next model id over a weekend.
Pin the tested id until the new id has its own eval.
Copying a budget_tokens body onto Opus 5.5.
Use adaptive thinking and effort. Rerun the set at the new default.
Treating a single nicer transcript as a model decision.
Sampling can flatter one call. The set is the comparison.
Practice questions
Original questions for this topic. They are study items, not questions from the live exam.
A classifier labels 2 million tickets a day. On the eval set, Haiku 4.5 matches the larger models. The bill is dominated by this route. Which deployment fits?
Refund decisions fail the policy set on Sonnet 5.5 and pass on Opus 5.5 at medium effort. Latency is acceptable. What should production pin?
An Opus 4.8 request disables thinking and sets effort to max. The same body returns 400 on Opus 5. What changed?
Opus 5.5 passes the coding eval and the product wants tokens on screen faster on the Claude API. Price may rise. What matches the constraint?
Scenario questions
The alias that moved
Production reads the model from a convenience alias. Last quarter's eval used a dated Sonnet snapshot. This week the alias points at a newer Sonnet. Support tickets mention a different tone and a tool the new model calls more often. The team asks whether to raise temperature or to add more examples first.
What should happen before any prompt edit?
Build exercise
Pick a rung and a pin
Intermediate · 40 minutes
What you will learn
- How volume and a failing set point at different families.
- What you re-test when the id changes.
- Where effort and fast mode sit relative to a rename.
Step 1
List three routes
Write a classifier, a customer answer, and a long coding agent. Add a daily volume and whether a person waits.
Why: One product is often three models.
You should see: Three rows.
Step 2
Assign a family
Put Haiku, Sonnet, or Opus on each row, and one sentence that would falsify the choice on an eval.
Why: The choice is a hypothesis about the set.
You should see: A model and a failure that would promote or demote it.
Step 3
Add the inside levers
For the Opus row, note the default effort on Opus 5.5 and whether fast mode is even a candidate. For Haiku, note that effort is absent.
Why: The stem will offer effort on a model that rejects it.
You should see: Medium as the Opus 5.5 default, and no effort field on Haiku 4.5.
Step 4
Write the release rule
Name the pinned id string you would store, and the checks you rerun before replacing it: the behavior set, the thinking fields, and the platform id.
Why: A release is a change, not a refresh.
You should see: A pin and a three-item rerun list.
Review checklist
Checks are saved in this browser.
Key takeaways
- Haiku, Sonnet, and Opus are a ladder of speed, price, and reasoning. Serve the smallest model that passes.
- Effort tunes a model. Fast mode speeds supported Opus models at a premium. Neither one is a new eval.
- A new id can reject old thinking settings and can change behavior. Pin the snapshot, then score the next one.
Sources
- Models overview — Current Haiku, Sonnet, and Opus prices, windows, thinking, and default effort.
- Choosing a model — How speed, effort, and fast mode sit next to the model family.
- Build with Claude — The request that carries the model id.
- CCDV-F blueprint notes — Domain 5 selection skill and weight. Study notes, not exam items.