CCDV-F · Study Guide

← Domain 5: Model Selection and Optimization

5.3 · 2.7% of the exam · Topic 3 of 4

Model Selection and Tradeoffs

Haiku, Sonnet, and Opus are a speed, quality, and price ladder. Effort and fast mode move a call along that ladder without a new name. A new model id is a behavior change: pin the snapshot you evaluated, and rerun the eval before you serve it.

Learning objectives

  • Place Haiku, Sonnet, and Opus on quality, latency, and cost.
  • Prefer effort, routing, and fast mode before a jump to a larger model.
  • Treat a model release as a change that can break prompts, tools, and thinking settings.
  • Pin the model id that the evaluation actually ran.

Detailed theory

What this skill covers

Model selection and tradeoffs are 2.7% of the exam. The skill is the Opus, Sonnet, and Haiku use cases, the quality, latency, and cost tradeoff, and behavior that changes between releases.

The current catalog also includes models outside those three names. The decision the exam is practicing is still which rung the task needs, and what you re-check when the id changes.

The three families

Haiku is the fastest current family member and the one you point at high volume and simpler work. Haiku 4.5 is priced at $1 per million input tokens and $5 per million output tokens, with a 200,000-token window. Its thinking mode is extended thinking. It does not take an effort level. Reach for it when the eval set already passes and the bill or the latency is dominated by easy calls.

Sonnet is the balance of speed and intelligence. Sonnet 5.5 is $2 and $10 per million input and output tokens, with a 1 million token window, adaptive thinking, and a default effort of high. It is the usual home for a product task that is harder than a label and cheaper than your hardest reasoning.

Opus is the deep reasoning and long agentic work. Opus 5.5 is $4 and $20 per million input and output tokens, with a 1 million token window. Adaptive thinking is always on, and the default effort is medium. The models overview says to start with Opus 5.5 when you are unsure. That is a starting point for quality, not a reason to put it on every route. A model above Opus exists in the current lineup for work whose evals still fail on Opus 5.5. You still justify the step with those evals.

Quality, latency, and cost

The three move together. Opus costs more per token and spends more time, especially when it thinks. Haiku returns sooner and costs less, and it will lose on tasks that need the longer reasoning. Sonnet sits between them. Output tokens cost more than input tokens on every rung, so a chatty answer dominates a short prompt.

Inside one model, effort is the finer knob. Drop it where the eval still passes. Raise it before you rename the model. Fast mode, on supported Opus models on the Claude API, buys output speed at a premium price. It is how you keep Opus when the eval needs Opus and the product needs tokens sooner. It is not a substitute for Haiku on a task Haiku already passes.

Route the easy step and the hard step apart. A classifier in front of a policy decision should not inherit the decision's model. Subagents that search or label can be Haiku while the parent stays on Opus. Measure the split on the eval set. A demo transcript is one sample.

Releases change behavior

A new model id can follow the same prompt differently: tool choice, format, tone, how much it thinks, and which parameters it accepts. Extended thinking with budget_tokens is rejected on models after the 4.6 generation. Opus 5 and later return a 400 if you disable thinking at effort xhigh or max. Opus 5.5's default effort is medium, where Opus 5's default was high. Carrying a configuration across that line without an eval is how a silent quality shift ships.

Current Claude API model ids are pinned snapshots, including the dateless ids from the 4.6 generation onward. Pin the id you tested. An alias that floats to a future release is a different deployment. The same string is not the same snapshot on every platform. Bedrock, Google Cloud, Foundry, and the Claude API publish their own ids. Pin the one on the platform that served the eval.

Retirement dates are part of the choice. Haiku 4.5's commitment runs no sooner than October 2026 on Anthropic-operated platforms. The Opus 5.5 and Sonnet 5.5 commitments run no sooner than September 2027. A pinned id still ends. Plan the next eval before the date, not after the traffic has moved.

A selection order

Write the behaviors the product promises. Run them on the candidate. Start from the docs' suggestion when you have no data, then move down a rung wherever the set still passes. Adjust effort on the rung you kept. Turn on fast mode only when that rung is required and latency is still the complaint, and only where the preview exists.

When a release notes page says the default effort, the thinking mode, or the sampling parameters changed, rerun the set. Update the request the new model rejects. Ship the new id when the set, not a single chat, says the behaviors still hold.

Core concepts

Haiku

What
The fastest family, for high volume and simpler tasks that already pass an eval.
Why
It is the low price and the low latency, with a smaller window on Haiku 4.5.
When
The task is a label, a short extract, or a fan-out, and the set passes.
When not
The set fails and the failure is reasoning depth. A lower price will not invent that depth.

Sonnet

What
The balance of speed and intelligence for product work between a label and the hardest reasoning.
Why
It is the middle of the price and latency ladder.
When
Haiku misses the eval and Opus is more than the task needs.
When not
You have not measured. The name is not a default for every route.

Opus

What
The family for hard reasoning and long agentic work.
Why
Quality on those tasks is why the higher price exists.
When
The eval fails on smaller models, or the work is a long tool loop that needs the deeper model.
When not
A cheap classification you attached out of habit. Route that call down.

Pinned model id

What
The exact id whose eval you are willing to serve.
Why
A floating alias or a platform's different id is a different behavior until you test it.
When
Production, and any job that quotes a previous score.
When not
A local sketch where you intend to throw the transcript away.

Release regression

What
A behavior change that arrives with a new id: prompts, tools, thinking settings, or defaults.
Why
The previous eval described the previous id.
When
You move from one snapshot to the next, including a default effort that shifted.
When not
You are comparing samples of the same pinned id. That spread is sampling.

Practical examples

One model on every route

A support product sends classification, retrieval ranking, and the customer-facing answer to Opus 5.5 at high effort. The answer is a small share of the calls and most of the complaints. The classifier is most of the tokens.

Keep Opus on the answer if that is what passes. Move the classifier to Haiku, or to a lower effort on a model that has effort, after the category eval passes. The bill follows the tokens you stop sending to the top rung.

The weekend upgrade

Production pins an older Opus id. A deploy swaps in the next id because the alias was convenient. Refund tone and tool choice shift on Monday. No eval ran.

Put the tested id back. Run the behavior set on the new snapshot, including the thinking and effort fields the new model accepts. Ship the new id with the diff of that set.

Claude-specific considerations

  • Haiku 4.5: $1 / $5 per million tokens, 200k window, extended thinking, no effort parameter.
  • Sonnet 5.5: $2 / $10, 1M window, adaptive thinking, default effort high.
  • Opus 5.5: $4 / $20, 1M window, adaptive thinking always on, default effort medium.
  • The models overview suggests starting with Opus 5.5 when you are unsure, then proving a smaller model with evals.
  • Fast mode is an Opus speed upgrade on the Claude API, at a premium price, not a Haiku rename.
  • Dateless ids from the 4.6 generation on are pinned snapshots. Older aliases can move. Pin what you tested.
  • Disabling thinking at xhigh or max fails on Opus 5 and later. budget_tokens fails on models that only accept adaptive thinking.

Architecture decisions

SituationChooseBecause
A high-volume label already passes on Haiku.Haiku.The extra Opus tokens buy nothing the set can see.
A customer-facing answer fails the policy eval on Sonnet.Opus, then an effort sweep.The failure is quality. Latency comes after a passing set.
Opus passes and the product is still slow.Lower effort if the set allows it, or fast mode where it exists.Both keep the model that passed. They change spend and speed.
A new snapshot is available and the old one is pinned.An eval on the new id before production traffic.Defaults, thinking, and tool behavior can move.
The same prompt was tested on the Claude API and will run on Bedrock.The Bedrock id, with the eval repeated there.The platform id is part of the deployment.

Tradeoffs

Moving up the family buys a better chance at hard work and a higher price on every token. Moving down buys speed and a risk the eval will catch. Effort and fast mode sit inside that choice.

AxisFaster and cheaperMore capable on the hard set
FamilyHaiku for volume that already passes.Opus for reasoning that does not.
Inside OpusLower effort, or fast mode if latency remains.Higher effort when the set moves.
ReleaseThe pinned id you scored.A new id after a fresh score.
RouteA small model on the easy step.The large model on the step that fails without it.

Quick reference

  • Haiku: fastest, cheapest, high volume, simpler tasks. Haiku 4.5 has a 200k window and no effort parameter.
  • Sonnet: speed and intelligence in the middle. Sonnet 5.5 defaults to high effort and a 1M window.
  • Opus: hard reasoning and long agentic work. Opus 5.5 defaults to medium effort. Adaptive thinking stays on.
  • Start from a measured set. The smallest passing model wins.
  • Effort is the lever inside a model. Fast mode is a premium speed on supported Opus models.
  • A new id can change tool use, format, and which thinking fields are legal.
  • Pin the snapshot you evaluated. Repeat the eval on a new id and on a different platform's id.
  • Output tokens cost more than input tokens on every family.
  • Do not promote a model from one flattering transcript.

Decision rules for the exam

If the question says…The answer is likely…
"high volume, already accurate on the small model"Haiku
"the policy answer fails on Sonnet"Opus, then remeasure
"Opus is correct and too slow"Lower effort if the set holds, or fast mode
"use the alias, it will track the latest"Pin the evaluated snapshot
"we tested the Claude API id and deployed the Bedrock id"Test the id you serve
"budget_tokens on a current Opus"That model expects adaptive thinking and effort
"one demo looked better"Compare the eval set
"put Opus on the classifier to be safe"Route the easy call down

Common exam traps

TrapCorrect answer
The largest model is the safe default for every route.Use the smallest model that passes the set.
Fast mode is how you turn Opus into Haiku prices.Fast mode is a premium speed on Opus.
A new release honors the old request byte for byte.Thinking fields and default effort can change and reject or shift.
Sonnet, Opus, and Haiku share one context window.Haiku 4.5 is 200k. Sonnet 5.5 and Opus 5.5 are 1M.
Quality, latency, and cost can all be maxed.The family is a tradeoff. Effort only tunes inside it.

Open the Domain 5 sheet

Exam tips

  • The stem usually names a constraint: volume, a failing eval, or a release. Answer that constraint.
  • A request for the newest model without an eval is the wrong change.
  • Keep Haiku's missing effort parameter in mind when the stem sets a rung on every model.

Common mistakes

  • One Opus route for labels and decisions.

    Split the calls. Spend Opus where the set requires it.

  • Shipping the next model id over a weekend.

    Pin the tested id until the new id has its own eval.

  • Copying a budget_tokens body onto Opus 5.5.

    Use adaptive thinking and effort. Rerun the set at the new default.

  • Treating a single nicer transcript as a model decision.

    Sampling can flatter one call. The set is the comparison.

Practice questions

Original questions for this topic. They are study items, not questions from the live exam.

A classifier labels 2 million tickets a day. On the eval set, Haiku 4.5 matches the larger models. The bill is dominated by this route. Which deployment fits?

Choose one answer

Refund decisions fail the policy set on Sonnet 5.5 and pass on Opus 5.5 at medium effort. Latency is acceptable. What should production pin?

Choose one answer

An Opus 4.8 request disables thinking and sets effort to max. The same body returns 400 on Opus 5. What changed?

Choose one answer

Opus 5.5 passes the coding eval and the product wants tokens on screen faster on the Claude API. Price may rise. What matches the constraint?

Choose one answer

Scenario questions

The alias that moved

Production reads the model from a convenience alias. Last quarter's eval used a dated Sonnet snapshot. This week the alias points at a newer Sonnet. Support tickets mention a different tone and a tool the new model calls more often. The team asks whether to raise temperature or to add more examples first.

What should happen before any prompt edit?

Choose one answer

Build exercise

Pick a rung and a pin

Intermediate · 40 minutes

What you will learn

  • How volume and a failing set point at different families.
  • What you re-test when the id changes.
  • Where effort and fast mode sit relative to a rename.
  1. Step 1

    List three routes

    Write a classifier, a customer answer, and a long coding agent. Add a daily volume and whether a person waits.

    Why: One product is often three models.

    You should see: Three rows.

  2. Step 2

    Assign a family

    Put Haiku, Sonnet, or Opus on each row, and one sentence that would falsify the choice on an eval.

    Why: The choice is a hypothesis about the set.

    You should see: A model and a failure that would promote or demote it.

  3. Step 3

    Add the inside levers

    For the Opus row, note the default effort on Opus 5.5 and whether fast mode is even a candidate. For Haiku, note that effort is absent.

    Why: The stem will offer effort on a model that rejects it.

    You should see: Medium as the Opus 5.5 default, and no effort field on Haiku 4.5.

  4. Step 4

    Write the release rule

    Name the pinned id string you would store, and the checks you rerun before replacing it: the behavior set, the thinking fields, and the platform id.

    Why: A release is a change, not a refresh.

    You should see: A pin and a three-item rerun list.

Review checklist

Checks are saved in this browser.

Key takeaways

  • Haiku, Sonnet, and Opus are a ladder of speed, price, and reasoning. Serve the smallest model that passes.
  • Effort tunes a model. Fast mode speeds supported Opus models at a premium. Neither one is a new eval.
  • A new id can reject old thinking settings and can change behavior. Pin the snapshot, then score the next one.

Sources

Domain 5 overview · Quick reference

View progress