CCDV-F · Study Guide

← Domain 5: Model Selection and Optimization

5.4 · 2.8% of the exam · Topic 4 of 4

Cost and Token Management

The bill is the usage object times the price card. Output and uncached input dominate. Prompt caching stores a shared prefix at a cache checkpoint so later calls read it at a fraction of the input price. The checkpoint only hits when that prefix, including thinking and effort, stays the same.

Learning objectives

  • Read input, output, cache write, and cache read from a usage object.
  • Sketch a call's cost from those counts and the model's prices.
  • Place a cache checkpoint at the end of the stable prefix.
  • Name the edits that move the checkpoint and drop the read.

Detailed theory

What this skill covers

Cost and token management is 2.8% of the exam: usage tracking, cost modeling, prompt caching, and cache checkpoints.

The previous topics decide how many tokens a design will generate. This topic is how those tokens show up in usage and what a second call can avoid paying for again.

What usage counts

A successful message reports usage. input_tokens are the uncached input billed at the full input price. output_tokens are the generated tokens, including thinking, billed at the output price. cache_creation_input_tokens are the prefix written into the cache. cache_read_input_tokens are the prefix served from a hit.

A word count will not reconcile an invoice. The token counting endpoint can estimate input before you spend the call. It does not know the output you have not generated, and it does not replace the usage on the response you already received.

A conversation resends the history. Turn ten pays input on everything you failed to cache, plus output for the new answer. The expensive pattern is a long thread with no cache and a large output. The report you want per route is calls, input, output, cache writes, and cache reads, summed to a dollar figure with the price card for that model id.

A cost sketch

Price cards are dollars per million tokens, and input and output differ. On the current lineup, output is five times the input price: Haiku 4.5 is $1 and $5, Sonnet 5.5 is $2 and $10, Opus 5.5 is $4 and $20. Multiply each usage bucket by its rate. Do not multiply a single blended rate by the sum of input and output.

Standard cache writes cost 1.25 times the base input price. Standard cache reads cost 0.1 times the base input price. The models overview notes cheaper reads on some current models: 5% of base input on Opus 5.5, and 2.5% on Claude Fable 5.1 and Claude Mythos 5.1. Use the card for the id you pin.

The Message Batches API is half price for work that can wait and does not need a multi-turn tool loop inside the batch. Fast mode uses its own premium card, so a model priced at $4 and $20 is not the fast-mode price. A cost model that forgets the mode will underbill the interactive route.

cost = input_tokens * input_rate
     + cache_creation_input_tokens * write_rate
     + cache_read_input_tokens * read_rate
     + output_tokens * output_rate

Prompt caching and the checkpoint

A cache hit resumes a prefix. Anything before the checkpoint must match the bytes that were written, in order. Stable material goes first: tools, the system prompt, policies, and examples. The user text that changes every call goes after the checkpoint.

Automatic caching sets cache_control to type ephemeral at the top level of the request. The service places the breakpoint on the last cacheable block and moves it forward as a conversation grows. That is the right default for a multi-turn thread whose history should be reused.

An explicit checkpoint puts cache_control on one content block. You may place up to four. Use them when the shared prefix ends before the last block, when you want a different lifetime on different blocks, or when you are pre-warming. A pre-warm must put the breakpoint on the last shared block, usually the system prompt or the tools. If automatic caching lands the breakpoint on a placeholder user message, the real user message is a different prefix and the read misses.

A prefix under the model's minimum length does not cache. Creation tokens stay at zero and the call still succeeds. Published minima include 1,024 tokens for Sonnet and 4,096 for Opus and for Haiku 4.5. A short system prompt will not show a write no matter how correctly you spelled cache_control.

What drops the read

The default lifetime is five minutes, and a hit refreshes it. A one-hour lifetime is available at a higher write price. A quiet route that calls every hour will miss a five-minute entry and pay to write again. A noisy route keeps the five-minute entry warm for free on each hit.

These changes start a new prefix: editing the system text or the tools, reordering blocks, switching model, switching fast mode, changing the thinking configuration, and changing top-level effort. Setting a parameter explicitly to its default is the same as omitting it and does not, by itself, break the prefix. On models that support per-message effort, an effort change carried in a system turn inside messages can leave the cached prefix intact. A top-level effort change does not.

Thinking blocks that you pass back unchanged stay compatible with the cache. Dropping or editing them changes the prefix from that point on, and it is also the 400 described in debugging. Tool results become part of the next cached prefix once you send them. That is useful, and it means a different tool result is a different suffix.

Using the numbers

Track usage in production per route and per model id. Alert on output tokens and on a cache read rate that fell, not only on the dollar total. A read rate that falls to zero after a prompt edit is a checkpoint bug. A dollar spike with a healthy read rate is usually more output or more calls.

Model the steady state, not the first call. The first call writes. Later calls read, if the prefix held. A demo that runs once will look like caching failed, because a write is supposed to happen once. Compare the second call's cache_read_input_tokens to the prefix length.

Core concepts

Usage object

What
The token buckets on a response: uncached input, output, cache creation, and cache read.
Why
The invoice is these counts times the price card.
When
You are explaining a bill or checking that a cache hit happened.
When not
You are estimating a prompt you have not sent. Use the token counting endpoint, and still expect output to be unknown.

Cache checkpoint

What
The cache_control breakpoint at the end of the prefix you want reused.
Why
Only the bytes up to that point can hit. A later edit of those bytes misses.
When
Tools, system text, or history repeat across calls.
When not
The shared text is below the model's minimum cache length. The marker will not write.

Automatic caching

What
A top-level ephemeral cache_control. The breakpoint rides the last cacheable block and moves forward.
Why
A growing conversation can reuse the prefix without you moving a marker.
When
The whole history so far is the shared prefix.
When not
You are pre-warming a system prompt and the last block is a dummy user message. Use an explicit breakpoint before that dummy.

Cache read

What
Tokens served from a matching prefix, billed at a fraction of the input price.
Why
This is the saving. A write on every call is not a saving.
When
cache_read_input_tokens is non-zero on the second call.
When not
The first call, which should show creation tokens instead.

Price card

What
The input, output, cache write, cache read, batch, and fast-mode rates for one model id.
Why
Families differ, and fast mode does not use the standard card.
When
You turn usage counts into money.
When not
You compare quality. A cheaper card does not describe the eval.

Practical examples

The write that never became a read

Call one shows cache_creation_input_tokens and no reads. Call two, a minute later with the same tools and system prompt, shows another write and zero reads. The only edit was output_config.effort from medium to high.

Top-level effort is part of the prefix. The second call wrote a new entry. Put effort back, or keep one effort for the life of the cached conversation. Then the second usage line should show cache reads.

The short policy

A 400-token system prompt carries cache_control. Usage shows input tokens and zeros in both cache fields. The call succeeds.

The prefix is under the model minimum, so nothing is written. Lengthen the stable prefix with the material you were already sending, such as the tools and the examples, until it clears the minimum. A longer dummy sentence that you do not need is the wrong way to get there.

Claude-specific considerations

  • Read usage on the response. The four buckets are uncached input, output, cache creation, and cache read.
  • Output is priced higher than input. Thinking is output.
  • Standard cache writes are 1.25 times base input. Standard cache reads are 0.1 times base input, with a lower read rate on some current model ids.
  • Automatic caching is top-level cache_control ephemeral. Explicit breakpoints go on blocks, up to four.
  • Minima include 1,024 tokens for Sonnet and 4,096 for Opus and Haiku 4.5. Below that, creation stays zero.
  • The default TTL is five minutes and a hit refreshes it. A one-hour TTL costs more to write.
  • A top-level effort change, a thinking-config change, a model change, and a fast-mode change all start a new prefix.
  • Batch pricing is half, for work that can wait. Fast mode has its own card.

Architecture decisions

SituationChooseBecause
The same tools and policy go out on every call.A cache checkpoint at the end of that prefix.The second call can read it.
A multi-turn chat grows by one user message each time.Automatic caching.The breakpoint moves forward with the history.
You want to warm the cache before real users arrive.An explicit breakpoint on the shared system or tools, with the same thinking and effort the traffic will use.A dummy user message at the breakpoint will not match the real one.
Reads dropped to zero after a one-line effort change.Restore the effort, or accept a new write and then stable reads.Top-level effort is inside the prefix.
Overnight extraction can wait and does not need a live tool loop.The Batches API, at half the token price.The discount follows the latency you can give up.

Tradeoffs

A cache write costs more than a normal input token and pays off on the next reads. A longer TTL costs more to write and covers a quiet route. A checkpoint that is too early leaves savings on the table. A checkpoint that includes a changing token never hits.

AxisPay full input againPay a write, then read cheaply
Repeated prefixNo cache_control.A checkpoint after the stable blocks.
ConversationA marker you forget to move.Automatic caching, which moves it.
LifetimeFive minutes, refreshed by traffic.One hour, when calls are sparse.
First callA write, which looks expensive alone.Reads on the calls that share the prefix.

Quick reference

  • Cost uses four buckets: uncached input, cache write, cache read, and output.
  • Output tokens, including thinking, use the output price. Do not blend them with input.
  • Current sticker prices: Haiku 4.5 $1/$5, Sonnet 5.5 $2/$10, Opus 5.5 $4/$20 per million input and output tokens.
  • Standard cache write is 1.25× input. Standard cache read is 0.1× input. Check the card for model-specific read rates.
  • The checkpoint is the end of the shared prefix. Stable tools and system text first. The changing user text after.
  • Automatic caching moves the breakpoint forward. Explicit breakpoints: up to four, on the blocks you name.
  • Under the minimum length, the call succeeds and the cache fields stay zero.
  • Five-minute TTL by default, refreshed on a hit. A one-hour TTL costs more to write.
  • Changing model, fast mode, thinking config, or top-level effort misses the old prefix.
  • The first call writes. Judge the hit on the second call's cache_read_input_tokens.

Decision rules for the exam

If the question says…The answer is likely…
"cache_read is zero on the first call"Expected. Look at the second call
"reads died after an effort edit"Top-level effort is part of the prefix
"the marker is set and creation is zero"The prefix is under the minimum length
"the user id is in the system prompt"That token changes the prefix. Move it after the checkpoint
"warm the cache with a dummy user message at the end"Put the explicit breakpoint before the dummy
"hourly job, five-minute TTL"The entry expires. Use the longer TTL or accept a rewrite
"fast mode for the demo, standard speed in prod"The speed change misses the other prefix
"bill equals tokens times one rate"Split input, output, write, and read

Common exam traps

TrapCorrect answer
Caching makes output tokens cheaper.Reads discount the prefix. Output is still output.
A cache_control field guarantees a write.The prefix must clear the model's minimum length.
Automatic caching is right for a pre-warm that ends in a placeholder.The breakpoint would sit on the placeholder. Use an explicit one.
Batch half-price applies to a user who is waiting.Batch is for work that can wait. Interactive calls use the standard card.
Word count times a dollar figure is a cost model.Use the usage buckets and the price card.

Open the Domain 5 sheet

Exam tips

  • If the stem shows a usage object, do the arithmetic with separate rates.
  • If the second call writes again, something in the prefix changed, or the TTL elapsed, or the text never cleared the minimum.
  • A checkpoint question is about order: shared bytes first.

Common mistakes

  • Putting the account id in the system prompt you cache.

    Keep the system prompt identical. Put the account id in the user turn after the checkpoint.

  • Declaring caching broken because the first response has creation tokens and no reads.

    Send the same prefix again and read cache_read_input_tokens.

  • Changing effort between turns of a cached Opus chat.

    Keep top-level effort stable, or use a per-message effort change on a model that preserves the prefix.

  • Modeling fast-mode traffic at the standard Opus price.

    Use the fast-mode card for any call that sets speed fast.

Practice questions

Original questions for this topic. They are study items, not questions from the live exam.

A response reports 200 cache_read_input_tokens, 50 input_tokens, 0 cache_creation_input_tokens, and 80 output_tokens. Which bucket uses the output price?

Choose one answer

The system prompt and tool list are identical on two calls a minute apart. The second call shows a fresh cache write and zero reads. The only request change is output_config.effort from medium to high. Why did the read miss?

Choose one answer

A 300-token system prompt on Sonnet includes cache_control and the call succeeds. Both cache fields are zero. What happened?

Choose one answer

You want the tools and the policy cached before traffic arrives. The warm-up request ends with a placeholder user message of warmup. Where does the breakpoint go?

Choose one answer

Scenario questions

The invoice that doubled

A nightly job sends the same 20,000-token policy and a new contract each time. Cache reads were healthy. This week the job appends the contract id to the system prompt so logs are easier to read. cache_read_input_tokens is now zero, cache writes equal the whole prefix, and the bill jumped. Output length is unchanged.

What restored the old cost?

Choose one answer

Build exercise

Price a prefix and a hit

Intermediate · 35 minutes

What you will learn

  • How to turn usage buckets into a cost.
  • Where a checkpoint sits.
  • Which edit forces a new write.
  1. Step 1

    Write two usage lines

    Invent a first call that writes 8,000 cache tokens and generates 500 output tokens, and a second call that reads those 8,000 and generates 500 more, with a 200-token uncached tail.

    Why: The hit is visible only on the second line.

    You should see: Two usage sketches with the four bucket names.

  2. Step 2

    Apply a card

    Pick Sonnet 5.5's $2 and $10 per million, a 1.25× write, and a 0.1× read. Compute both calls and the pair.

    Why: Blended rates hide the expensive bucket.

    You should see: A dollar figure per call, with output larger than the cached tail.

  3. Step 3

    Place the checkpoint

    Order tools, policy, three examples, and the contract text. Mark the breakpoint after the examples.

    Why: The contract changes. The examples do not.

    You should see: A four-part order with the marker before the contract.

  4. Step 4

    Break it on purpose

    Name two edits that would force call two to write again: a top-level effort change, and appending an id to the policy.

    Why: Most production misses are prefix edits.

    You should see: Two edits, each mapped to a usage change.

Review checklist

Checks are saved in this browser.

Key takeaways

  • The bill is usage times the price card: uncached input, cache write, cache read, and output.
  • A cache checkpoint marks the end of the shared prefix. The next call reads it only when those bytes match.
  • Automatic caching follows a growing transcript. An explicit breakpoint is how you stop the marker before a placeholder or a varying field.
  • The first call writes. Effort, thinking, model, and fast mode changes start a new prefix.

Sources

  • Prompt caching — Automatic caching, explicit breakpoints, TTL, and what invalidates a prefix.
  • Models overview — Input and output prices, and the model-specific cache-read rates.
  • Build with Claude — Usage on the Messages response and the request fields that affect it.
  • CCDV-F blueprint notes — Domain 5 cost skill and weight. Study notes, not exam items.

Domain 5 overview · Quick reference

View progress