5.4 · 2.8% of the exam · Topic 4 of 4
Cost and Token Management
The bill is the usage object times the price card. Output and uncached input dominate. Prompt caching stores a shared prefix at a cache checkpoint so later calls read it at a fraction of the input price. The checkpoint only hits when that prefix, including thinking and effort, stays the same.
Learning objectives
- Read input, output, cache write, and cache read from a usage object.
- Sketch a call's cost from those counts and the model's prices.
- Place a cache checkpoint at the end of the stable prefix.
- Name the edits that move the checkpoint and drop the read.
Detailed theory
What this skill covers
Cost and token management is 2.8% of the exam: usage tracking, cost modeling, prompt caching, and cache checkpoints.
The previous topics decide how many tokens a design will generate. This topic is how those tokens show up in usage and what a second call can avoid paying for again.
What usage counts
A successful message reports usage. input_tokens are the uncached input billed at the full input price. output_tokens are the generated tokens, including thinking, billed at the output price. cache_creation_input_tokens are the prefix written into the cache. cache_read_input_tokens are the prefix served from a hit.
A word count will not reconcile an invoice. The token counting endpoint can estimate input before you spend the call. It does not know the output you have not generated, and it does not replace the usage on the response you already received.
A conversation resends the history. Turn ten pays input on everything you failed to cache, plus output for the new answer. The expensive pattern is a long thread with no cache and a large output. The report you want per route is calls, input, output, cache writes, and cache reads, summed to a dollar figure with the price card for that model id.
A cost sketch
Price cards are dollars per million tokens, and input and output differ. On the current lineup, output is five times the input price: Haiku 4.5 is $1 and $5, Sonnet 5.5 is $2 and $10, Opus 5.5 is $4 and $20. Multiply each usage bucket by its rate. Do not multiply a single blended rate by the sum of input and output.
Standard cache writes cost 1.25 times the base input price. Standard cache reads cost 0.1 times the base input price. The models overview notes cheaper reads on some current models: 5% of base input on Opus 5.5, and 2.5% on Claude Fable 5.1 and Claude Mythos 5.1. Use the card for the id you pin.
The Message Batches API is half price for work that can wait and does not need a multi-turn tool loop inside the batch. Fast mode uses its own premium card, so a model priced at $4 and $20 is not the fast-mode price. A cost model that forgets the mode will underbill the interactive route.
cost = input_tokens * input_rate
+ cache_creation_input_tokens * write_rate
+ cache_read_input_tokens * read_rate
+ output_tokens * output_ratePrompt caching and the checkpoint
A cache hit resumes a prefix. Anything before the checkpoint must match the bytes that were written, in order. Stable material goes first: tools, the system prompt, policies, and examples. The user text that changes every call goes after the checkpoint.
Automatic caching sets cache_control to type ephemeral at the top level of the request. The service places the breakpoint on the last cacheable block and moves it forward as a conversation grows. That is the right default for a multi-turn thread whose history should be reused.
An explicit checkpoint puts cache_control on one content block. You may place up to four. Use them when the shared prefix ends before the last block, when you want a different lifetime on different blocks, or when you are pre-warming. A pre-warm must put the breakpoint on the last shared block, usually the system prompt or the tools. If automatic caching lands the breakpoint on a placeholder user message, the real user message is a different prefix and the read misses.
A prefix under the model's minimum length does not cache. Creation tokens stay at zero and the call still succeeds. Published minima include 1,024 tokens for Sonnet and 4,096 for Opus and for Haiku 4.5. A short system prompt will not show a write no matter how correctly you spelled cache_control.
What drops the read
The default lifetime is five minutes, and a hit refreshes it. A one-hour lifetime is available at a higher write price. A quiet route that calls every hour will miss a five-minute entry and pay to write again. A noisy route keeps the five-minute entry warm for free on each hit.
These changes start a new prefix: editing the system text or the tools, reordering blocks, switching model, switching fast mode, changing the thinking configuration, and changing top-level effort. Setting a parameter explicitly to its default is the same as omitting it and does not, by itself, break the prefix. On models that support per-message effort, an effort change carried in a system turn inside messages can leave the cached prefix intact. A top-level effort change does not.
Thinking blocks that you pass back unchanged stay compatible with the cache. Dropping or editing them changes the prefix from that point on, and it is also the 400 described in debugging. Tool results become part of the next cached prefix once you send them. That is useful, and it means a different tool result is a different suffix.
Using the numbers
Track usage in production per route and per model id. Alert on output tokens and on a cache read rate that fell, not only on the dollar total. A read rate that falls to zero after a prompt edit is a checkpoint bug. A dollar spike with a healthy read rate is usually more output or more calls.
Model the steady state, not the first call. The first call writes. Later calls read, if the prefix held. A demo that runs once will look like caching failed, because a write is supposed to happen once. Compare the second call's cache_read_input_tokens to the prefix length.
Core concepts
Usage object
- What
- The token buckets on a response: uncached input, output, cache creation, and cache read.
- Why
- The invoice is these counts times the price card.
- When
- You are explaining a bill or checking that a cache hit happened.
- When not
- You are estimating a prompt you have not sent. Use the token counting endpoint, and still expect output to be unknown.
Cache checkpoint
- What
- The cache_control breakpoint at the end of the prefix you want reused.
- Why
- Only the bytes up to that point can hit. A later edit of those bytes misses.
- When
- Tools, system text, or history repeat across calls.
- When not
- The shared text is below the model's minimum cache length. The marker will not write.
Automatic caching
- What
- A top-level ephemeral cache_control. The breakpoint rides the last cacheable block and moves forward.
- Why
- A growing conversation can reuse the prefix without you moving a marker.
- When
- The whole history so far is the shared prefix.
- When not
- You are pre-warming a system prompt and the last block is a dummy user message. Use an explicit breakpoint before that dummy.
Cache read
- What
- Tokens served from a matching prefix, billed at a fraction of the input price.
- Why
- This is the saving. A write on every call is not a saving.
- When
- cache_read_input_tokens is non-zero on the second call.
- When not
- The first call, which should show creation tokens instead.
Price card
- What
- The input, output, cache write, cache read, batch, and fast-mode rates for one model id.
- Why
- Families differ, and fast mode does not use the standard card.
- When
- You turn usage counts into money.
- When not
- You compare quality. A cheaper card does not describe the eval.
Practical examples
The write that never became a read
Call one shows cache_creation_input_tokens and no reads. Call two, a minute later with the same tools and system prompt, shows another write and zero reads. The only edit was output_config.effort from medium to high.
Top-level effort is part of the prefix. The second call wrote a new entry. Put effort back, or keep one effort for the life of the cached conversation. Then the second usage line should show cache reads.
The short policy
A 400-token system prompt carries cache_control. Usage shows input tokens and zeros in both cache fields. The call succeeds.
The prefix is under the model minimum, so nothing is written. Lengthen the stable prefix with the material you were already sending, such as the tools and the examples, until it clears the minimum. A longer dummy sentence that you do not need is the wrong way to get there.
Claude-specific considerations
- Read usage on the response. The four buckets are uncached input, output, cache creation, and cache read.
- Output is priced higher than input. Thinking is output.
- Standard cache writes are 1.25 times base input. Standard cache reads are 0.1 times base input, with a lower read rate on some current model ids.
- Automatic caching is top-level cache_control ephemeral. Explicit breakpoints go on blocks, up to four.
- Minima include 1,024 tokens for Sonnet and 4,096 for Opus and Haiku 4.5. Below that, creation stays zero.
- The default TTL is five minutes and a hit refreshes it. A one-hour TTL costs more to write.
- A top-level effort change, a thinking-config change, a model change, and a fast-mode change all start a new prefix.
- Batch pricing is half, for work that can wait. Fast mode has its own card.
Architecture decisions
Tradeoffs
A cache write costs more than a normal input token and pays off on the next reads. A longer TTL costs more to write and covers a quiet route. A checkpoint that is too early leaves savings on the table. A checkpoint that includes a changing token never hits.
Quick reference
- Cost uses four buckets: uncached input, cache write, cache read, and output.
- Output tokens, including thinking, use the output price. Do not blend them with input.
- Current sticker prices: Haiku 4.5 $1/$5, Sonnet 5.5 $2/$10, Opus 5.5 $4/$20 per million input and output tokens.
- Standard cache write is 1.25× input. Standard cache read is 0.1× input. Check the card for model-specific read rates.
- The checkpoint is the end of the shared prefix. Stable tools and system text first. The changing user text after.
- Automatic caching moves the breakpoint forward. Explicit breakpoints: up to four, on the blocks you name.
- Under the minimum length, the call succeeds and the cache fields stay zero.
- Five-minute TTL by default, refreshed on a hit. A one-hour TTL costs more to write.
- Changing model, fast mode, thinking config, or top-level effort misses the old prefix.
- The first call writes. Judge the hit on the second call's cache_read_input_tokens.
Decision rules for the exam
Common exam traps
Exam tips
- If the stem shows a usage object, do the arithmetic with separate rates.
- If the second call writes again, something in the prefix changed, or the TTL elapsed, or the text never cleared the minimum.
- A checkpoint question is about order: shared bytes first.
Common mistakes
Putting the account id in the system prompt you cache.
Keep the system prompt identical. Put the account id in the user turn after the checkpoint.
Declaring caching broken because the first response has creation tokens and no reads.
Send the same prefix again and read cache_read_input_tokens.
Changing effort between turns of a cached Opus chat.
Keep top-level effort stable, or use a per-message effort change on a model that preserves the prefix.
Modeling fast-mode traffic at the standard Opus price.
Use the fast-mode card for any call that sets speed fast.
Practice questions
Original questions for this topic. They are study items, not questions from the live exam.
A response reports 200 cache_read_input_tokens, 50 input_tokens, 0 cache_creation_input_tokens, and 80 output_tokens. Which bucket uses the output price?
The system prompt and tool list are identical on two calls a minute apart. The second call shows a fresh cache write and zero reads. The only request change is output_config.effort from medium to high. Why did the read miss?
A 300-token system prompt on Sonnet includes cache_control and the call succeeds. Both cache fields are zero. What happened?
You want the tools and the policy cached before traffic arrives. The warm-up request ends with a placeholder user message of warmup. Where does the breakpoint go?
Scenario questions
The invoice that doubled
A nightly job sends the same 20,000-token policy and a new contract each time. Cache reads were healthy. This week the job appends the contract id to the system prompt so logs are easier to read. cache_read_input_tokens is now zero, cache writes equal the whole prefix, and the bill jumped. Output length is unchanged.
What restored the old cost?
Build exercise
Price a prefix and a hit
Intermediate · 35 minutes
What you will learn
- How to turn usage buckets into a cost.
- Where a checkpoint sits.
- Which edit forces a new write.
Step 1
Write two usage lines
Invent a first call that writes 8,000 cache tokens and generates 500 output tokens, and a second call that reads those 8,000 and generates 500 more, with a 200-token uncached tail.
Why: The hit is visible only on the second line.
You should see: Two usage sketches with the four bucket names.
Step 2
Apply a card
Pick Sonnet 5.5's $2 and $10 per million, a 1.25× write, and a 0.1× read. Compute both calls and the pair.
Why: Blended rates hide the expensive bucket.
You should see: A dollar figure per call, with output larger than the cached tail.
Step 3
Place the checkpoint
Order tools, policy, three examples, and the contract text. Mark the breakpoint after the examples.
Why: The contract changes. The examples do not.
You should see: A four-part order with the marker before the contract.
Step 4
Break it on purpose
Name two edits that would force call two to write again: a top-level effort change, and appending an id to the policy.
Why: Most production misses are prefix edits.
You should see: Two edits, each mapped to a usage change.
Review checklist
Checks are saved in this browser.
Key takeaways
- The bill is usage times the price card: uncached input, cache write, cache read, and output.
- A cache checkpoint marks the end of the shared prefix. The next call reads it only when those bytes match.
- Automatic caching follows a growing transcript. An explicit breakpoint is how you stop the marker before a placeholder or a varying field.
- The first call writes. Effort, thinking, model, and fast mode changes start a new prefix.
Sources
- Prompt caching — Automatic caching, explicit breakpoints, TTL, and what invalidates a prefix.
- Models overview — Input and output prices, and the model-specific cache-read rates.
- Build with Claude — Usage on the Messages response and the request fields that affect it.
- CCDV-F blueprint notes — Domain 5 cost skill and weight. Study notes, not exam items.