2.3 · 6.8% of the exam · Topic 3 of 6
Claude API Mechanics
The Messages API is stateless: the client sends the history every time. Streaming changes when tokens arrive, not what they cost. Batch is for work nobody is waiting on and that does not need a multi-turn tool loop. Prompt caching pays when the stable prefix comes first.
Learning objectives
- Explain that the Messages API stores no conversation and that the client resends history.
- Use stop_reason to tell a finished turn from a tool request, and say who executes the tool.
- Separate streaming, which changes perceived latency, from prompt caching, which can cut latency and cost.
- Choose the Message Batches API only when nobody is waiting and the work does not need multi-turn tool calling.
- Place stable content before variable content when caching, and know that the same API is also offered by other vendors.
Detailed theory
One stateless request
Claude API mechanics is the 6.8% skill in this domain. The Messages API does not keep a session. Your application holds the messages and sends the system prompt plus every turn so far on each call. The response is one assistant turn, which you append if you call again.
That is why a long transcript costs more on the next turn than a short one. The server did not remember the cheap prefix unless you cached it. System content is where durable behavior goes. User and assistant turns are the dialogue.
Tools and stop_reason
A response can ask for a tool. stop_reason tool_use means the application executes the tool and sends the result back as a user turn. stop_reason end_turn means the model is finished. stop_reason max_tokens means the turn was cut off. The model does not run the tool. The history grows because the tool result is another message you must send next time.
Whether the application should be a workflow or an agent is Domain 1. This topic is the request itself: the tools array, the message list, and the stop reason the client branches on.
if (response.stop_reason === "end_turn") return textFrom(response);
if (response.stop_reason === "tool_use") {
messages.push({ role: "assistant", content: response.content });
messages.push({ role: "user", content: runTools(response) });
}Streaming, vision, and thinking
Streaming sends tokens as they are produced. Public notes for this exam say it cuts perceived latency, not the token cost. Use it when a person is watching the reply form.
Vision and other non-text input arrive as content blocks on a user message, alongside text. Thinking is a request control for harder problems. It spends more of the turn on reasoning and more latency. It is not a substitute for sending the right history.
Batch, caching, and vendors
The Message Batches API is for bulk work that can wait. Public notes give the shape: about half the real-time cost, results inside a 24-hour window, results correlated with a custom id, and no multi-turn tool calling. The decision is whether anyone is waiting. If they are, use the Messages API. If nobody is waiting but the job needs a tool loop across turns, the notes still keep it on the real-time API. Cost is the consequence of that choice.
Prompt caching matches a prefix. Put the stable system prompt, policies, and examples first, and the variable user content last. A user message placed at the front changes the prefix every call, so the rest of the prompt cannot be reused. Caching is the mechanic that can cut both latency and cost. Reordering is often the whole fix.
The same model is not only on Anthropic's own API. The published skill includes third-party vendors. The request is still messages, tools, and the vendor's auth and endpoint. Pin the model id you verified, whichever vendor serves it.
Core concepts
Messages API
- What
- A stateless HTTPS request. The client sends system content and the full message list.
- Why
- Nothing on the server reconstructs the dialogue for you.
- When
- Interactive work, or any job that needs another turn after a tool result.
- When not
- A large pile of independent requests that nobody is waiting on. That is a batch candidate.
stop_reason
- What
- The field that says the turn finished, wants a tool, or was truncated.
- Why
- The first content block can be text while a later block is a tool call.
- When
- Every client that loops.
- When not
- You scan the assistant's prose for a sentence that says it is done.
Streaming
- What
- Tokens delivered as they are generated.
- Why
- A person who is watching should not wait for the full body.
- When
- Someone is waiting on the reply.
- When not
- You are trying to reduce the token bill. Streaming does not do that.
Message Batches
- What
- An asynchronous bulk API. Notes: lower cost, up to about a day, custom ids, no multi-turn tool use.
- Why
- Independent overnight work should not pay real-time rates.
- When
- Nobody is waiting, and each item can finish in one request.
- When not
- A user is watching, or the item must call a tool and then continue.
Prompt caching
- What
- Reuse of a stable prefix.
- Why
- Policies and examples that never change should not be reprocessed at full price.
- When
- The prefix is identical across calls and placed first.
- When not
- The user text, the date, or a retrieved document sits at the front of the cached region.
Practical examples
Overnight summaries
One hundred thousand documents, no one waiting, no tool follow-up, a strict budget. The notes' question is whether anyone is waiting. They are not, and the work does not need another turn after a tool. Message Batches fits. Streaming does not, because nobody is watching tokens arrive. A real-time loop would work and would spend more.
A cache that never hits
The same policy and examples are sent on every call, but the client puts the user document first. The prefix changes every time, so the policy is never reused. Move the document after the stable block. The bytes of the policy did not need to change.
Claude-specific considerations
- Tool results go back as a user message. That is why the next request is larger.
- max_tokens truncation is stop_reason max_tokens. Raising the limit or splitting the task is the mechanic. It is not end_turn.
- Image inputs are content blocks. They count toward the same request as the text.
- Vendor SDKs wrap the same ideas: messages, tools, streaming, and a model id. Auth, regions, and the exact model string differ by vendor. Pin the string you tested.
Architecture decisions
Tradeoffs
Real time spends more and answers while someone watches. Batch spends less and answers later, without a multi-turn tool loop. Streaming changes arrival, not price. Caching changes price when the prefix is stable.
Quick reference
- The client stores history. Each Messages call resends it.
- end_turn stops. tool_use means your code runs the tool and sends a user turn back. max_tokens is truncation.
- Streaming lowers perceived latency, not token cost.
- Vision and other media are content blocks on the message.
- Batch when nobody is waiting and the item does not need a multi-turn tool loop.
- Cache the stable prefix first. Variable content last.
Decision rules for the exam
Common exam traps
Exam tips
- This repository's architect mock bank is not the question set for this skill. No practice items are wired here.
- Public notes: batch versus real time turns on whether anyone is waiting, then on whether the work needs another turn after a tool.
- A cache miss with unchanged policy text is an ordering problem.
Common mistakes
Keeping only the latest assistant message.
Resend the history the next decision depends on.
Using batch for a chat window.
A person is waiting. Use the Messages API.
Caching a prefix that includes the user payload.
Stable instructions first, the payload after.
Practice questions
Original questions for this topic. They are study items, not questions from the live exam.
Scenario questions
Build exercise
Classify six calls: real time, stream, batch, or cache
Intermediate · 35 minutes
What you will learn
- Who stores the conversation.
- When batch is allowed.
- What streaming does and does not change.
- How to order a cached prompt.
Step 1
Write six requests
Include a chat reply, a tool follow-up, an overnight corpus with no tools, a UI that shows tokens, a repeated policy plus a new document, and a request that forgot earlier turns.
Why: The skill is the classification, not a single demo call.
You should see: Six one-line requests.
Step 2
Mark real time or batch
For each request, answer whether anyone is waiting and whether a tool result must be sent back for another model turn.
Why: Those are the two notes questions for batch.
You should see: Only the independent overnight work marked batch.
Step 3
Mark streaming and cache order
Turn streaming on only where a person watches tokens. For the repeated policy, write the block order.
Why: Streaming and caching solve different costs.
You should see: Policy, then examples, then the document.
Step 4
Fix the amnesiac client
Show the messages array the client must store after a tool result.
Why: The server will not reconstruct it.
You should see: Assistant tool_use, then a user tool result, ready for the next call.
Review checklist
Checks are saved in this browser.
Key takeaways
- Resend history. The Messages API does not store it.
- The application runs tools. tool_use continues. end_turn stops.
- Stream for a watching user. Batch when nobody waits and the item is one shot.
- Cache a stable prefix. Put variable content after it.
Sources
- CCDV-F exam guide, Domain 2 skill weights — Public blueprint summary. API notes: messages, tools, streaming, vision, thinking, caching, vendors, batch versus real time.
- Messages API — Request and response fields