CCAR-P · Study Guide

← Domain 2: Claude Models, Prompting & Context Engineering

2.1 · Lesson 1 of 5

Select appropriate Claude models based on trade-offs

What You Need to Know

Model selection is a trade-off among capability, latency, unit cost, and risk — not a contest to always pick the highest-capability tier. Match relative tiers (fast/cheap, balanced, highest-capability) to task difficulty and blast radius. High-volume, low-risk paths usually want the lightest model that clears an explicit accuracy floor and latency/cost SLAs. Escalate harder or higher-risk traffic when evals say the lighter tier fails.

Price the decision at expected volume. A prototype that looks fine at tiny traffic can break budget or p95 at production QPS. Treat context-window sizes as scenario constraints when the stem gives them — not as trivia to memorize. Revisit selection when evals, SLAs, volume, or risk mix change. “Always the largest model” is the exam’s favorite wrong answer.

Selection inputs

  • Task difficulty and failure cost (risk class)
  • Accuracy floor from owned evals
  • Latency and unit-cost SLAs at expected volume
  • Traffic mix: easy bulk vs hard escalations
  • Revisit triggers when evidence moves

Decision rules

  • Smallest model that clears accuracy floor + latency/cost SLAs.
  • Match tier to difficulty and risk — not prestige.
  • Price unit cost and latency at expected volume.
  • Escalate hard/high-risk traffic; do not over-provision the bulk path.
  • Revisit when evals, SLAs, volume, or risk mix change.

Relative tiers, not absolute dogma

Prefer relative language — fast/cheap vs balanced vs highest-capability — over brittle absolute prices or model nicknames as forever-facts. When a scenario states a window or latency budget, treat those numbers as constraints for this workload. The decision rule stays the same: clear the floor and SLAs with the lightest tier that works.

Review checklist

  • Are accuracy floor and latency/cost SLAs written?
  • Is volume priced into unit cost?
  • Is the bulk path on the lightest clearing tier?
  • Is there an escalation path for hard/high-risk cases?
  • Are revisit triggers defined?

Stems that praise “always largest for safety” reward naming the over-provisioning trap and returning to floor + SLA fit.

Exam application

Prefer answers that select the smallest sufficient tier and escalate only with evidence. Distractors: always-largest, freeze forever, ignore volume, or run every model on every request.

Exam traps

  • Always pick the largest model

    Capability without latency and unit-cost accounting fails high-volume paths. Exam stems reward the smallest tier that clears floors and SLAs.

  • Ignoring expected volume

    A prototype that looks cheap at 10 req/day can blow the budget at 10k/hour. Price unit cost at the traffic you will actually serve.

  • One-time selection forever

    Evals, SLAs, and mix of hard vs easy tasks change. Revisit when evidence moves — not only when marketing asks.

  • Memorizing absolute prices or window sizes as fixed facts

    Treat tiers relatively (fast/cheap vs balanced vs highest-capability) and treat window sizes as scenario constraints, not trivia to memorize.

Practice scenario

A high-volume FAQ deflector must hold p95 under a tight latency SLA and already clears its accuracy floor on a mid-tier model. Leadership wants the highest-capability model on every request “to be safe.” What is the architect’s best default stance?

Choose one answer

Build exercise

Select model tiers for a high-volume FAQ deflector

35 minutes

What you'll learn

  • Write accuracy floor and latency/cost SLAs
  • Map traffic mix to relative model tiers
  • Price unit cost at expected volume
  • Define revisit triggers when evidence changes
  1. Step 1

    Name the accuracy floor and latency/cost SLAs

    Write the minimum quality bar, p95 latency, and unit-cost ceiling for the FAQ path before comparing model tiers.

    Why: Trade-offs are meaningless without explicit bars. Exam items punish capability-only picks.

    You should see: A one-page SLA card: accuracy floor, p95, cost/request, risk class.

  2. Step 2

    Map workload mix to relative tiers

    Classify traffic as easy/high-volume vs hard/high-risk. Assign relative tiers (fast/cheap, balanced, highest-capability) — not brand names as dogma.

    Why: Most volume often clears a lighter tier; hard cases may need escalation.

    You should see: A tier table: traffic class → tier → why → escalate when.

  3. Step 3

    Estimate unit cost and latency at expected volume

    Multiply projected QPS by per-request tokens and relative unit cost. Compare against the cost ceiling and p95 budget.

    Why: Volume turns small per-call differences into real budget and SLA risks.

    You should see: A rough volume sheet: QPS × tokens × relative cost vs budget.

  4. Step 4

    Define revisit triggers

    Document when selection must be re-checked: eval regression, SLA breach, volume spike, or new high-risk intents.

    Why: Model choice is continuous with evidence — not a one-time prototype decision.

    You should see: Revisit checklist owners and thresholds.

Sources