CCAR-P · Study Guide

← Domain 4: Evaluation, Testing & Optimization

4.3 · Lesson 3 of 6

Conduct A/B testing and iterative improvements

What You Need to Know

A/B testing and iterative improvement turn Claude config changes into evidence, not opinions. Professional architects isolate the primary factor when possible — prompt, retrieval, or model settings — power the experiment with enough traffic, and ship only when quality, cost, and other SLAs clear agreed thresholds. Confounded releases that change everything at once cannot attribute a win to a single tweak.

Iterate with a fixed golden set plus online metrics. Anecdotes and demo vibes are not ship criteria. Watch for confounders (seasonality, traffic mix, simultaneous feature launches) and underpowered samples that look significant by chance. Attribute carefully in write-ups: what changed, what was held fixed, and which gates passed.

Ingredients of a controlled iteration

  • One primary factor (or an explicit multi-factor design)
  • Powered traffic split and a pre-set runtime window
  • Ship gates for quality, cost, and latency/SLA ceilings
  • Golden-set regression plus online outcome metrics
  • Clear attribution — no rewriting history after a multi-change release

Decision rules

  • Change one primary factor at a time when possible (or design a proper multi-factor test).
  • Power and traffic matter for significance — anecdotes are not enough.
  • Ship only when quality, cost, and other SLAs clear agreed thresholds.
  • Attribute wins carefully; confounded releases that change everything fail exam items.
  • Iterate with a fixed golden set plus online metrics, not vibes alone.

Why isolation beats kitchen-sink releases

Shipping a new model, prompt, and retriever together may improve metrics while teaching you nothing about which lever worked — or which one will break next time. On the exam, attributing a multi-change win to \"the prompt\" is a classic wrong answer. Architects either isolate or declare a factorial design; they do not invent causation after the fact.

Review checklist

  • Is one primary factor isolated, or is multi-factor design explicit?
  • Is the experiment powered with enough traffic and a fixed window?
  • Are quality, cost, and SLA ship thresholds pre-registered?
  • Are golden-set and online metrics both in the decision?
  • Would attributing a kitchen-sink release to one tweak fail review?

Exam distractors push vibes, underpowered anecdotes, or confounded multi-change attribution. The professional move is a controlled plan with isolation, power, and multi-axis ship criteria.

Exam application

Anecdotal preference and confounded multi-change releases are wrong; controlled tests with ship criteria win. Distractors: ship on demo vibes, change everything and credit one factor, ignore cost, or declare significance on tiny traffic. Prefer isolation + powered A/B + golden set and online gates.

Exam traps

  • Shipping on vibes or demo applause

    Anecdotal preference is not an A/B result. Exam items reward golden-set plus online metrics with thresholds.

  • Changing everything in one release

    Confounded multi-factor ships block learning. Isolate the primary factor or design a real multi-factor test.

  • Ignoring cost and SLA thresholds

    Quality wins that breach cost or latency ceilings should not ship. Ship criteria cover quality, cost, and SLAs together.

  • Underpowered traffic and calling it significant

    Tiny samples and anecdotes produce false wins. Power and traffic matter for trustworthy decisions.

Practice scenario

A team changes the model, the system prompt, and the retrieval config in one release. Online metrics improve. They attribute the win entirely to the new prompt. What is wrong with this conclusion?

Choose one answer

Build exercise

Design an A/B plan for a prompt change with ship criteria and isolation

35 minutes

What you'll learn

  • Isolate the prompt factor and state a hypothesis
  • Power traffic and set a fixed experiment window
  • Pre-register quality, cost, and SLA ship gates
  • Close the loop with golden set plus online metrics
  1. Step 1

    Isolate the primary factor and write a hypothesis

    For a prompt-only change, hold model and retrieval fixed. State the hypothesis (e.g. clearer refuse rules reduce policy errors without raising latency) and the metrics that would confirm or reject it.

    Why: Isolation is how architects learn. Confounded releases fail exam items on attribution.

    You should see: A one-page experiment brief: factor under test, controls held fixed, hypothesis, metrics.

  2. Step 2

    Power the experiment and split traffic

    Estimate sample size for the minimum detectable effect on your primary metric. Assign users or sessions randomly to control vs treatment with enough traffic and a fixed runtime window.

    Why: Underpowered tests produce vibes dressed as science. Traffic and power are part of architect competence.

    You should see: A traffic plan: allocation %, expected N, primary metric, MDE, stop date.

  3. Step 3

    Define ship criteria across quality, cost, and SLAs

    Write pass/fail thresholds before results land: quality must clear the floor, cost and latency must stay under ceilings, safety/security must not regress. No ship if any critical gate fails.

    Why: Iterate toward production only when agreed bars clear — not when one chart moves.

    You should see: A pre-registered ship checklist with numeric thresholds for quality, cost, and SLAs.

  4. Step 4

    Iterate with golden set plus online metrics

    After the A/B, keep the fixed golden set for regression and watch online metrics for drift. Attribute carefully in the write-up: what changed, what was held fixed, what shipped and why.

    Why: Continuous improvement needs both offline anchors and online truth — not vibes alone.

    You should see: Post-experiment report with isolation statement, results vs ship criteria, and next iteration.

Sources