4.3 · Lesson 3 of 6
Conduct A/B testing and iterative improvements
What You Need to Know
A/B testing and iterative improvement turn Claude config changes into evidence, not opinions. Professional architects isolate the primary factor when possible — prompt, retrieval, or model settings — power the experiment with enough traffic, and ship only when quality, cost, and other SLAs clear agreed thresholds. Confounded releases that change everything at once cannot attribute a win to a single tweak.
Iterate with a fixed golden set plus online metrics. Anecdotes and demo vibes are not ship criteria. Watch for confounders (seasonality, traffic mix, simultaneous feature launches) and underpowered samples that look significant by chance. Attribute carefully in write-ups: what changed, what was held fixed, and which gates passed.
Ingredients of a controlled iteration
- One primary factor (or an explicit multi-factor design)
- Powered traffic split and a pre-set runtime window
- Ship gates for quality, cost, and latency/SLA ceilings
- Golden-set regression plus online outcome metrics
- Clear attribution — no rewriting history after a multi-change release
Decision rules
- Change one primary factor at a time when possible (or design a proper multi-factor test).
- Power and traffic matter for significance — anecdotes are not enough.
- Ship only when quality, cost, and other SLAs clear agreed thresholds.
- Attribute wins carefully; confounded releases that change everything fail exam items.
- Iterate with a fixed golden set plus online metrics, not vibes alone.
Why isolation beats kitchen-sink releases
Shipping a new model, prompt, and retriever together may improve metrics while teaching you nothing about which lever worked — or which one will break next time. On the exam, attributing a multi-change win to \"the prompt\" is a classic wrong answer. Architects either isolate or declare a factorial design; they do not invent causation after the fact.
Review checklist
- Is one primary factor isolated, or is multi-factor design explicit?
- Is the experiment powered with enough traffic and a fixed window?
- Are quality, cost, and SLA ship thresholds pre-registered?
- Are golden-set and online metrics both in the decision?
- Would attributing a kitchen-sink release to one tweak fail review?
Exam distractors push vibes, underpowered anecdotes, or confounded multi-change attribution. The professional move is a controlled plan with isolation, power, and multi-axis ship criteria.
Exam application
Anecdotal preference and confounded multi-change releases are wrong; controlled tests with ship criteria win. Distractors: ship on demo vibes, change everything and credit one factor, ignore cost, or declare significance on tiny traffic. Prefer isolation + powered A/B + golden set and online gates.
Exam traps
Shipping on vibes or demo applause
Anecdotal preference is not an A/B result. Exam items reward golden-set plus online metrics with thresholds.
Changing everything in one release
Confounded multi-factor ships block learning. Isolate the primary factor or design a real multi-factor test.
Ignoring cost and SLA thresholds
Quality wins that breach cost or latency ceilings should not ship. Ship criteria cover quality, cost, and SLAs together.
Underpowered traffic and calling it significant
Tiny samples and anecdotes produce false wins. Power and traffic matter for trustworthy decisions.
Practice scenario
A team changes the model, the system prompt, and the retrieval config in one release. Online metrics improve. They attribute the win entirely to the new prompt. What is wrong with this conclusion?
Build exercise
Design an A/B plan for a prompt change with ship criteria and isolation
35 minutes
What you'll learn
- Isolate the prompt factor and state a hypothesis
- Power traffic and set a fixed experiment window
- Pre-register quality, cost, and SLA ship gates
- Close the loop with golden set plus online metrics
Step 1
Isolate the primary factor and write a hypothesis
For a prompt-only change, hold model and retrieval fixed. State the hypothesis (e.g. clearer refuse rules reduce policy errors without raising latency) and the metrics that would confirm or reject it.
Why: Isolation is how architects learn. Confounded releases fail exam items on attribution.
You should see: A one-page experiment brief: factor under test, controls held fixed, hypothesis, metrics.
Step 2
Power the experiment and split traffic
Estimate sample size for the minimum detectable effect on your primary metric. Assign users or sessions randomly to control vs treatment with enough traffic and a fixed runtime window.
Why: Underpowered tests produce vibes dressed as science. Traffic and power are part of architect competence.
You should see: A traffic plan: allocation %, expected N, primary metric, MDE, stop date.
Step 3
Define ship criteria across quality, cost, and SLAs
Write pass/fail thresholds before results land: quality must clear the floor, cost and latency must stay under ceilings, safety/security must not regress. No ship if any critical gate fails.
Why: Iterate toward production only when agreed bars clear — not when one chart moves.
You should see: A pre-registered ship checklist with numeric thresholds for quality, cost, and SLAs.
Step 4
Iterate with golden set plus online metrics
After the A/B, keep the fixed golden set for regression and watch online metrics for drift. Attribute carefully in the write-up: what changed, what was held fixed, what shipped and why.
Why: Continuous improvement needs both offline anchors and online truth — not vibes alone.
You should see: Post-experiment report with isolation statement, results vs ship criteria, and next iteration.
Sources
- Increase output consistency — docs.anthropic.com — iterate with measurable consistency
- Test and evaluate overview — docs.anthropic.com
- Prompt engineering overview — docs.anthropic.com — isolate prompt changes when testing