4.4 · Lesson 4 of 6
Diagnose system issues (prompt failure, hallucinations, model mismatch)
What You Need to Know
Diagnosing Claude systems means separating prompt failure, hallucination or grounding failure, retrieval miss, and model mismatch before you change configuration. Professional architects reproduce with fixed inputs and traces first. Jumping to a larger model without that classification is one of the most common wrong answers on the exam.
Prompt failure shows as format or instruction misses on tasks the model can otherwise do. Retrieval miss means the needed evidence never entered context. Hallucination or grounding failure means claims are not supported by what was retrieved. Model mismatch shows as consistent capability gaps after retrieval and prompt look fine — not as one noisy chat.
Failure-mode signals
- Cited sources absent from retrieved chunk IDs → provenance / grounding path
- Empty or wrong top-k with correct query intent → retrieval miss
- Correct context, wrong schema or ignored constraints → prompt failure
- Healthy context and instructions, persistent skill gaps → model mismatch
Decision rules
- Reproduce with fixed inputs and traces before changing config.
- Check retrieval provenance before blaming the model for bad citations.
- Model mismatch shows as consistent capability gaps after retrieval looks fine.
- Prompt failure shows as format/instruction misses on otherwise capable tasks.
- Jumping to a larger model without diagnosis is a common wrong answer.
Why classification comes before capacity
Upgrading the model cannot invent documents that were never retrieved. Lowering temperature cannot fix a broken filter. Architects treat diagnosis as a decision tree: pin the failure, name the class, apply one fix, re-measure. Temperature and model tier are late levers, not first guesses.
Review checklist
- Can you replay the failure with identical inputs and a saved trace?
- Do cited claims map to retrieved chunk IDs?
- Is the miss format/instruction, missing evidence, unsupported claim, or capability?
- Have you held other variables fixed while testing one fix class?
- Would a model upgrade still be justified if retrieval were perfect?
Exam stems that show invented citations with unchanged uptime usually key on provenance and grounding — not on swapping models or blaming temperature alone.
Exam application
Prefer reproduce → classify → fix the named layer. Reject always-upgrade, temperature-first, ignore-retrieval, and skip-reproduction distractors.
Exam traps
Always upgrade the model first
Larger models do not fix missing chunks or broken citations. Exam items reward classification of failure mode before capacity spend.
Blame temperature first
Temperature can add variance, but consistent invented citations with empty retrieval traces are grounding or retrieval issues — not sampling noise.
Ignore retrieval provenance
If the answer cites docs that never entered context, the pipeline failed before generation. Check chunk IDs and scores before rewriting prompts.
Skip reproduction with fixed inputs
One-off chats cannot separate prompt failure from hallucination or model mismatch. Architects pin inputs and traces before changing config.
Practice scenario
A RAG support assistant starts citing policy sections that never appear in retrieved chunks. Latency and HTTP success rates are unchanged. Stakeholders want an immediate model upgrade. What should the architect do first?
Build exercise
Diagnose a failing golden set with a decision tree
40 minutes
What you'll learn
- Reproduce failures with fixed inputs and traces
- Classify prompt vs retrieval vs hallucination vs model mismatch
- Verify provenance before escalating model tier
- Apply one fix class and re-measure the golden set
Step 1
Reproduce failures on a fixed golden set with traces
Collect failing cases into pinned inputs. Capture prompt, retrieved chunks (IDs/scores), tool results, and model output for each. Do not change config yet.
Why: Diagnosis without reproduction confounds environment drift with true failure modes.
You should see: A folder of fixtures plus one trace dump per failure labeled with request_id.
Step 2
Classify: prompt vs retrieval vs hallucination/grounding vs model
Ask whether instructions/format were missed, whether needed evidence was retrieved, whether claims lack support in context, or whether the model consistently lacks the capability after context looks fine.
Why: Each class implies a different fix. Mixing them leads to wrong upgrades and wasted eval churn.
You should see: A decision-tree worksheet with one primary class per golden failure.
Step 3
Verify provenance before any model swap
For citation and factual failures, confirm chunk presence and scores. Fix indexing, filters, or grounding checks before escalating model tier.
Why: Model mismatch is diagnosed only after retrieval and prompt look healthy. Skipping provenance is a classic wrong answer.
You should see: Before/after tables showing cited claim → supporting chunk ID or "not retrieved."
Step 4
Apply one fix class and re-run the golden set
Change only the classified layer (prompt, retriever, grounding gate, or model). Re-measure the same fixtures. Document rejected "just upgrade" options.
Why: Architects close the loop with evidence. Multi-change releases hide which diagnosis was right.
You should see: Eval delta: pass rate by failure class; ADR note of rejected model upgrade if provenance was the root cause.
Sources
- Evaluate prompts and outputs — docs.anthropic.com — fixed evals for diagnosis
- Prompt engineering overview — docs.anthropic.com — prompt vs model failure signals
- Models overview — docs.anthropic.com — when mismatch justifies a tier change