5.2 · Lesson 2 of 5
Identify risks, limitations, and failure modes of LLM systems
What You Need to Know
Identifying risks for LLM systems means writing them down: hallucination and grounding failures, quality drift, tool misuse, outages, injection, and policy gaps — each with user impact, detection, mitigation, and residual risk. “Claude will handle it” is not a risk posture. Architects treat model and tool limitations (non-determinism, soft prompt adherence, brittle tools) as first-class design constraints.
After mitigations, leftover risk must be documented and accepted by named owners — not buried or assigned vaguely to a vendor. Silent failures matter as much as crashes: a refund agent that slowly drifts on edge policies harms customers without a red dashboard. Tie high-severity modes to monitoring and eval fixtures so the register is operational, not ceremonial.
Failure-mode categories to cover
- Hallucination / grounding failure — invented facts or citations
- Tool misuse — wrong args, unauthorized actions, partial side effects
- Outages and dependency failure — model, tools, retrieval, auth
- Drift and policy gaps — silent quality or coverage holes
- Adversarial input — jailbreaks, injection, social engineering
Decision rules
- Enumerate failure modes with user impact — no happy-path-only risk docs.
- State model and tool limitations explicitly as design constraints.
- Pair each risk with a mitigation and a residual-risk owner.
- Include silent failures (drift, partial tool success, policy gaps).
- Reject “the model will handle it” as architecture.
Residual risk
Perfect elimination is rare. Professional designs show what remains after guardrails, HITL, and monitoring — and who accepts it. Hiding residual risk from stakeholders or deleting it from ADRs fails both exams and audits. Owners review residual risk on a cadence tied to incident and eval evidence.
Review checklist
- Are major harm classes listed with impact and detection?
- Are limitations named (not implied)?
- Does every high-severity mode have a mitigation?
- Is residual risk owned and scheduled for review?
- Do monitoring and evals cover the top modes?
Exam stems often tempt you with optimism or vendor blame. Prefer answers that enumerate modes, assign owners, and refuse to ship without residual-risk acceptance.
Exam application
Correct options list concrete modes (hallucinated balances, tool misuse, outages, policy gaps) and residual ownership. Distractors: happy-path screenshots only, CSS-only audits, “LLMs do not fail,” or residual risk with no product owner.
Exam traps
“The model will handle it”
Optimistic answers without enumerated modes, mitigations, and residual risk fail exam stems.
Happy-path-only risk docs
Demos and screenshots do not surface hallucination, tool misuse, drift, or outages.
Hiding residual risk
After mitigations, leftover risk must be documented and accepted by named owners — not deleted from ADRs.
Ignoring silent failures
Quality drift, partial tool success, and policy gaps harm users without obvious crashes.
Practice scenario
A team designs a Claude refund agent and claims “Claude is careful, so we do not need a failure-mode list.” Stakeholders ask for residual risk. What should the architect deliver?
Build exercise
Build a failure-mode register for a Claude refund agent
35 minutes
What you'll learn
- Enumerate modes with user impact
- Name model/tool limitations as design inputs
- Assign mitigations and residual-risk owners
- Connect top modes to detection and evals
Step 1
Brainstorm failure modes by harm class
For the refund agent, list hallucination (wrong balances), tool misuse (wrong account), outages, retrieval misses, injection, and policy gaps. Attach user impact to each.
Why: Enumeration precedes mitigation. Exam items punish vague “be careful” language.
You should see: A risk register with columns: mode, impact, likelihood, detection signal.
Step 2
Name model and tool limitations explicitly
Record non-determinism, knowledge limits, tool brittleness, and that prompts are soft. Tie each limitation to a design consequence (HITL, allowlists, grounding checks).
Why: Limitations are design inputs, not footnotes. Stakeholders need them visible.
You should see: A short limitations appendix referenced from the ADR.
Step 3
Pair mitigations with residual risk owners
For each mode, name a control (guardrail, HITL, monitoring) and who accepts leftover risk. Escalate undocumented residual risk before ship.
Why: Mitigations without owners leave holes. Residual risk must be accepted, not hidden.
You should see: Risk register updated with mitigation, residual, owner, review date.
Step 4
Wire detection into monitoring and evals
Map each high-severity mode to an alert or eval fixture. Confirm silent modes (drift, policy gaps) have a signal.
Why: A risk list that never becomes detection is theater.
You should see: Three modes with matching canary, alert, or golden-case IDs.
Sources
- Reduce hallucinations — docs.anthropic.com — grounding and failure modes
- Test and evaluate overview — docs.anthropic.com — turning risks into evals
- Tool use overview — docs.anthropic.com — tool misuse as a failure class