7.3 · Lesson 3 of 3
Support debugging and operational issue resolution
What You Need to Know
Supporting debugging and operational issue resolution for Claude-powered developer and agent systems means leading with evidence. Reproduce failures with logs, traces, and minimal fixtures. Separate product bugs from prompt, tool, or MCP bugs before rewriting the wrong layer. Capture durable fixes as runbooks so the next on-call does not start from zero.
Guessing prompt text without traces is the signature distractor. Intermittent agent failures after a Claude Code or MCP rollout are classic Domain 7 stems: the right move is instrumentation and isolation, not thrash. Observability here is operational enablement — request IDs, tool results, and redacted payloads that make classification possible.
Debug loop
- Identify: request IDs, tool/MCP names, errors, versions
- Reproduce: minimal fixtures, not only full prod wait-and-see
- Classify: product vs prompt/skill vs tool/AuthZ vs MCP/config
- Fix the owning layer with the smallest change that clears the fixture
- Record: runbook + optional team skill for the next incident
Decision rules
- Reproduce with logs/traces and minimal fixtures before rewriting prompts.
- Separate product bugs from prompt, tool, and MCP bugs.
- Assign owning lane before broad code or schema changes.
- Capture fixes as reusable runbooks with verify steps.
- Reject guessing without traces as an ops strategy.
Product vs prompt/tool/MCP
A product bug is wrong business behavior even when the agent path and tools are healthy. A prompt or skill bug misroutes judgment with working tools. A tool/AuthZ bug fails schemas, permissions, or side effects. An MCP/config bug is connectivity, secrets, or host wiring. Architects who skip this split burn the wrong owners and miss the exam’s isolation requirement.
Review checklist
- Are request/tool traces available for the failing path?
- Is there a minimal fixture that reproduces the issue?
- Is the fault classified before a broad rewrite?
- Is a runbook updated with symptoms and verify steps?
- Are secrets redacted in logs and incident scraps?
Stems that praise prompt guessing or deleting observability reward naming the evidence-free debug trap.
Exam application
Prefer answers that reproduce with traces/fixtures, isolate the layer, and write a runbook. Distractors: guess prompts, disable logging, restart unrelated infra only, or treat every failure as a product rewrite.
Exam traps
Guessing without traces
Prompt thrash without request IDs, tool results, or fixtures wastes on-call and hides the real layer. Exam distractors love “just tweak the prompt.”
Collapsing all failures into one bucket
Product bugs, prompt bugs, tool/AuthZ bugs, and MCP connectivity bugs need different fixes. Mis-labeling burns the wrong team.
Fix once, never write it down
Without a runbook, the next on-call repeats the same archaeology. Capture steps, signals, and owners.
Reproducing only in full production traffic
Prefer minimal fixtures that isolate the failing tool call or prompt path — safer and faster than waiting for the next intermittent blast.
Practice scenario
An on-call engineer sees intermittent agent failures after a Claude Code + MCP rollout. The team keeps rewriting the system prompt from gut feel; there are no request IDs, tool traces, or minimal fixtures. What should the architect push for first?
Build exercise
Build an evidence-first debug loop for agent and MCP failures
40 minutes
What you'll learn
- Collect identifying logs and traces safely
- Reproduce with minimal fixtures
- Classify product vs prompt/tool/MCP bugs
- Publish a reusable ops runbook
Step 1
Capture identifying signals
Pull request IDs, timestamps, tool names, MCP server, error codes, and redacted input/output snippets. Confirm logging is on for the failing path.
Why: Without identifiers you cannot reproduce or classify the fault layer.
You should see: An incident scrap with IDs and a redaction note.
Step 2
Reproduce with minimal fixtures
Build the smallest input that triggers the failure locally or in a staging harness. Prefer one tool call or one transcript over full prod replay.
Why: Minimal fixtures turn intermittent pain into a repeatable debug loop.
You should see: A fixture file checked into the debug harness with expected vs actual.
Step 3
Separate product vs prompt/tool/MCP bugs
Classify: wrong product behavior with a correct agent path; bad prompt/skill; tool schema/AuthZ; or MCP connectivity/config. Assign the owning lane before changing code broadly.
Why: Wrong-layer fixes (rewrite CRM for an MCP timeout) fail ops items.
You should see: A one-page fault tree with four buckets and owners.
Step 4
Capture the fix as a reusable runbook
Document symptoms, signals, repro fixture, classification, fix, and verification. Link it from on-call docs and team Claude skills if useful.
Why: Runbooks turn one-off heroics into operational enablement.
You should see: A runbook entry with owner and last-verified date.
Sources
- Claude Code overview — docs.anthropic.com — developer tooling context for ops issues
- Model Context Protocol (MCP) — docs.anthropic.com — tool-server failures and config surfaces
- Tool use overview — docs.anthropic.com — tool errors and debugging surfaces
- Working with messages — docs.anthropic.com — request/response shapes for logging