4.6 · Lesson 6 of 6
Monitor system performance using logging and observability tools
What You Need to Know
Monitoring LLM system performance means operating a logging and observability loop that catches quality and cost regressions — not only HTTP uptime. Alert on quality and cost, keep enough redacted prompt and tool context to debug, join request IDs to outcomes, and close every serious alert into a runbook and an eval fixture.
Domain 3 covers sampling and SLIs at scale; this lesson focuses on the operational loop: detect silent quality or cost drift, debug with safe context, and feed production misses back into evaluation. Monitoring that only tracks HTTP 500s misses the failure modes the exam cares about.
Ops loop essentials
- Quality and cost alerts with owners, not uptime alone
- Redacted prompt/tool/retrieval context on alerted requests
- Alert → runbook → eval fixture capture
- request_id joined to thumbs-down, reopen, and eval labels
Decision rules
- Alert on quality and cost, not only errors or uptime.
- Keep enough prompt/tool context (redacted) to debug regressions.
- Close the loop from alert to runbook to eval fixture.
- Monitoring that only tracks HTTP 500s misses LLM failure modes.
- Join request IDs to outcomes so silent quality dips are explainable.
Why the eval loop matters
Production will invent failure shapes your offline set never saw. If alerts do not deposit fixtures, the golden set stagnates and the same class of miss returns. Architects treat observability as an input to evaluation — not a separate ops hobby.
Review checklist
- Would a week-long quality drop page someone while uptime stays green?
- Can on-call pull redacted context by request_id within minutes?
- Does each quality/cost alert link to a runbook?
- Is there a path from confirmed incidents into the eval suite?
- Are outcomes joinable to requests for silent-dip analysis?
Prefer answers that add quality/cost monitors tied to evals. Reject HTTP-500-only, no debug context, alert-without-runbook, and no-outcome-join distractors.
Exam application
Green uptime + falling quality → missing quality/cost alerts and the alert→eval fixture loop. Uptime-only monitoring is incomplete.
Exam traps
HTTP-500-only monitoring
LLM systems fail with 200s and wrong answers. Uptime without quality and cost SLIs misses the exam's operational loop.
Alerts without debug context
A page that cannot show redacted prompt, tool, and retrieval context cannot diagnose regressions quickly.
Alerts without runbooks or eval fixtures
Firing without a next step leaves quality debt. Close the loop: alert → runbook → capture fixture into the golden set.
No join from request_id to outcomes
Without linking requests to thumbs-down, reopen, or eval labels, silent quality dips stay unexplained.
Practice scenario
An LLM support bot shows green HTTP uptime for a week while golden-set quality and thumbs-down rates quietly drop. On-call has no redacted prompt or tool context and no path from alerts into the eval suite. What is the primary monitoring gap?
Build exercise
Design monitoring for an LLM support bot
40 minutes
What you'll learn
- Define quality and cost SLIs beyond uptime
- Retain redacted prompt/tool context for debug
- Wire alert → runbook → eval fixture capture
- Join request_id to outcomes for silent dips
Step 1
Define quality and cost SLIs beside availability
Add canary or golden pass rate, cost per resolution, p95 latency, tool-error rate, and user outcome signals. Set alert thresholds with owners.
Why: Operational monitoring for LLM performance is not only "is the process up."
You should see: An SLI table: metric, source, threshold, page vs ticket, owner.
Step 2
Retain redacted prompt and tool context for debug
On alerted or sampled failures, keep enough redacted prompt, tool I/O, and retrieval IDs to reproduce. Strip secrets, tokens, and regulated PII before export.
Why: Quality alerts without context force guesswork. Redaction makes debug storage acceptable.
You should see: A debug payload schema with redaction rules checked into the logging config.
Step 3
Wire alert → runbook → eval fixture capture
Each quality or cost alert links to a runbook step that triages class of failure and, when confirmed, adds a minimized fixture to the eval suite.
Why: Monitoring that never feeds evals repeats the same production miss. Architects close the loop.
You should see: Runbook links on alert rules plus a fixture PR or upload path from a real incident.
Step 4
Join request_id to outcomes and review weekly
Propagate request_id through logs, traces, feedback, and tickets. Review unexplained quality dips in a recurring ops ritual with eval owners.
Why: Outcome join turns silent regressions into explainable incidents instead of anecdotes.
You should see: A dashboard joining request_id → outcome label → eval fixture status.
Sources
- Streaming and operational response patterns — docs.anthropic.com — runtime behavior for ops context
- Evaluate prompts and outputs — docs.anthropic.com — closing alert→fixture loop
- Working with messages — docs.anthropic.com — request/response shapes for logging