CCAR-P · Study Guide

← Domain 4: Evaluation, Testing & Optimization

4.6 · Lesson 6 of 6

Monitor system performance using logging and observability tools

What You Need to Know

Monitoring LLM system performance means operating a logging and observability loop that catches quality and cost regressions — not only HTTP uptime. Alert on quality and cost, keep enough redacted prompt and tool context to debug, join request IDs to outcomes, and close every serious alert into a runbook and an eval fixture.

Domain 3 covers sampling and SLIs at scale; this lesson focuses on the operational loop: detect silent quality or cost drift, debug with safe context, and feed production misses back into evaluation. Monitoring that only tracks HTTP 500s misses the failure modes the exam cares about.

Ops loop essentials

  • Quality and cost alerts with owners, not uptime alone
  • Redacted prompt/tool/retrieval context on alerted requests
  • Alert → runbook → eval fixture capture
  • request_id joined to thumbs-down, reopen, and eval labels

Decision rules

  • Alert on quality and cost, not only errors or uptime.
  • Keep enough prompt/tool context (redacted) to debug regressions.
  • Close the loop from alert to runbook to eval fixture.
  • Monitoring that only tracks HTTP 500s misses LLM failure modes.
  • Join request IDs to outcomes so silent quality dips are explainable.

Why the eval loop matters

Production will invent failure shapes your offline set never saw. If alerts do not deposit fixtures, the golden set stagnates and the same class of miss returns. Architects treat observability as an input to evaluation — not a separate ops hobby.

Review checklist

  • Would a week-long quality drop page someone while uptime stays green?
  • Can on-call pull redacted context by request_id within minutes?
  • Does each quality/cost alert link to a runbook?
  • Is there a path from confirmed incidents into the eval suite?
  • Are outcomes joinable to requests for silent-dip analysis?

Prefer answers that add quality/cost monitors tied to evals. Reject HTTP-500-only, no debug context, alert-without-runbook, and no-outcome-join distractors.

Exam application

Green uptime + falling quality → missing quality/cost alerts and the alert→eval fixture loop. Uptime-only monitoring is incomplete.

Exam traps

  • HTTP-500-only monitoring

    LLM systems fail with 200s and wrong answers. Uptime without quality and cost SLIs misses the exam's operational loop.

  • Alerts without debug context

    A page that cannot show redacted prompt, tool, and retrieval context cannot diagnose regressions quickly.

  • Alerts without runbooks or eval fixtures

    Firing without a next step leaves quality debt. Close the loop: alert → runbook → capture fixture into the golden set.

  • No join from request_id to outcomes

    Without linking requests to thumbs-down, reopen, or eval labels, silent quality dips stay unexplained.

Practice scenario

An LLM support bot shows green HTTP uptime for a week while golden-set quality and thumbs-down rates quietly drop. On-call has no redacted prompt or tool context and no path from alerts into the eval suite. What is the primary monitoring gap?

Choose one answer

Build exercise

Design monitoring for an LLM support bot

40 minutes

What you'll learn

  • Define quality and cost SLIs beyond uptime
  • Retain redacted prompt/tool context for debug
  • Wire alert → runbook → eval fixture capture
  • Join request_id to outcomes for silent dips
  1. Step 1

    Define quality and cost SLIs beside availability

    Add canary or golden pass rate, cost per resolution, p95 latency, tool-error rate, and user outcome signals. Set alert thresholds with owners.

    Why: Operational monitoring for LLM performance is not only "is the process up."

    You should see: An SLI table: metric, source, threshold, page vs ticket, owner.

  2. Step 2

    Retain redacted prompt and tool context for debug

    On alerted or sampled failures, keep enough redacted prompt, tool I/O, and retrieval IDs to reproduce. Strip secrets, tokens, and regulated PII before export.

    Why: Quality alerts without context force guesswork. Redaction makes debug storage acceptable.

    You should see: A debug payload schema with redaction rules checked into the logging config.

  3. Step 3

    Wire alert → runbook → eval fixture capture

    Each quality or cost alert links to a runbook step that triages class of failure and, when confirmed, adds a minimized fixture to the eval suite.

    Why: Monitoring that never feeds evals repeats the same production miss. Architects close the loop.

    You should see: Runbook links on alert rules plus a fixture PR or upload path from a real incident.

  4. Step 4

    Join request_id to outcomes and review weekly

    Propagate request_id through logs, traces, feedback, and tickets. Review unexplained quality dips in a recurring ops ritual with eval owners.

    Why: Outcome join turns silent regressions into explainable incidents instead of anecdotes.

    You should see: A dashboard joining request_id → outcome label → eval fixture status.

Sources