CCAR-P · Study Guide

← Domain 5: Governance, Safety & Risk Management

5.1 · Lesson 1 of 5

Implement guardrails and safety controls

What You Need to Know

Guardrails for Claude systems are layered controls: prompt guidance, input and output filters, and deterministic policy on the tool path. When a requirement says the system must always or must never do something — especially irreversible actions like transfers, deletes, or privilege changes — architects put hard blocks in code, hooks, or a policy engine. Softer wording in the system prompt alone does not satisfy that bar.

Place filters where risk lives: jailbreak and injection screening on intake; secret, PII, and disallowed-content checks on outputs; allowlists and authZ binding on tools. Log every trigger with enough redacted context to review and improve rules. Defense in depth means prompt + filter + tool denylist + escalation for residual risk — not a single clever instruction.

Control layers

  • Prompt / policy text — steers behavior; never the only “must always” control
  • Input filters — jailbreaks, injection, oversized or disallowed payloads
  • Output filters — secrets, PII, banned topics, schema violations
  • Tool-path guards — allowlists, arg validation, session-bound authZ, pre-exec hooks
  • Logging and review — trigger rates, false positives, adversarial regressions

Decision rules

  • “Must always / must never” → deterministic control on the tool or output path.
  • Prompts alone do not satisfy hard safety language.
  • Layer prompt + I/O filters + tool denylists + escalation.
  • Block irreversible actions before execution, not after the fact.
  • Log and review guardrail triggers; feed failures into evals.

Why prompt-only fails

Models can be steered, jailbroken, or simply wrong. A refund agent that “knows” not to wire outside policy can still emit a tool call if the only barrier is text. Architects therefore treat prompts as necessary guidance and enforcement as mandatory for high-severity rules. Temperature, few-shot jokes, and longer sermons are distractors.

Review checklist

  • Which actions are irreversible or privacy-sensitive?
  • For each “must always,” is there a non-prompt enforcer?
  • Are inputs and outputs both filtered where needed?
  • Are tool args bound to verified session entitlements?
  • Are triggers logged, reviewed, and covered by adversarial evals?

Exam stems often show a careful system prompt next to a must-always requirement. The professional move is to add a deterministic guard on the path that can cause harm — then keep the prompt as the soft layer, not the only layer.

Exam application

Prefer answers that name hooks, policy checks, allowlists, output filters, or DLP-style blocks. Distractors: longer prompts, higher temperature, hope, or removing all structure so humans copy-paste secrets. Layered defenses beat single soft controls.

Exam traps

  • Prompt-only “must always” rules

    System prompts are soft. Exam stems that say must always / must never reward code-path blocks, hooks, and policy engines.

  • Guarding inputs but not outputs or tools

    Jailbreak filters without output DLP or tool denylists leave exfiltration and irreversible actions open.

  • Silent guardrail triggers

    Blocks without logging and review cannot be tuned, audited, or turned into eval fixtures.

  • Relying on temperature or “be careful” tone

    Sampling knobs and politeness do not enforce policy. Deterministic controls do.

Practice scenario

A bank refund agent "must always" refuse wire transfers outside an allowlisted set of accounts. The system prompt already says never to transfer outside policy. What primary control should the architect add?

Choose one answer

Build exercise

Design layered guardrails for a bank refund agent with wire-transfer tools

35 minutes

What you'll learn

  • Map irreversible actions to hard controls
  • Layer prompt guidance with I/O and tool enforcement
  • Log triggers and define residual-risk paths
  • Prove blocks with adversarial eval cases
  1. Step 1

    Inventory irreversible and sensitive actions

    List tools and outputs that can move money, change privileges, delete data, or expose secrets/PII. Mark each as hard-block, filter, or escalate.

    Why: Guardrails start from harm classes, not from prompt style. Exam items reward mapping risk to control type.

    You should see: A table: action → harm → control type (block / filter / HITL) → owner.

  2. Step 2

    Layer prompt guidance with deterministic enforcement

    Keep clear system-prompt rules, then implement matching checks: input filters, output filters, tool allow/deny lists, and pre-execution policy hooks.

    Why: Defense in depth. Prompts steer; code enforces “must always.”

    You should see: Diagram showing prompt → model → output filter → tool policy → execute/deny.

  3. Step 3

    Log triggers and route residual risk

    Record every block/escalate with redacted context. Define what happens when the model still attempts a forbidden action (deny, HITL, alert).

    Why: Triggers without review never improve. Residual attempts need an explicit path.

    You should see: Sample log fields + runbook for repeated policy attempts.

  4. Step 4

    Prove the control with adversarial evals

    Add jailbreak, injection, and tool-misuse cases that must be blocked. Ship only when hard blocks hold under attack cases, not only happy path.

    Why: Friendly demos do not prove guardrails. Adversarial coverage is the gate.

    You should see: Five adversarial cases tied to the allowlist/deny rules with expected block outcomes.

Sources