← Quick reference
Domain 5: Governance, Safety & Risk Management
Cheat sheet for Governance: guardrails, risks, HITL, compliance, ethics.
Guardrails & safety controls (5.1)
“Must always / must never” → deterministic tool-path or output controls. Layer prompt + I/O filters + allow/deny lists + escalation. Log and review triggers.
- Prompts alone are soft — not enough for hard safety language
- Block irreversible actions before execution
- Adversarial evals prove the blocks
Risks & failure modes (5.2)
- Enumerate: hallucination, tool misuse, outages, drift, policy gaps
- Name model/tool limitations as design constraints
- Mitigation + residual risk owner for each high-severity mode
- Reject “the model will handle it”
Human-in-the-loop (5.3)
- Measurable triggers (risk class, policy gap, user ask, stuck) — not confidence alone
- Structured handoff: facts, suggestion, uncertainty, decision
- Calibrate auto vs mandatory review by risk tier
- Close the loop: rejects → eval fixtures
Regulatory compliance (5.4)
Map data class + region + deployment to GDPR / HIPAA-oriented / FedRAMP-oriented (as applicable). Minimize sensitive data in prompts/logs; own controls with evidence.
- No generic slogans without regime mapping
- Debug logging of identifiers is a compliance decision
- Auditors want boundaries and evidence, not vibes
Ethical AI (5.5)
- Measure disparate outcomes where the use case warrants it
- Disclose AI involvement and limitations when required
- Provide recourse (appeal / human review)
- Concrete controls beat ethics slogans