7.1 · 3.0% of the exam · Topic 1 of 4
AI Application Security
Protect Claude applications from prompt injection, jailbreaks, untrusted input, data leakage, and unauthorized access while preserving confidentiality, privacy, and integrity.
Learning objectives
- Recognize prompt injection and jailbreak risks.
- Handle untrusted input safely.
- Prevent sensitive data and PII leakage.
- Apply authentication and authorization controls.
- Protect confidentiality, privacy, and integrity.
Detailed theory
Prompt injection and untrusted input
Treat external content as untrusted data rather than instructions.
Prompt injection can attempt to override application rules or make Claude disclose information or perform unauthorized actions.
- Separate trusted instructions from untrusted content.
- Validate and constrain tool inputs.
- Do not allow retrieved content to redefine system policy.
Data leakage and PII
Claude applications should minimize the sensitive information exposed to the model and to downstream systems.
PII and confidential data should be handled according to the application's security and privacy requirements.
- Minimize sensitive data exposure.
- Redact or filter unnecessary PII.
- Restrict access to confidential information.
- Avoid returning secrets in model output.
Authentication and authorization
Authentication establishes who the requester is. Authorization determines what that requester is allowed to access or perform.
Claude should not be treated as the authorization boundary for sensitive operations.
- Authenticate users and services.
- Authorize every sensitive action.
- Apply least privilege.
- Validate permissions before executing tools.
Core concepts
Prompt injection
- What
- Untrusted content attempts to influence instructions or behavior.
- Why
- It can cause unauthorized actions or disclosure.
- When
- Processing user input or external content.
- When not
- Trusted application instructions that your system controls.
Jailbreak
- What
- An attempt to bypass model or application safety constraints.
- Why
- It can produce unsafe or unauthorized behavior.
- When
- Users intentionally try to circumvent restrictions.
- When not
- Normal requests that follow the application's rules.
Least privilege
- What
- Giving a user, service, or tool only the permissions it needs.
- Why
- Limits the impact of compromised or incorrect behavior.
- When
- Designing authentication and tool access.
- When not
- Never grant broad access simply for convenience.
Practical examples
Untrusted document content
A Claude application processes a user-uploaded document that contains instructions telling the model to ignore application rules.
The document should be treated as untrusted data. Application instructions and authorization controls remain authoritative.
Sensitive tool access
A support assistant can retrieve customer records through a tool.
Authentication identifies the user, while authorization determines which customer data and actions that user may access.
Claude-specific considerations
- Treat user-provided and external content as untrusted.
- Prompt instructions should not be the only security boundary for sensitive actions.
- Authentication establishes identity; authorization determines permitted actions.
- Apply least privilege to Claude tools and connected systems.
- Minimize unnecessary PII and confidential data exposure.
Architecture decisions
Tradeoffs
Security controls add validation and authorization steps, but reduce the impact of prompt injection, incorrect model behavior, and compromised inputs.
Quick reference
- Treat user and external content as untrusted.
- Prompt injection attempts to influence application behavior through untrusted content.
- Authentication identifies the requester.
- Authorization determines what the requester may access or perform.
- Use least privilege for tools and connected systems.
- Minimize PII and confidential data exposure.
Decision rules for the exam
Common exam traps
Exam tips
- Separate trusted application instructions from untrusted external content.
- For sensitive actions, look for an application-level authorization control.
- Authentication answers who the requester is; authorization answers what they can do.
Common mistakes
Treating retrieved content as trusted instructions.
Keep external content separate from trusted application policy.
Giving Claude broad permissions for convenience.
Apply least privilege to tools and connected systems.
Sending unnecessary PII to the model.
Minimize and filter sensitive data before exposure.
Practice questions
Original questions for this topic. They are study items, not questions from the live exam.
Scenario questions
Build exercise
Threat-model a Claude application
Intermediate · 30 minutes
What you will learn
- Identify untrusted inputs.
- Find potential prompt injection paths.
- Identify sensitive data exposure.
- Apply authentication and authorization boundaries.
Step 1
Map the inputs
List every user, document, webpage, and tool result that enters the Claude application.
Why: Security starts by identifying what is trusted and untrusted.
You should see: A list separating trusted instructions from external content.
Step 2
Find sensitive actions
List the tools and operations that can read, modify, or delete sensitive data.
Why: Sensitive actions need explicit authorization.
You should see: A small list of protected operations.
Step 3
Add security boundaries
Define authentication, authorization, input validation, and least-privilege controls.
Why: Model instructions alone are not a complete security boundary.
You should see: A control mapped to each sensitive action.
Review checklist
Checks are saved in this browser.
Key takeaways
- Treat user and external content as untrusted.
- Do not rely on prompts as the only security boundary.
- Minimize sensitive data and PII exposure.
- Authorize sensitive actions before execution.
- Use least privilege to reduce blast radius.
Sources
- Anthropic Safety — Anthropic safety and responsible AI information.
- Claude Documentation — Claude platform documentation and application guidance.
- CCDV-F blueprint notes — Developer certification study reference.