CCAR-P · Study Guide

← Domain 2: Claude Models, Prompting & Context Engineering

2.5 · Lesson 5 of 5

Implement prompt reuse strategies (caching, modular prompts, Skills)

What You Need to Know

Prompt reuse keeps quality consistent without flooding every turn. Cache stable prefixes — system instructions, schemas, and other unchanging blocks — where the platform supports prompt caching. Modularize shared instructions so multiple agents import one versioned source of truth instead of forking private copies.

Skills (and similar isolated workflows) hold verbose reusable playbooks and bring them in when needed, so live context stays lean. The anti-pattern is inlining the full skill text into every user message. Architects design for reuse mechanisms that reduce repeated tokens and drift — not for paste-everything convenience.

Reuse toolkit

  • Prompt caching for stable prefixes
  • Modular shared instructions with owners/versions
  • Skills isolating verbose reusable workflows
  • Lean live turns: invoke depth only when needed
  • Measure cache hits and tokens per turn after reuse

Decision rules

  • Cache stable prefixes where the platform supports it.
  • Modularize shared instructions — one source of truth.
  • Use Skills to isolate verbose reusable workflows.
  • Keep live context lean — do not inline full skill text every turn.
  • Prefer reuse that cuts tokens and drift, not private forks.

Lean live context

The hot path should carry what this turn needs plus stable cached structure. Deep playbooks belong behind a Skill or module boundary. Exam stems that show a huge paste every run reward caching, modularization, and isolation — not “keep pasting so nothing is missed.”

Review checklist

  • Are stable vs volatile parts identified?
  • Is the stable prefix cache-eligible and measured?
  • Do agents share versioned modules instead of forks?
  • Is verbose workflow isolated (Skill) rather than inlined?
  • Did tokens/latency per turn drop after reuse?

Prefer answers that cache, modularize, and isolate. Distractors celebrate inlining, private duplicates, or discarding shared config.

Exam application

Choose reuse that keeps live context lean: caching, modules, Skills. Reject full inline every turn and unmanaged private copies.

Exam traps

  • Inlining full skill / playbook text every turn

    Verbose reusable workflows should be isolated and invoked — not pasted into live context each time.

  • Skipping cache on stable prefixes

    Unchanging system instructions and tool schemas are prime cache candidates where the platform supports prompt caching.

  • One monolithic prompt with no modules

    Shared rules duplicated across agents drift. Modular blocks keep one source of truth.

  • Private fork copies per engineer

    Reuse fails when everyone maintains a slightly different prompt. Prefer shared modules and Skills.

Practice scenario

A team pastes a 4k-token review playbook into every user message. The playbook rarely changes. Live turns are slow and expensive. What reuse strategy should the architect prefer?

Choose one answer

Build exercise

Refactor a review agent for caching, modules, and Skills

40 minutes

What you'll learn

  • Split stable prefixes from volatile turn content
  • Apply prompt caching to the stable prefix
  • Modularize shared instructions across agents
  • Isolate the verbose playbook as a Skill
  1. Step 1

    Identify stable vs volatile prompt parts

    Split the review agent into stable prefix (role, policies, schemas) versus per-turn volatile content (ticket text, tool results).

    Why: Only stable prefixes are good cache and module candidates.

    You should see: A two-list breakdown: stable blocks vs turn-specific inputs.

  2. Step 2

    Enable caching for the stable prefix

    Structure requests so the large unchanging prefix is cache-eligible per platform rules, and measure cache hit rate on the hot path.

    Why: Prompt caching cuts repeated cost and latency for stable tokens.

    You should see: A request shape diagram marking the cached prefix boundary.

  3. Step 3

    Modularize shared instructions

    Extract shared policy and format blocks into versioned modules imported by FAQ, review, and escalation agents.

    Why: One edited module beats three drifting copies.

    You should see: A module map with owners and versions.

  4. Step 4

    Isolate the verbose workflow as a Skill

    Move the long review playbook into a Skill (or equivalent isolated workflow) invoked when needed — do not inline the full text every chat turn.

    Why: Skills keep live context lean while preserving reusable depth.

    You should see: Skill entrypoint + when-to-invoke note; main chat without the 4k paste.

Sources