2.3 · Lesson 3 of 5
Apply prompt engineering techniques (zero-shot, few-shot, chain-of-thought)
What You Need to Know
Prompt engineering techniques are tools matched to failure modes. Zero-shot works when the task is clear and well-specified — a sharp rubric often beats ceremony. Few-shot helps when format consistency or edge-case behavior matters: a few strong exemplars teach structure better than more adjectives. Chain-of-thought (or explicit scratchpads) helps multi-step reasoning when the extra tokens improve measured accuracy enough to justify latency and cost.
The architect’s rule: pick the lightest technique that fixes the named failure. Do not default to CoT on latency-sensitive simple tasks. Do not skip few-shot when the bug is format drift. Technique shopping without evals is theater.
Technique map
- Zero-shot — clear task, clear rubric, no format pain yet
- Few-shot — schema/format drift or edges need demonstration
- Chain-of-thought — multi-step reasoning failures with measured gain
- Always: name failure mode → try lightest fix → measure
- Hot path: respect latency SLAs when adding reasoning tokens
Decision rules
- Zero-shot for clear, well-specified tasks.
- Few-shot when format or edge cases are the failure mode.
- CoT when multi-step reasoning earns its tokens on evals.
- Lightest technique that fixes the failure — not the heaviest by default.
- Do not default CoT on simple latency-sensitive paths.
Earning tokens
Every technique spends context and often latency. Few-shot examples and CoT traces are not free. Report Δ quality against Δ latency/cost before promoting a heavier technique to the production hot path. A hard-slice-only CoT route can be smarter than CoT everywhere.
Review checklist
- Is the failure mode named?
- Was zero-shot with a sharp rubric tried first when appropriate?
- Do few-shot examples teach format/edges without contradiction?
- Is CoT justified with measured accuracy gain and SLA fit?
- Is the hot path still on the lightest clearing technique?
Stems that push CoT “for quality” on a simple classifier reward refusing unnecessary reasoning tokens.
Exam application
Match technique to failure mode and prefer the lightest fix. Distractors: always-CoT, few-shot for unrelated problems, deleting rubrics, or raising temperature instead of structure.
Exam traps
Default CoT on simple, latency-sensitive tasks
Extra reasoning tokens cost latency and money. Exam stems reward lightest technique that fixes the failure.
Zero-shot when format or edge cases are the real bug
Clear instructions without examples often still drift on schema. Few-shot examples of the desired format fix format failure modes.
Few-shot examples that teach the wrong pattern
Misleading or contradictory exemplars poison behavior. Choose examples that cover edges you care about.
Technique shopping without naming the failure mode
Pick zero-shot, few-shot, or CoT because of a measured problem — not because a blog listed them.
Practice scenario
A classifier must map tickets to five labels using a clear rubric. Latency is tight. The team proposes chain-of-thought on every request “for quality.” What is the best first move?
Build exercise
Choose PE techniques for a latency-sensitive ticket classifier
35 minutes
What you'll learn
- Name the failure mode before picking a technique
- Prove zero-shot with a sharp rubric when possible
- Add few-shot for format and edges only
- Justify CoT with measured gain vs latency
Step 1
Name the failure mode
For a ticket classifier, record whether failures are unclear task specification, format drift, edge-case confusion, or multi-step reasoning errors.
Why: Technique choice follows the failure — not habit.
You should see: A short failure-mode card with one primary label.
Step 2
Try zero-shot with a sharp rubric
Write a precise task + acceptance criteria and run the gold set before adding examples or reasoning scaffolding.
Why: Many well-specified tasks clear the bar without extra tokens.
You should see: Zero-shot prompt + eval score on the gold set.
Step 3
Add few-shot only for format or edges
If format or rare edges fail, add a few high-quality examples that show the schema and those edges — not a giant noisy set.
Why: Few-shot earns tokens when examples teach structure or edges the model misses.
You should see: 3–5 exemplars covering happy path + two hard edges.
Step 4
Justify CoT with measured gain
Enable chain-of-thought (or scratchpad) only on the subset where multi-step reasoning fails and measure Δ accuracy vs Δ latency/cost.
Why: CoT must earn its tokens against an SLA — especially on hot paths.
You should see: A/B note: CoT vs no-CoT on the hard slice with p95 impact.
Sources
- Prompt engineering overview — docs.anthropic.com
- Multishot prompting (few-shot) — docs.anthropic.com — examples for format and edges
- Chain of thought — docs.anthropic.com — when reasoning scaffolds earn tokens