Domain 4 · 20% of the exam
Prompt Engineering & Structured Output
Craft effective prompts, implement structured output patterns, and apply prompt engineering techniques for production Claude applications.
Task statements
- 4.1
System Prompts with Explicit Criteria
Vague instructions such as "be conservative" and "only report high-confidence findings" give the model no decision boundary. Explicit categorical criteria define what to flag and what to skip, severity is calibrated with code examples, and a high false-positive category is disabled until its prompt is fixed.
- 4.2
Few-Shot Prompting
When detailed instructions still produce inconsistent formatting, ambiguous judgement calls, or empty fields for data that exists, few-shot examples are the first technique. Use 2-4 targeted examples that include reasoning. Malformed JSON, fabricated missing fields, and sum mismatches need other techniques.
- 4.3
Structured Output with Tool Use
tool_use with a JSON schema eliminates JSON syntax errors. Prompt-based JSON does not. tool_choice auto may return text, any forces some tool call when the document type is unknown, and a named tool forces one step. The schema does not prevent semantic errors, and nullable fields are what stop fabrication.
- 4.4
Validation, Retry, and Feedback Loops
Retry-with-error-feedback sends the original document, the failed extraction, and the specific validation error. Retries fix format mismatches, structural errors, misplaced values, and mathematical mistakes. They cannot create information that is absent from the source. tool_use removes schema syntax errors; semantic checks, including Pydantic validators, feed the retry loop.
- 4.5
Batch Processing Strategies
The Message Batches API saves 50% against synchronous calls, with a processing window of up to 24 hours and no latency SLA. custom_id matches each request to its result. Blocking work such as a pre-merge check stays synchronous. Overnight debt reports, weekly audits, and nightly test generation use batch. Refine prompts on a sample before you submit, and resubmit only the failed items.
- 4.6
Multi-Instance and Multi-Pass Review
A model that reviews its own output in the same session keeps the reasoning that produced it and tends to confirm those choices. An independent instance judges the output fresh. Large reviews split into a per-file local pass plus a cross-file integration pass so attention is not diluted. Confidence scores route findings only after labelled sets calibrate the threshold.