CCAR-P · Study Guide

← Domain 3: Integration

3.6 · Lesson 6 of 8

Apply retrieval strategies matched to data shape and query pattern

What You Need to Know

Retrieval strategy is not a brand preference — it is a match between data shape and query pattern. Dense embeddings excel at paraphrased intent. Sparse or keyword signals excel at exact SKUs, error codes, and rare tokens. Structured filters reduce candidate sets when catalogs have fields. Hybrid systems combine signals when traffic is mixed.

Architects classify real queries, route or configure retrieval accordingly, and evaluate per pattern. Increasing top-k or model size does not fix a fundamental mismatch between dense-only search and identifier lookup.

Decision rules

  • Classify queries before picking a stack.
  • Use sparse/hybrid for exact identifiers; dense for paraphrased intent.
  • Apply structured filters when metadata fields exist.
  • Evaluate recall and grounding per query pattern, not only overall averages.
  • Preserve provenance regardless of which ranker wins.

Pattern → strategy map

  • Exact identifiers (SKU, error code, ticket ID) → sparse or hybrid; consider filters
  • Paraphrased how-to / intent → dense semantic ranking
  • Faceted browse (region, status, product) → structured filters first, then rank
  • Mixed traffic → router or hybrid default with per-bucket evals

Why macro metrics lie

Overall recall can look healthy while the exact-ID bucket collapses. Architects publish scorecards per pattern. Raising top-k on a mismatched ranker mostly adds noise. Prefer fixing the signal (sparse terms, filters) over drowning the model in extra chunks.

Whatever retrieves must still return provenance. Operators will not trust a SKU answer they cannot open in the source system of record.

Exam application

Match mode to query pattern. Exact IDs failing under dense-only → sparse/hybrid and filters. Distractors: raise temperature to invent IDs, delete metadata, or rely on macro recall alone.

Exam traps

  • One retrieval stack for every query type

    Paraphrased intent favors dense; SKUs and error codes favor sparse/hybrid; catalogs favor filters first.

  • Semantic similarity for exact identifiers

    Embeddings can blur rare tokens. Exact match paths exist for a reason.

  • Skipping structured filters when fields exist

    Product, region, and status filters cut noise before ranking and are often the highest-ROI change.

  • Confusing more top-k with better matching

    If the ranking signal is wrong for the query pattern, larger k mostly adds noise.

Practice scenario

Warehouse operators query an assistant with exact SKUs, error codes, and part numbers. A dense-embedding-only retriever frequently misses the right row even when the string exists in the corpus. Which change best matches data shape and query pattern?

Choose one answer

Build exercise

Route retrieval by query pattern

40 minutes

What you'll learn

  • Bucket production queries by pattern
  • Match dense, sparse, hybrid, and filters
  • Score per-bucket retrieval quality
  • Keep citations on every path
  1. Step 1

    Classify queries by pattern

    Sample production queries into buckets: exact ID lookup, paraphrased how-to, filtered browse, and multi-hop. Measure failure rates per bucket.

    Why: Strategy follows query shape. Architects who skip this pick a fashionable stack and miss.

    You should see: A labeled query sample with per-bucket failure rates.

  2. Step 2

    Match retrieval mode to each bucket

    Wire sparse/keyword or hybrid for ID buckets; dense for paraphrases; apply structured filters when metadata exists; use routing rules rather than one global mode.

    Why: Exam items key on matching strategy to data and query pattern.

    You should see: Router or config that selects mode by query features (regex for IDs, else hybrid, etc.).

  3. Step 3

    Evaluate per pattern, not only macro accuracy

    Report recall@k and grounded answer rate for each bucket. Do not declare victory on an average that hides ID lookup failures.

    Why: Macro metrics hide pattern-specific failures the exam loves to surface.

    You should see: A scorecard with per-bucket metrics before/after the hybrid change.

  4. Step 4

    Keep provenance on whatever retrieves

    Whether sparse or dense wins a hit, return source metadata so operators can verify the row.

    Why: Retrieval strategy changes do not remove the need for citations and auditability.

    You should see: Answers citing doc/row IDs for both hybrid and dense paths.

Sources