CCAR-P · Study Guide

← Domain 3: Integration

3.5 · Lesson 5 of 8

Design a RAG pipeline with appropriate chunking and indexing strategies

What You Need to Know

A RAG pipeline is only as good as its chunking and indexing. Chunks that ignore document structure lose clause boundaries and table meaning. Indexes that skip metadata lose filters and provenance. Pipelines that upload files without re-embedding leave the live index stale — a frequent cause of wrong answers after a documentation refresh.

Architects design chunk size and boundaries from data shape and query needs, attach provenance fields, choose index types that match expected queries, and operate a freshness swap with smoke tests. When answers go wrong after a corpus change, investigate retrieval and index health before swapping models or raising temperature.

Decision rules

  • Let document structure and query needs drive chunk boundaries — not a single global token size.
  • Attach provenance metadata (source, section, version, date) to every chunk.
  • Treat re-embed + index promote as a required deploy step after corpus refresh.
  • Debug wrong post-refresh answers via retrieval provenance before generation changes.
  • Keep overlap modest and intentional; overlap is not a substitute for structure-aware splits.

Chunking follows meaning

A chunk should be the smallest unit that still answers a plausible query without requiring its neighbor. Policies need clause integrity. FAQs need Q+A integrity. Tables need header context with row groups. Blind fixed windows optimize for tokenizer convenience, not for retrieval. Modest overlap can bridge boundaries, but it cannot repair a strategy that bisects every rule mid-sentence.

Index freshness is an ops concern

  • Detect corpus change (hash, version, effective date)
  • Re-chunk and re-embed into a shadow index
  • Smoke-test known queries for new and deprecated content
  • Atomically promote; alert on index age
  • Keep provenance so wrong answers are debuggable

Sample cue: wrong answers after a documentation refresh → investigate retrieval and indexing before generation knobs or model upgrades.

Exam application

Structure-aware chunks and provenance metadata beat fixed windows. After refresh failures, check index freshness and top-k provenance before temperature or model swaps. Upload ≠ index.

Exam traps

  • One-size fixed token chunks for every corpus

    Legal clauses, FAQs, and tables need structure-aware splits. Fixed windows that break meaning hurt retrieval.

  • Blaming the model first after a corpus refresh

    If prompts are unchanged and docs moved, check indexing and retrieval provenance before changing model tier.

  • Indexes without metadata

    Missing source, section, version, and date fields blocks filters and auditability.

  • Assuming upload equals searchable

    Objects in a bucket are not an embedding index. Pipelines must re-chunk, embed, and swap the live index.

Practice scenario

After a documentation refresh, a RAG assistant starts answering policy questions incorrectly even though the new PDFs are in the object store. Generation prompts were unchanged. What should the architect investigate first?

Choose one answer

Build exercise

Design a policy-doc RAG chunk and index pipeline

45 minutes

What you'll learn

  • Structure-aware chunking for policies
  • Provenance metadata on chunks
  • Freshness-safe index promotion
  • Retrieval-first diagnosis after refresh
  1. Step 1

    Choose chunking that follows document structure

    For policies and contracts, split on headings/clauses with modest overlap. For FAQs, keep Q+A pairs intact. Avoid blind fixed windows that cut mid-rule.

    Why: Chunk shape drives whether retrieval returns a complete, usable unit of meaning.

    You should see: Sample chunks that start/end on structural boundaries with overlap policy documented.

  2. Step 2

    Attach provenance metadata at index time

    Store source URI, title, section path, version, and effective date on each chunk. Require citations to return these fields.

    Why: Metadata enables filters, debugging, and trust. Exam items treat provenance as first-class.

    You should see: A retrieved hit showing source and section, not only raw text.

  3. Step 3

    Build an index swap with freshness checks

    Re-embed changed docs, build a shadow index, run retrieval smoke tests, then atomically promote. Alert if live index age exceeds policy.

    Why: Refresh failures are a top cause of "suddenly wrong" RAG answers.

    You should see: Pipeline logs showing build → smoke → promote, plus an index age metric.

  4. Step 4

    Verify grounding before changing generation config

    For failing questions, inspect top-k chunks. If the right text is missing, fix retrieval/index. Only then tune prompts or model.

    Why: Diagnosis order matters: retrieval first after corpus change.

    You should see: A failing case with empty/wrong retrieval fixed by re-index, quality restored without model change.

Sources