TOPIC #166Advanced 12 min read

Chunking Strategies for RAG

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

Chunking means slicing documents into retrieval-sized pieces, and the slice you choose decides what your search can ever find: one vector can only carry about 512 tokens of meaning, but every cut can sever a cross-reference. This topic compares fixed, recursive, structural, semantic, late, and parent-child schemes.

Ingestion: One Document, Many Chunking Regimes

Parsing normalizes structure first; the strategy choice fixes the granularity of everything downstream. Parent-child linking decouples retrieval precision from prompt context.

Ingestion: One Document, Many Chunking Regimes
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: One Vector Cannot Hold a Whitepaper

Topic 162 said: an embedding turns text into coordinates, and similar meanings land close together.

But embedders are trained on short spans — mostly 256–512 tokens. Their notion of "meaning" was built at paragraph size.

Now feed one a 20-page whitepaper.

You get one vector for ~10,000 tokens. What does it point at? Everything and nothing: a bag of averages — a blurry coordinate that is vaguely about pricing and architecture and security and the FAQ, and therefore precisely similar to no specific question.

Try the query "What is the API rate limit?" against that single vector: cosine similarity lands around 0.3. Against the one paragraph that actually states the limit: 0.8. The document contained the answer. The retrieval unit couldn't show it.

So the question becomes

Insight

How do we slice documents so each piece is small enough to have one clear meaning — without cutting the meaning out of any piece?

Because there is a second law fighting you: a query can only match a chunk if everything relevant to it lives inside that chunk. Split a contract mid-clause and the chunk that says "as defined in §4.2" can never satisfy "what counts as Confidential Information?" — the definition lives across the seam.

Chunking is the compromise between these two laws. This whole topic is picking where to cut.

02.The Idea in Plain Words: Slice to the Size of a Question

Chunking = splitting each document into small, self-contained retrieval units ("chunks"), embedding each chunk separately, and indexing them all.

Three plain-word rules:

  • The chunk is the retrieval unit — the thing search compares against a query. It should ideally hold one topic, sized to the embedder's trained window (256–512 tokens).
  • Every cut is a seam, and seams amputate cross-references: pronouns ("as shown above"), clause pointers ("per §4.2"), and running explanations get orphaned.
  • The cardinal rule of 2024–2026 RAG practice:
Insight

Search small, show big. Optimize chunks for precise matching, then hand the generator more context than you searched with.

That last move has a name — small-to-big retrieval (parent-child): you index fine chunks, but when one wins, you fetch its parent (the whole section) into the prompt. Search granularity and prompt granularity are decoupled, and each can be set to what it does best.

Defaults that survive contact with reality: 256–512 tokens for Q&A over long docs, 10–20% overlap as a seam tax, and never chunk code or tables by raw character count.

03.A Simple Worked Example: Where the Knife Falls

A Kafka runbook section reads (450 tokens total):

code
"## Message Retention
Kafka stores messages for 7 days by default...
... tuning discussion ...
As noted above, the limit is 500/s per broker."

Query: "What is the default retention?"

Cut #1 — fixed 200-token windows. The answer sentence "stores messages for 7 days" survives in chunk 1 → retrieved (cos 0.84). Good.

Cut #2 — same windows, query is "what is the per-broker limit?" The sentence lands as "As noted above, the limit is 500/s" — "above" refers to text that is now in a different chunk. The vector for this chunk is about nothing in particular; cos ≈ 0.31. Missed. This is the seam problem in one line.

Fixes ranked by cost:

  • Overlap 40 tokens (20%): the reference sentence rides along with a copy of its context → cos ≈ 0.6. Partial heal.
  • Header splitter: cut at "## " boundaries so each chunk keeps its section title — the title itself disambiguates ("Retention" + "limit 500/s" in one unit) → cos ≈ 0.8.
  • Context header (Anthropic-style): prepend "This excerpt is from the retention section of the Kafka runbook" to the chunk before embedding → cos ≈ 0.82, at ingestion LLM cost.

One document, three cuts, three very different answers to the same question. That is why chunking is a design decision, not a detail.

04.Visual Intuition: The Seam and the Ladder

A document as a strip; each strategy draws the knife differently:

code
doc:      [ intro | retention | tuning | limits | FAQ ]

fixed:    [intro|reten][tion|tuni][ng|limit][s|FAQ]   ← cuts
                                                   slice mid-topic
recursive:[intro][retention][tuning|limits][FAQ]      ← cut at
                                                   paragraphs first
structural:## Intro ## Retention ## Limits ## FAQ     ← cut at headers

small-to-big:  index  [ret-sent-1][ret-sent-2][ret-sent-3]
                 ↑ hit one sentence...
                 ...prompt returns  [ whole ## Retention section ]

And the ladder picture (RAPTOR / hierarchical): sentences → chunks → section summaries → document summaries, every level embedded. Local questions hit low rungs; "what are the main themes of the whole corpus?" hits a high rung that no leaf chunk could answer.

05.The Analogy: Cutting a Magazine into Note Cards

Carry one analogy through: your document set is a magazine archive, and you are cutting it into index cards for the open-book exam from topic 161.

  • One card, one fact. A card with three unrelated clippings glued on helps the student find none of them — that is the whole-document "bag of averages" vector.
  • Never cut mid-sentence. A card that reads "...as shown above, the limit is 500/s" is useless in isolation — "above" was on the previous card. This is the seam problem; every strategy below is a different answer to "where does the scissor go?"
  • Write the page number on each card. Metadata — "Kafka runbook > Retention, p.4" — turns a vague clipping into a self-describing one. This is exactly what context headers and structure-aware splits do.
  • Small-to-big is the pull-out section. The card gets you to the right page; you then hand the student the whole page, because the full answer needs surrounding paragraphs. You searched cards, you grade pages.
  • Late chunking is reading before cutting: you skim the whole magazine first, so each card you write can silently carry the context of the articles around it.
  • RAPTOR is a summary card on top of the pile: "Cover story: why retention matters" — coarse questions get answered without hunting dozens of leaf cards.

The exam never changes. Only how you cut — and what the student can therefore find.

06.The Strategy Menu

Fixed-size: split every N tokens. Fast, deterministic, a strong baseline — and the strategy sophisticated systems keep failing to beat. Ignores sentence and topic boundaries; the seam problem is maximal.

Recursive character splitting: try paragraph breaks, then newlines, then sentences, then characters, until a chunk fits the size budget. The LangChain/LlamaIndex default; respects natural boundaries at modest cost.

Structure-aware splitting: parse to markdown/HTML/AST first and split on document structure — headings, sections, list items, table rows, code functions. For Markdown use a header splitter (e.g., MarkdownHeaderParentTextSplitter); for code, split on top-level functions/classes so retrieval units are semantically complete symbols (tree-sitter based splitters).

Semantic chunking: embed sentences in a sliding window and cut where consecutive similarity drops (a topic shift). Better topic purity, costs an O(n) embedding pass at ingestion; empirically inconsistent (see the warning callout above) and slow on huge corpora.

Late chunking (Jina AI, 2024): invert the order — run the Transformer over the WHOLE document to get contextual token embeddings, then mean-pool per planned chunk. Each chunk vector carries full-document context, so "…as shown above, the limit is 500/s" resolves correctly even in a small chunk. Requires long-context embedders, and any text edit forces re-chunking (the vectors of siblings depended on it).

Hierarchical (RAPTOR): recursively summarize chunks into parent summaries, embedding every level of the tree. Wins on corpus-level questions ("what are the main themes?") at the cost of an LLM-heavy build.

python— Recursive splitting with structure awareness + parent linkage (small-to-big)
from langchain_text_splitters import (
    MarkdownHeaderTextSplitter, RecursiveCharacterTextSplitter)

hdr = MarkdownHeaderTextSplitter(
    headers_to_split_on=[("##", "h2"), ("###", "h3")],
    strip_headers=False)

sub = RecursiveCharacterTextSplitter(
    chunk_size=350, chunk_overlap=50,
    separators=["\n\n", "\n", ". ", " "])

parents, child_to_parent = [], {}
for section in hdr.split_text(raw_markdown):
    pid = len(parents); parents.append(section)      # h2 section = parent
    for c in sub.split_text(section.page_content):
        child_to_parent[len(children)] = pid          # link for small-to-big
        children.append(c)
# Index ONLY the children; at query time fetch parent by id,
# dedupe, and stuff parents into the prompt.

07.Matching Chunking to Content Type and Query Type

There is no universal chunk size — the right granularity is a function of document shape and question shape:

  • Prose / FAQ / support docs: 256–512 tokens, recursive or header-based; questions are local, so small precise chunks win.
  • Legal / contracts: clause-level units plus defined-term registers; cross-references ("per §4.2") demand parent-child or late chunking, because the definition and the usage never live in one clause.
  • Code: function/method-level chunks, and embed a metadata header inside the chunk text itself — "In file auth/jwt.py, function refresh_token(): ..." — so retrieval knows what scope a snippet lives in (a large repo has many identically named helpers).
  • Tables / aggregations: "total spend in 2025" is not chunk-retrievable — no slice of a 400-row table contains the sum. Summarize tables into sentences for the long tail, and route aggregate queries to a SQL / text-to-SQL tool instead (topic 168).
  • Multi-document synthesis: RAPTOR hierarchies or GraphRAG-style entity graphs, which replace "chunks" with communities/relations as the retrieval unit.

And remember: metadata is half the chunk. Attach source, heading path ("Runbook > Kafka > Retention"), page, and version — the LLM uses it for citations, and the filter layer uses it for scoping (topic 163).

08.In Practice: The Decision You Actually Ship

The consensus recipe, stated as one sentence: structure-aware recursive split at 256–512 tokens, 10% overlap, parent-child linking, metadata embedded in the text — then measure before you fancy it up.

  • Try the cheap ladder first: fixed → recursive → header-based. Each step is minutes of engineering; only move to semantic/late/RAPTOR when your eval set proves the simpler cut loses.
  • Anthropic's Contextual Retrieval (real-world example below) is the strongest "context header" result published: prepending ~100 LLM-generated context tokens per chunk cut retrieval failure by ~49%, and ~67% combined with reranking — but budget the ingestion LLM calls.
  • Tiny corpora that fit whole documents in context? Skip chunking — retrieve whole docs or paste everything (topic 161's "when to avoid").

Chunking is the cheapest, highest-leverage dial in the whole RAG stack: it is the only one that changes what is findable at all.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Small chunks: precise matching, fine-grained citations, cheaper reranking pools.
  • Late chunking / parent-child: recover document-level context without coarse vectors.
  • Structure-aware splits keep tables, code, and clauses semantically intact.

Trade-offs & Constraints

  • Seams destroy cross-reference answers; overlap only partially heals them.
  • Semantic chunking and RAPTOR add real ingestion cost for inconsistent gains.
  • Small retrieval units mean the prompt must reassemble context — more plumbing, more failure modes.
Production Implementation in Big Tech
Anthropic — Contextual Retrieval (2024)• Making standalone chunks self-describing

Anthropic's published pattern prepends ~100 tokens of LLM-generated context ("This excerpt is from the retention section of the Kafka runbook…") to every chunk before embedding with a Matryoshka model, combined with BM25 hybrid retrieval and reranking; internal evals reported up to ~49% (and ~67% combined) reductions in retrieval failure rate.

Staff+ Engineering Takeaways

  • Chunking fixes the retrieval unit; embedders lose fidelity beyond ~512-token trained windows, so long docs must be sliced.
  • Recursive + structure-aware splitting is the workhorse; fixed windows are the baseline that beats much fancier schemes per dollar.
  • Small-to-big (search chunks, prompt parents) and late chunking reconcile precise matching with rich context — "search small, show big."
  • Content type dictates strategy: code by symbol, contracts by clause, tables need summarization or a SQL tool, not chunks.
  • Metadata (source, heading path, version) embedded with each chunk powers citations, filtering, and eval debugging.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

What problem do "small-to-big" (parent-child) retrieval and late chunking each solve?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?