TOPIC #164Advanced 12 min read

Semantic Search

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

Semantic search matches what a query MEANS, not which letters it shares, by comparing query and content as coordinates in one learned meaning-space. In production it is not one model call but a funnel: query understanding, hybrid multi-recall, fusion, reranking, and business signals.

A Modern Semantic Search Stack (2025)

Semantic search is not one model call — it is a multi-recall funnel where embeddings feed fusion, reranking, and business logic.

A Modern Semantic Search Stack (2025)
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: Characters Are Not Meaning

A user types into your site search:

Insight

"laggy checkout on phone"

Somewhere in your blog is the perfect post. Its title:

Insight

"Cutting render delay from your mobile payment flow"

It never says "laggy". It never says "checkout". It never says "phone".

Keyword search answers a question users never ask: "which pages contain these tokens?" It compares characters, so it shrugs at synonyms and paraphrases.

Semantic search answers the question users actually have: "which pages satisfy this intent?"

It works because of embeddings (topic 162, in one sentence here: an embedding assigns every text coordinates in a learned space so similar meanings land close together). "Laggy" and "slow" and "render delay" sit in the same neighborhood of that space, so the query reaches the post even with zero shared words.

But there is a flip side, and it hurts. Because semantic matching ignores exact tokens, it loses on the things that are precisely tokens: order IDs, error codes, SKUs, rare proper names. "ERRX-401" is not a meaning — it is a literal string, and generalizing embedding models blur exactly such rare strings away.

So the real product is never "AI search = one vector call." It is a stack — and this topic is the stack.

02.The Idea in Plain Words: Meaning Matching, Organized as a Funnel

The mechanism in one line:

Insight

Encode the query and all content in the SAME learned coordinate space, then retrieve by closeness (cosine similarity).

But production semantic search wraps that mechanism in stages, because no single model call is both fast over billions of pages AND precise about relevance:

  1. Query understanding — fix spelling, expand synonyms/aliases, detect language, classify intent, maybe rewrite the query into several variants.
  2. Multi-recall — run a dense path (embedding ANN over vectors, topics 162/163) and a lexical path (BM25 exact-term matching, topic 165) in parallel; variants and hypothetical answers (HyDE, see callout) feed the dense path.
  3. Fusion — merge the two result lists into one candidate pool (Reciprocal Rank Fusion, topic 165).
  4. Reranking — score the ~100 survivors with a precise but slow cross-encoder (topic 167).
  5. Business signals — freshness, authority, access control, personalization produce the final order.

Dense and lexical are complementary, not rivals: the embedding catches "laggy → slow", BM25 catches "ERRX-401". Treating hybrid as optional is the classic failure of pure-vector "AI search" launches.

03.A Simple Worked Example: One Query, End to End

Follow "laggy checkout on phone" through the funnel.

Understand. Spellfix: nothing. Rewrite variants (LLM or rules):

code
v1: laggy checkout on phone
v2: slow mobile payment page
v3: reduce checkout load time mobile

Recall. Each variant embeds → ANN fetches top-50; BM25 on v1 fetches its own list.

code
dense  rank 1: "Cutting render delay from
                your mobile payment flow"  cos 0.81
dense  rank 2: "Phone battery drain fixes" cos 0.55
BM25   rank 1: "Checkout page caching 101"  (shares
        "checkout" but is about desktop)

Note the two paths disagree — exactly as designed. Dense found the paraphrase BM25 could not see; BM25 found a term-match dense scored midly.

Fuse (RRF, topic 165). Ranks only: render-delay is dense#1 → big fused score; caching-101 is BM25#1 but dense-nothing → lower.

Rerank (cross-encoder, topic 167). Pairwise attention confirms: render-delay post genuinely answers (score 4.1); caching-101 is topically adjacent but answerless (0.2).

Business layer. Freshness boost, ACL filter (drop internal docs), personalization.

Final: the mobile-payment post at #1, with a snippet highlighting "render delay" — a result with zero words in common with the query. That is the whole trick.

04.Visual Intuition: Wide and Cheap → Narrow and Careful

The funnel shape explains every design decision:

code
        billions of pages
     ┌────────────────────┐
     │  recall (dense+BM25)│  cheap per doc,
     └─────────┬──────────┘  slightly fuzzy
          ~100-150            (coverage matters)
     ┌─────────┴──────────┐
     │  cross-encoder      │  expensive per pair,
     │  rerank             │  very precise
     └─────────┬──────────┘  (ordering matters)
           5-10 results
     ┌─────────┴──────────┐
     │  business signals   │  rules, ACLs,
     └─────────┬──────────┘  freshness, prefs
            final page
  • Early stages answer "did we find it?" → measured by recall@k.
  • Late stages answer "did we place it well?" → measured by nDCG@10.
  • The expensive stage only ever touches ~100 candidates — which is precisely why the funnel exists.

05.The Analogy: The Cab Driver Who Knows the City

Carry one analogy through: keyword search is a phone book; semantic search is a cab driver who knows the city.

  • You tell the phone book "Main Street 12" and it finds the exact string — or nothing. It has no idea where you actually want to go.
  • You tell the cab driver "take me to that place with the best biryani" and he knows. Intent, not letters. That is dense retrieval.
  • But say "pick me up at 42 Ellipse Road" — the driver may fumble a landmark nobody names that way, while the phone book nails it. That is BM25. Serious dispatchers keep both (hybrid).
  • A good driver also clarifies garbled requests ("you said laggy — you mean slow-loading?"). That is query understanding, and you want it fast because every passenger needs it — cache the frequent destinations.
  • At rush hour the dispatcher pulls candidate addresses from two books and merges the lists (fusion), then calls each place to confirm they are actually open before driving you there (reranking).
  • Finally, company policy: prefer the branch you always use, never enter restricted zones (business signals and ACL filters).

One more quirk worth memorizing: if a passenger gives an address-shaped destination but you send it to the "biryani neighborhood" brain, you get a beautiful wrong guess. Rare exact strings belong to the phone-book leg. The dispatch lesson is the same: route intent to the driver, literals to the book, and merge.

06.Query Understanding: The Half You Can Cache

Before embedding anything, normalize intent: spelling correction, synonym/alias expansion (product names, internal jargon — "FusionPro 3000" → canonical SKU family), language detection, and query classification (navigational "log in page" vs informational "what is an API key" vs transactional "buy").

The 2024–2026 shift: a small LLM now performs query rewriting — decomposing "compare our SLA vs competitor X for EU data residency" into 2–3 focused sub-queries executed in parallel and merged. Compound questions are single queries only in appearance; retrieval likes them split.

The economics split cleanly:

  • Query-side transforms run per request → keep them cheap (small models, tight prompts) and cache aggressively: rewrites, expansions, even the final ranking for popular "head" queries are all cacheable by (query, user-context-hash).
  • Document-side work (parsing, chunking, embedding) happens in the ingestion pipeline → it can be slow and thorough. You re-run it whenever the embedder model is versioned (topic 162).
python— Multi-query recall with fusion — the semantic search loop in miniature
def semantic_search(query, index, llm, embedder, k_final=10):
    # 1) LLM query understanding (cacheable)
    variants = llm(
        f"Rewrite for search. Return 3 diverse queries, "
        f"one line each:\n{query}")
    variants.append(query)                # always keep the original

    # 2) Multi-recall: embed each variant, fetch top 50
    pool = {}
    for v in variants:
        for doc, score in index.search(embedder(v), k=50):
            pool[doc.id] = max(pool.get(doc.id, 0), score)

    # 3) Rerank survivors with a cross-encoder (topic 167)
    ranked = cross_encoder.rerank(query, list(pool)[:50])
    return ranked[:k_final]

07.Metrics, Multilingual, and Keeping the Index Fresh

Semantic search is judged by ranking metrics, each tied to a funnel stage:

  • recall@k — did retrieval find the right page at all?
  • nDCG@10 — graded relevance weighted by position: was it placed well?
  • MRR — the rank of the first relevant hit (1/rank; rewards "answer #1" over "answer #7").

Behavioral metrics close the loop with real users: click-through on top results, query reformulation rate (a user quickly editing their query = "dwell and pivot", a strong miss signal), and zero-result rate.

Multilingual retrieval no longer requires translate-then-search pipelines: modern embedders (e5-m, bge-m3, multilingual-e5) map 100+ languages into one shared space, so a Spanish query retrieves English docs. Verify, however, that the model saw your language pair and domain vocabulary during training.

Freshness and change data capture: re-embedding entire corpora on every model update is expensive. Track the embedder version per document, queue re-index jobs wherever versions skew, and serve stale-but-tagged vectors in the interim. It is the classic cache-invalidation problem — with a cosine-similarity tax added.

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Matches intent, not tokens — handles synonyms, abbreviations, cross-language queries.
  • Same embeddings power related-items, dedup, and RAG recall — one artifact, many features.
  • Query rewriting + multi-recall demonstrably lifts nDCG on long/compound queries.

Trade-offs & Constraints

  • Fails on exact identifiers, rare entities, and freshly-coined terms until the embedder learns them (training cutoff applies to embedders too).
  • Every query pays embedding + ANN latency; relevance regressions are silent without eval sets.
  • The full stack (rewrite → hybrid → rerank) adds LLM cost and p95 to reach ~900 ms, versus ~100 ms classic keyword search.
Production Implementation in Big Tech
Perplexity / Bing Copilot-style answer engines• Web-scale semantic retrieval feeding generated answers

User queries are expanded and rewritten; a dense semantic index plus traditional lexical/web ranking produce candidate pages, which are reranked by quality/relevance models before an LLM composes a cited answer. Semantic recall is the differentiator versus classic SERPs: it fetches pages that answer the intent even when wording is unrelated.

Staff+ Engineering Takeaways

  • Semantic search retrieves by meaning via shared query/document embedding geometry.
  • Production systems are multi-stage: query understanding, hybrid multi-recall, fusion, reranking, business signals.
  • Pure semantic recall fails on exact tokens (IDs, codes); hybridization with lexical search is mandatory.
  • Measure recall, ranking (nDCG/MRR), and behavioral signals separately per funnel stage.
  • Embeddings are versioned artifacts: model upgrades require corpus re-embedding with freshness tracking.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

A product-search team finds queries like "ERRX-401" and "SKU 8812" return irrelevant semantic results. The best fix is:

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?