Reranking: Cross-Encoders & LLM Listwise Ranking
Reranking means taking the ~100 candidates a fast first-stage search returns and re-scoring them carefully, pair by pair, with a model that lets the query and document attend to each other. It is the single highest-leverage quality upgrade in a RAG or search stack — paid for by only ever running it on a small pool.
01.The Problem: The First Stage Never Let the Two Texts Meet
Embedding-based retrieval (topic 162) works by crushing an entire document into ONE vector, and the query into ONE vector, and comparing two lists of numbers.
Read that again: the two texts were never in the same room. No word of your query ever "looked at" any word of the document. This is called independent encoding — and it is exactly what makes billion-scale ANN search possible.
It is also why dense retrieval tops out. Some judgments structurally require the texts to interact:
Query: "In what year did Priya Sharma join the company?"
- Passage A: "Priya Sharma joined in 2021 and now leads infrastructure." — states the fact.
- Passage B: "Priya Sharma's career at the company has been marked by rapid promotion since her arrival." — about the topic, but no year.
Both passages are clearly "about Priya joining" — a bi-encoder happily rates both near the query. Only actually reading A and B against the question exposes the difference: does this passage state the answer, or merely discuss the topic? Negation ("does NOT rotate keys"), entity confusion ("Sharma joined, not Chen"), and answerability are invisible to independent vectors.
So the question becomes
Can we afford a careful reading — and if so, of how many documents?
Two-Stage Retrieve-then-Rerank Funnel
Two-Stage Retrieve-then-Rerank Funnel
Recall optimizes coverage at billions scale; reranking optimizes precision on ~100 candidates where full cross-attention is affordable.
Unlock Topic #167: Reranking: Cross-Encoders & LLM Listwise Ranking
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?