Recommendation System Funnel Lab (Interactive)
Score a 500M-item catalog with one model or a two-tower funnel and compare latency and GPU spend. Walk the retrieval-filtering-ranking-diversity funnel, compute per-request items scored, end-to-end milliseconds, and the GPU fleet a monolithic deep ranker would demand.
Two-Tower Funnel vs Monolithic Ranker
Price the retrieval→ranking funnel against scoring the whole catalog per request.
How It Works Under the Hood
Production recommenders cannot push a hundred-million-item catalog through a deep cross network on every request, so they funnel. A two-tower model precomputes item embeddings into an HNSW vector index; at request time only the user tower runs, and one approximate nearest-neighbor search returns hundreds of candidates in milliseconds. A heavy ranker scores that shortlist with thousands of cross features, and a diversity re-rank emits the final ten. The split works because item vectors are static while user context is not, decoupling offline indexing from online inference.
Core Architectural Principles
- Two-tower retrieval shrinks the scoring set roughly six orders of magnitude before ranking.
- Static item embeddings are indexed nightly; only the user tower runs inside the request path.
- Feature stores hydrate online features under 2 ms to prevent train/serve skew.
Never propose one monolithic model for a large catalog. Draw the funnel with explicit scales: 100M to 500 by ANN, 500 to 30 by deep ranking, 30 to 10 by diversity, each with a latency budget. Then handle cold start with content embeddings and bandit exploration, and mention re-ranking constraints like dedup and freshness that rankers alone ignore.
Funnel stages buy latency and compute but cascade recall loss if early retrieval misses relevant items.