Vector ANN Recall vs Latency Lab (Interactive)
Probe clusters in an embedding projection and dial nprobe, k, and corpus size to trade recall for latency. Brute-force k-NN versus Approximate Nearest Neighbor search computed live, with recall@K, visited fraction, and distance metrics.
ANN Vector Search: Recall vs Latency
Exact k-NN scans every embedding; the ANN index probes only nearby clusters — tune the tradeoff live.
Recall@5
100%
Vectors distance-scored
128 / 1,200
11% of corpus visited
Brute-force k-NN
1.2 ms
O(N) — unusable past millions
ANN latency
0.58 ms
2x faster
Real RAG systems live on this curve: HNSW (or IVF like modeled here) trades exactness for O(log N) traversal. nprobe=1 is sub-millisecond but drops true neighbors in un-probed clusters; raising it converges toward 100% recall and brute-force cost. Pinecone/Qdrant tune these graph parameters per index; pgvector defaults fit <10M vectors.
How It Works Under the Hood
Embedding models map meaning to high-dimensional coordinates where proximity equals semantic similarity, but exact k-NN over one hundred million vectors is linear per query, far too slow for RAG. Approximate Nearest Neighbor structures like HNSW layer navigable graphs, or IVF cluster centroids, to visit a fraction of the corpus in logarithmic time. Tuning probe breadth moves you along the recall-latency frontier that every retrieval-augmented generation pipeline must choose deliberately.
Core Architectural Principles
- Cosine versus Euclidean metrics score vector proximity differently on unnormalized embeddings.
- ANN visits only nearby clusters, trading a measurable percent of recall for orders of magnitude in speed.
- pgvector fits millions-scale corpora in Postgres; Pinecone or Qdrant shard billions of vectors.
In RAG designs state the retrieval mechanics: chunk documents, embed with the same model at query time, ANN-search top-K context. Quantify the ANN contract, about ninety-nine percent recall for hundred-fold speedup, and mention hybrid search fusing BM25 exactness with dense semantics via reciprocal rank fusion for SKUs and names.
Millisecond semantic retrieval at slightly imperfect recall versus brute-force exactness measured in seconds.