TOPIC #163Advanced 13 min read

Vector Databases: Pinecone, Chroma, Weaviate, pgvector

AI
AI & ML Editorial
Report an issue
Key takeawayCore Concept Summary

A vector database stores embedding coordinates and answers "what is nearest to this query?" fast enough for live traffic. It does so with approximate indexes (HNSW graphs, IVF-PQ clusters) that trade a little accuracy for 100–1000x speed, plus metadata filtering — the feature that decides whether your RAG is usable or a data leak.

Filtered ANN Query Path in a Vector Database

Modern engines interleave metadata filtering with graph traversal; filter selectivity decides pre-filtering vs in-graph filtering.

Filtered ANN Query Path in a Vector Database
100%
Touchpad: Pinch to zoom • Drag to pan
Rendering visual architecture flowchart...

01.The Problem: You Cannot Check Every Shelf

Topic 162 gave every text a coordinate — an address in meaning-space.

Now a user asks a question, and you need the 20 closest addresses out of your whole corpus.

The honest way — exact k-nearest neighbors (kNN) — compares the query against all N vectors. Cost per query: O(N·d) multiplications.

Put in numbers:

  • 100 million vectors × 1024 dimensions ≈ 100 billion multiply-adds per query.
  • Even one user gets 100+ ms of pure arithmetic; a thousand users hit the same index at once.

Interactive search needs answers in tens of milliseconds. So exactness has to go.

But wait — there is a second problem. A real RAG question is rarely just "find similar."

Insight

"Only documents this user may read, updated after 2024, in product X."

Speed and filters. That combination — a storage engine for coordinates plus fast approximate search plus structured filtering — is what a vector database exists to solve.

02.The Idea in Plain Words: The ANN Bargain

Every vector database delivers

Insight

Approximate Nearest Neighbor (ANN) search: find the top-K results most of the time, 100–1000x faster than checking everything.

The bargain is stated precisely:

  • Recall = the fraction of the true nearest neighbors your fast search actually found. Return 9 of the real top-10 → recall@10 = 90%.
  • Latency = how long one query takes.
  • Memory/cost = how many bytes of RAM (or disk) the index needs.

You cannot max all three. Tuning any vector DB is turning this three-way dial — that is the entire engineering contract.

The two dominant index families, in plain words:

  • HNSW (graph-based, 2016 → the 2024–2026 default): build a network of shortcuts between vectors, layered like an airport map: a few far-apart "long-haul" connections on top, dense "city street" connections below. A query lands on the top layer, greedily hops toward the target, then descends layer by layer to finer and finer resolution. Excellent recall/latency, insertion-friendly; the cost is RAM — roughly 1.5x the vector bytes in link overhead. Two dials: M (connections per node — how many streets) and efSearch (beam width at query time — how many candidates you keep in play while descending).
  • IVF-PQ (quantization-based): first cluster the vectors with k-means into geographic cells; a query probes only nprobe nearby cells instead of the whole map. Then product quantization (PQ) compresses each vector from hundreds of bytes down to a few bytes by storing "which cluster-centroid is each block closest to." Massive memory savings at billion scale, lower recall. Common via Faiss/ScaNN and DiskANN-style SSD variants.

03.A Simple Worked Example: Turning the Dial

Same corpus, same true answer, four settings — watch the three-way dial move.

code
method              latency   recall@10   RAM
brute force (exact)   120 ms     100%     610 GB
HNSW M=16 ef=40        8 ms      88%     900 GB
HNSW M=16 ef=200      35 ms      98%     900 GB
IVF-PQ nprobe=8      12 ms      91%     160 GB

Read the rows like a purchase decision:

  • Brute force is perfect but slow and hot — fine at 100k vectors, dead at 100M.
  • Widening the beam (efSearch 40 → 200) buys recall with latency: you explore more candidate nodes during descent, approaching exact search while spending more CPU per query. Note that M and ef_construction are build-time dials; efSearch is the query-time one.
  • IVF-PQ trades recall for a 4x smaller memory bill — the cell-probing + compression route wins when RAM is the binding constraint.

Now add the filter. Query: "top-10, tenant = acme." If acme holds 50k of the 100M vectors, the planner can simply pre-filter and brute-force the 50k matching subset — exact, fast, no ANN compromise at all. Filter selectivity (what fraction survives) decides the best strategy. That is why filtering is not a bolt-on feature; it changes the search algorithm itself.

04.Visual Intuition: Highways, Then City Streets

HNSW descent, top layer coarse, bottom layer dense:

code
layer 2 (few nodes):   A ─────────── F            land anywhere,
                       │             │            hop far and fast
layer 1:               A ── C ── F──┤            refine
                       │    │    │  │
layer 0 (all nodes):   A─B─C─D─E─F─G─● ← target   crawl streets

IVF: the map is divided into cells; you only open a few.

code
  ┌───────┬───────┬───────┐
  │ cell1 │ cell2 │▓cell3 │   query lands near cell3,
  ├───────┼───────┼───────┤   probes ▓ cells only:
  │ cell4 │▓cell5 │ cell6 │   5 and 3 (nprobe = 2)
  ├───────┼───────┼───────┤
  │ cell7 │ cell8 │ cell9 │   ← other 7 cells skipped
  └───────┴───────┴───────┘

The filtered-query path in one line:

Insight

planner checks filter selectivity → pre-filter + exact scan if small, HNSW descent with filter bitmaps if large → verify metadata → global top-K.

05.The Analogy: A Warehouse Library with a Concierge

Carry one analogy through: your corpus is a warehouse-sized library, and the vector DB is a concierge who has walked it a million times.

  • Exact kNN = the concierge personally checking all 100 million books for each visitor. Perfect. Unemployable.
  • HNSW = her mental shortcuts: "the philosophy wing connects to law via three doors." She races through the long-haul corridors (top layers), then browses shelves (base layer). M = how many shortcuts she memorized per room; efSearch = how many shelves she re-checks before deciding.
  • IVF-PQ = the library divided into labeled rooms (clusters); she enters only 2 rooms near your request. PQ means even the catalog cards are compressed — "Room 4, block B, near the window" instead of a full description of each book.
  • Quantization = writing addresses in pencil shorthand: int8 fits 4 books per line of fp32; binary fits 32 — great until you must verify, which is the rescoring pass (fetch a few originals in full precision).
  • Metadata filters = "restricted section, keycard required." A good concierge never even shows you titles from a section you cannot enter (pre-filter / in-graph filtering). A naive one grabs the 20 best books first and then notices you were not allowed 12 of them — and hands you 8. That is post-filtering, and it is a classic bug.
  • Tombstones/compaction = withdrawn books leave stubs in the catalog; periodically she re-shelves everything, or the map rots.

The whole product decision reduces to: hire the right kind of concierge for the size of your library.

06.The Landscape: Specialized Engines vs Postgres Extensions

Pinecone — fully managed, serverless (compute and storage separated since 2024), namespaces for tenant isolation, named indexes with metadata filtering. Least operational work; highest lock-in and per-query cost.

Weaviate — open-source (BSL license) with a GraphQL/REST API, built-in hybrid BM25 + vector fusion (topic 165), vectorizer modules, multi-tenancy; a strong turnkey RAG backend.

Chroma — lightweight, embedded-first (Python/JS in-process or client/server), the default teaching/prototype store; HNSW under the hood; adequate to tens of millions of docs.

pgvector — a Postgres extension: exact and HNSW/IVFFlat indexes over vector/halfvec/bit columns. Your embeddings live next to your relational rows, so ACL joins, transactions, and operational tooling are all native. Since the 0.7/0.8 releases (halfvec + iterative index scans) it credibly serves hundreds of millions of vectors.

Also relevant: Qdrant (Rust, excellent filtering), Milvus (billions-scale, distributed), and Redis or SQLite-VSS for edge deployments.

The SQL below is the whole point of pgvector in one screen: similarity search, tenant filter, and freshness filter in one ordinary query on the database you already run.

sql— pgvector: hybrid-ish filtered semantic search inside ordinary Postgres
CREATE EXTENSION vector;

CREATE TABLE chunks (
  id        bigserial PRIMARY KEY,
  doc_id    text   NOT NULL,
  tenant_id uuid   NOT NULL,
  body      text   NOT NULL,
  embedding halfvec(1024) NOT NULL
);

-- HNSW on halfvec: ~2x smaller than float32, negligible recall loss
CREATE INDEX ON chunks
  USING hnsw (embedding halfvec_cosine_distance_ops)
  WITH (m = 16, ef_construction = 64);

SELECT id, doc_id, 1 - (embedding <=> $1::halfvec) AS score
FROM chunks
WHERE tenant_id = $2                      -- filtered search, one query
  AND updated_at > now() - interval '90 days'
ORDER BY embedding <=> $1::halfvec
LIMIT 20;

07.Metadata Filtering: The Feature That Selects the Product

RAG queries are almost never pure vector search. "Only docs this user may read, updated after 2024, in product X" is the normal shape — and filter-then-search vs search-then-filter changes both results and latency:

  • Pre-filter (brute force): when filters leave a small candidate set, scan it exactly — fast and high-recall. This is what the planner in section 3 chose for tenant=acme's 50k vectors.
  • Filter-in-graph: traverse HNSW but respect filter bitmaps while descending. pgvector historically degraded badly when filtered rows sat outside the explored graph; iterative index scans (0.8, 2024) fixed the pathology. Qdrant/Weaviate/Pinecone attach filter bitmaps to graph segments.
  • Post-filter: retrieve top-K, then drop disallowed rows — silently returns fewer than K results (or worse, returns them from the wrong tenant if you forget). A classic naive-implementation bug.

Tenant isolation patterns follow the same fork: namespaces (Pinecone), partitions/tenants (Weaviate, Qdrant payload partitioning), or row-level security (RLS) in Postgres — where per-tenant ACLs are just RLS policies, a huge compliance argument for pgvector.

08.In Practice: What Breaks and What to Measure

Three habits keep a vector DB honest:

  1. Measure recall against exact search. ANN always trades recall for speed, so dashboards lie unless you periodically re-run sampled queries as brute-force kNN and compute recall@k. That is the number behind every "tune efSearch" decision — and what ann-benchmarks.com compares.
  2. Watch the RAM bill and shrink it deliberately. The quantization ladder (fp32 → halfvec/int8 → binary + rescoring) plus Matryoshka truncation (topic 162: store only the first 256 dims for coarse recall, rescore at full size) cut memory 4–32x. Cost lives here, not in the license.
  3. Treat migration as a re-index event. Changing the embedding model invalidates every stored vector (topic 162), so store the model version beside each vector and plan dual-index cutovers.

And size the choice by volume, not prestige: under ~100k vectors, exact brute-force numpy in-process beats every dedicated system on simplicity, cost, and recall (100% by definition).

Architectural Trade-offs & Production Realities

Architectural Advantages

  • Sub-100 ms semantic recall over hundreds of millions of vectors with tuned HNSW.
  • Managed engines (Pinecone) remove ops entirely; pgvector removes a whole system from the architecture.
  • Quantization + MRL truncation (topic 162) cut RAM costs 4–32x with rescoring recovering accuracy.

Trade-offs & Constraints

  • ANN always trades recall for speed — dashboards lie unless recall@k is measured against exact search.
  • Specialized engines add a stateful system: replicas, segment compaction, version migrations.
  • Managed pricing (per-query/per-unit) can dominate costs at high QPS versus self-hosted Qdrant/pgvector.
Production Implementation in Big Tech
Many YC-era SaaS teams (pattern, ~2024–2026)• Per-tenant knowledge RAG shipped on Postgres

An app already storing documents in Postgres adds a pgvector column, embeds chunks with a hosted 1024-d model, and creates an HNSW index with halfvec. Tenant isolation rides existing row-level security; the entire RAG read path is one SQL query — no new infra until vector volume or QPS justifies extraction to Qdrant.

Staff+ Engineering Takeaways

  • Vector DBs answer ANN queries: recall vs latency vs memory is the fundamental three-way dial.
  • HNSW is the default index (layered shortcut graph, dials M and efSearch); IVF-PQ/quantization win at extreme scale or tight RAM budgets.
  • Filtered search correctness (pre-filter vs post-filter, iterative scans) is the make-or-break RAG feature.
  • pgvector keeps vectors transactional with your data; managed engines trade cost for zero ops.
  • Under ~100k vectors, exact brute-force search beats any dedicated system on simplicity and recall.

Topic Knowledge Check

Exercise 1 of 3 • Test your architectural comprehension.

Exercise 1 of 30 answered
1

Why does increasing efSearch in HNSW raise recall?

Rate This Architecture ChapterFeedback & Rating

How clear and actionable was this distributed systems breakdown?