Vector Databases: Pinecone, Chroma, Weaviate, pgvector
A vector database stores embedding coordinates and answers "what is nearest to this query?" fast enough for live traffic. It does so with approximate indexes (HNSW graphs, IVF-PQ clusters) that trade a little accuracy for 100–1000x speed, plus metadata filtering — the feature that decides whether your RAG is usable or a data leak.
Filtered ANN Query Path in a Vector Database
Modern engines interleave metadata filtering with graph traversal; filter selectivity decides pre-filtering vs in-graph filtering.
01.The Problem: You Cannot Check Every Shelf
Topic 162 gave every text a coordinate — an address in meaning-space.
Now a user asks a question, and you need the 20 closest addresses out of your whole corpus.
The honest way — exact k-nearest neighbors (kNN) — compares the query against all N vectors. Cost per query: O(N·d) multiplications.
Put in numbers:
- 100 million vectors × 1024 dimensions ≈ 100 billion multiply-adds per query.
- Even one user gets 100+ ms of pure arithmetic; a thousand users hit the same index at once.
Interactive search needs answers in tens of milliseconds. So exactness has to go.
But wait — there is a second problem. A real RAG question is rarely just "find similar."
"Only documents this user may read, updated after 2024, in product X."
Speed and filters. That combination — a storage engine for coordinates plus fast approximate search plus structured filtering — is what a vector database exists to solve.
02.The Idea in Plain Words: The ANN Bargain
Every vector database delivers
Approximate Nearest Neighbor (ANN) search: find the top-K results most of the time, 100–1000x faster than checking everything.
The bargain is stated precisely:
- Recall = the fraction of the true nearest neighbors your fast search actually found. Return 9 of the real top-10 → recall@10 = 90%.
- Latency = how long one query takes.
- Memory/cost = how many bytes of RAM (or disk) the index needs.
You cannot max all three. Tuning any vector DB is turning this three-way dial — that is the entire engineering contract.
The two dominant index families, in plain words:
- HNSW (graph-based, 2016 → the 2024–2026 default): build a network of shortcuts between vectors, layered like an airport map: a few far-apart "long-haul" connections on top, dense "city street" connections below. A query lands on the top layer, greedily hops toward the target, then descends layer by layer to finer and finer resolution. Excellent recall/latency, insertion-friendly; the cost is RAM — roughly 1.5x the vector bytes in link overhead. Two dials:
M(connections per node — how many streets) andefSearch(beam width at query time — how many candidates you keep in play while descending). - IVF-PQ (quantization-based): first cluster the vectors with k-means into geographic cells; a query probes only
nprobenearby cells instead of the whole map. Then product quantization (PQ) compresses each vector from hundreds of bytes down to a few bytes by storing "which cluster-centroid is each block closest to." Massive memory savings at billion scale, lower recall. Common via Faiss/ScaNN and DiskANN-style SSD variants.
03.A Simple Worked Example: Turning the Dial
Same corpus, same true answer, four settings — watch the three-way dial move.
codemethod latency recall@10 RAM brute force (exact) 120 ms 100% 610 GB HNSW M=16 ef=40 8 ms 88% 900 GB HNSW M=16 ef=200 35 ms 98% 900 GB IVF-PQ nprobe=8 12 ms 91% 160 GB
Read the rows like a purchase decision:
- Brute force is perfect but slow and hot — fine at 100k vectors, dead at 100M.
- Widening the beam (
efSearch40 → 200) buys recall with latency: you explore more candidate nodes during descent, approaching exact search while spending more CPU per query. Note thatMandef_constructionare build-time dials;efSearchis the query-time one. - IVF-PQ trades recall for a 4x smaller memory bill — the cell-probing + compression route wins when RAM is the binding constraint.
Now add the filter. Query: "top-10, tenant = acme." If acme holds 50k of the 100M vectors, the planner can simply pre-filter and brute-force the 50k matching subset — exact, fast, no ANN compromise at all. Filter selectivity (what fraction survives) decides the best strategy. That is why filtering is not a bolt-on feature; it changes the search algorithm itself.
04.Visual Intuition: Highways, Then City Streets
HNSW descent, top layer coarse, bottom layer dense:
codelayer 2 (few nodes): A ─────────── F land anywhere, │ │ hop far and fast layer 1: A ── C ── F──┤ refine │ │ │ │ layer 0 (all nodes): A─B─C─D─E─F─G─● ← target crawl streets
IVF: the map is divided into cells; you only open a few.
code┌───────┬───────┬───────┐ │ cell1 │ cell2 │▓cell3 │ query lands near cell3, ├───────┼───────┼───────┤ probes ▓ cells only: │ cell4 │▓cell5 │ cell6 │ 5 and 3 (nprobe = 2) ├───────┼───────┼───────┤ │ cell7 │ cell8 │ cell9 │ ← other 7 cells skipped └───────┴───────┴───────┘
The filtered-query path in one line:
planner checks filter selectivity → pre-filter + exact scan if small, HNSW descent with filter bitmaps if large → verify metadata → global top-K.
05.The Analogy: A Warehouse Library with a Concierge
Carry one analogy through: your corpus is a warehouse-sized library, and the vector DB is a concierge who has walked it a million times.
- Exact kNN = the concierge personally checking all 100 million books for each visitor. Perfect. Unemployable.
- HNSW = her mental shortcuts: "the philosophy wing connects to law via three doors." She races through the long-haul corridors (top layers), then browses shelves (base layer).
M= how many shortcuts she memorized per room;efSearch= how many shelves she re-checks before deciding. - IVF-PQ = the library divided into labeled rooms (clusters); she enters only 2 rooms near your request. PQ means even the catalog cards are compressed — "Room 4, block B, near the window" instead of a full description of each book.
- Quantization = writing addresses in pencil shorthand: int8 fits 4 books per line of fp32; binary fits 32 — great until you must verify, which is the rescoring pass (fetch a few originals in full precision).
- Metadata filters = "restricted section, keycard required." A good concierge never even shows you titles from a section you cannot enter (pre-filter / in-graph filtering). A naive one grabs the 20 best books first and then notices you were not allowed 12 of them — and hands you 8. That is post-filtering, and it is a classic bug.
- Tombstones/compaction = withdrawn books leave stubs in the catalog; periodically she re-shelves everything, or the map rots.
The whole product decision reduces to: hire the right kind of concierge for the size of your library.
06.The Landscape: Specialized Engines vs Postgres Extensions
Pinecone — fully managed, serverless (compute and storage separated since 2024), namespaces for tenant isolation, named indexes with metadata filtering. Least operational work; highest lock-in and per-query cost.
Weaviate — open-source (BSL license) with a GraphQL/REST API, built-in hybrid BM25 + vector fusion (topic 165), vectorizer modules, multi-tenancy; a strong turnkey RAG backend.
Chroma — lightweight, embedded-first (Python/JS in-process or client/server), the default teaching/prototype store; HNSW under the hood; adequate to tens of millions of docs.
pgvector — a Postgres extension: exact and HNSW/IVFFlat indexes over vector/halfvec/bit columns. Your embeddings live next to your relational rows, so ACL joins, transactions, and operational tooling are all native. Since the 0.7/0.8 releases (halfvec + iterative index scans) it credibly serves hundreds of millions of vectors.
Also relevant: Qdrant (Rust, excellent filtering), Milvus (billions-scale, distributed), and Redis or SQLite-VSS for edge deployments.
The SQL below is the whole point of pgvector in one screen: similarity search, tenant filter, and freshness filter in one ordinary query on the database you already run.
CREATE EXTENSION vector;
CREATE TABLE chunks (
id bigserial PRIMARY KEY,
doc_id text NOT NULL,
tenant_id uuid NOT NULL,
body text NOT NULL,
embedding halfvec(1024) NOT NULL
);
-- HNSW on halfvec: ~2x smaller than float32, negligible recall loss
CREATE INDEX ON chunks
USING hnsw (embedding halfvec_cosine_distance_ops)
WITH (m = 16, ef_construction = 64);
SELECT id, doc_id, 1 - (embedding <=> $1::halfvec) AS score
FROM chunks
WHERE tenant_id = $2 -- filtered search, one query
AND updated_at > now() - interval '90 days'
ORDER BY embedding <=> $1::halfvec
LIMIT 20;07.Metadata Filtering: The Feature That Selects the Product
RAG queries are almost never pure vector search. "Only docs this user may read, updated after 2024, in product X" is the normal shape — and filter-then-search vs search-then-filter changes both results and latency:
- Pre-filter (brute force): when filters leave a small candidate set, scan it exactly — fast and high-recall. This is what the planner in section 3 chose for tenant=acme's 50k vectors.
- Filter-in-graph: traverse HNSW but respect filter bitmaps while descending. pgvector historically degraded badly when filtered rows sat outside the explored graph;
iterative index scans(0.8, 2024) fixed the pathology. Qdrant/Weaviate/Pinecone attach filter bitmaps to graph segments. - Post-filter: retrieve top-K, then drop disallowed rows — silently returns fewer than K results (or worse, returns them from the wrong tenant if you forget). A classic naive-implementation bug.
Tenant isolation patterns follow the same fork: namespaces (Pinecone), partitions/tenants (Weaviate, Qdrant payload partitioning), or row-level security (RLS) in Postgres — where per-tenant ACLs are just RLS policies, a huge compliance argument for pgvector.
08.In Practice: What Breaks and What to Measure
Three habits keep a vector DB honest:
- Measure recall against exact search. ANN always trades recall for speed, so dashboards lie unless you periodically re-run sampled queries as brute-force kNN and compute recall@k. That is the number behind every "tune efSearch" decision — and what ann-benchmarks.com compares.
- Watch the RAM bill and shrink it deliberately. The quantization ladder (fp32 → halfvec/int8 → binary + rescoring) plus Matryoshka truncation (topic 162: store only the first 256 dims for coarse recall, rescore at full size) cut memory 4–32x. Cost lives here, not in the license.
- Treat migration as a re-index event. Changing the embedding model invalidates every stored vector (topic 162), so store the model version beside each vector and plan dual-index cutovers.
And size the choice by volume, not prestige: under ~100k vectors, exact brute-force numpy in-process beats every dedicated system on simplicity, cost, and recall (100% by definition).
Architectural Trade-offs & Production Realities
Architectural Advantages
- Sub-100 ms semantic recall over hundreds of millions of vectors with tuned HNSW.
- Managed engines (Pinecone) remove ops entirely; pgvector removes a whole system from the architecture.
- Quantization + MRL truncation (topic 162) cut RAM costs 4–32x with rescoring recovering accuracy.
Trade-offs & Constraints
- ANN always trades recall for speed — dashboards lie unless recall@k is measured against exact search.
- Specialized engines add a stateful system: replicas, segment compaction, version migrations.
- Managed pricing (per-query/per-unit) can dominate costs at high QPS versus self-hosted Qdrant/pgvector.
An app already storing documents in Postgres adds a pgvector column, embeds chunks with a hosted 1024-d model, and creates an HNSW index with halfvec. Tenant isolation rides existing row-level security; the entire RAG read path is one SQL query — no new infra until vector volume or QPS justifies extraction to Qdrant.
Staff+ Engineering Takeaways
- Vector DBs answer ANN queries: recall vs latency vs memory is the fundamental three-way dial.
- HNSW is the default index (layered shortcut graph, dials M and efSearch); IVF-PQ/quantization win at extreme scale or tight RAM budgets.
- Filtered search correctness (pre-filter vs post-filter, iterative scans) is the make-or-break RAG feature.
- pgvector keeps vectors transactional with your data; managed engines trade cost for zero ops.
- Under ~100k vectors, exact brute-force search beats any dedicated system on simplicity and recall.
Topic Knowledge Check
Exercise 1 of 3 • Test your architectural comprehension.
Why does increasing efSearch in HNSW raise recall?
How clear and actionable was this distributed systems breakdown?