NoSQL Workload Fit Lab (Interactive)
Race four workloads against KV, document, column-family, and graph engines. Sessions, catalog reads, IoT ingest, and friend-of-friend queries compute latency, RAM, and storage per NoSQL family, revealing which data shape each engine actually matches.
The 4 NoSQL Families vs Real Workloads
Compute what each storage engine actually pays for the same query — there is no best NoSQL, only best fit.
3-hop traversal from Alice over a graph with avg degree 30: frontier ≈ 2.7e+4 candidates.
Polyglot persistence: none of this says "drop PostgreSQL". Production shapes the trio — relational source of truth, CDC/Kafka fans changes out to Redis (sessions), Elasticsearch (search), and a graph or document store (relationships) — each family answering only the query it is physically built for.
How It Works Under the Hood
The NoSQL families are access-pattern specializations, not flavors. Key-value gives O(1) gets with TTLs — perfect for session caches, punishing for large values billed in RAM. Documents store nested JSON that matches read-an-entity workloads like product pages without join assembly. Column-family stores ride LSM-tree append speed for massive time-series and IoT write velocity. Graph engines follow index-free adjacency pointers so multi-hop relationship cost scales with the local neighborhood (degree^hops) while SQL pays N self-joins. Fit the data shape to the query shape.
Core Architectural Principles
- Key-value O(1) lookups and TTL win for session caches, but RAM pricing punishes large per-key values.
- Column-family LSM append speed absorbs IoT write rates that row-store fsync paths cannot.
- Graph traversals scale with degree^hops over the visited subgraph while SQL friend-of-friend explodes with N self-joins.
Pick storage by access pattern, not popularity: get-by-key, whole-entity reads, append-heavy time series, or multi-hop relationships. Name the matching family plus one honest con — Redis RAM cost, MongoDB join limits, Neo4j operational maturity. Never list four databases without tying each to the workload shape it exists for.
Each NoSQL family buys its ideal workload speed by giving up generality — a wrong-fit workload costs more than plain SQL would have.