Data Lake & Lakehouse Lab (Interactive)
Race Hive directory listings against Iceberg manifest pruning and stress OCC catalog commits. Size the query-planning gap between Hive Metastore S3 LIST pagination and Apache Iceberg snapshot-manifest metadata, then add concurrent writers to see atomic Compare-And-Swap commits.
Lakehouse Metadata Planning: Hive vs Iceberg
See how the manifest tree replaces S3 LIST storms and enables ACID commits via OCC.
How It Works Under the Hood
A raw data lake on S3 has no transactional truth: Hive maps partitions to directories, so planners issue thousands of paginated LIST calls and failed ETL jobs leave orphaned files readers can stumble into. Apache Iceberg replaces directories with an immutable metadata tree: the catalog pointer resolves a snapshot, a manifest list points to manifests carrying per-file column min/max bounds, and the planner prunes data files from metadata alone. Commits swap the pointer atomically with optimistic concurrency, giving snapshot isolation and time travel over open Parquet files.
Core Architectural Principles
- S3 LIST pagination costs roughly one round trip per 1,000 keys before any scan starts.
- Manifest files store exact Parquet URIs plus min/max bounds, enabling metadata-level file pruning.
- OCC writers stage immutable files, then race a catalog Compare-And-Swap, retrying on conflict.
For petabyte analytics designs, propose Iceberg or Delta over raw Hive and explain the mechanism: manifests eliminate LIST storms and provide atomic multi-file commits. Mention optimistic concurrency at the catalog pointer, time travel via historical snapshots, and multi-engine interoperability where Spark writes and Trino or Snowflake query the same Parquet without copying data.
Open ACID tables add metadata maintenance and commit-retry latency versus raw files or closed warehouse formats.