Home/Labs/Iceberg Metadata Planner
All 280 Labs
INTERACTIVE LAB🏔️

Data Lake & Lakehouse Lab (Interactive)

Race Hive directory listings against Iceberg manifest pruning and stress OCC catalog commits. Size the query-planning gap between Hive Metastore S3 LIST pagination and Apache Iceberg snapshot-manifest metadata, then add concurrent writers to see atomic Compare-And-Swap commits.

Lakehouse Metadata Planning: Hive vs Iceberg

See how the manifest tree replaces S3 LIST storms and enables ACID commits via OCC.

Query planning time
348 ms
metadata tree read only
Files opened by scan
8,000
of 400,000 total
Data actually scanned
1562.5 GB
200 MB avg file · 98.0% pruned
Atomic commit latency
195 ms
catalog CAS, 3 retry round(s)
The planner resolves the current snapshot pointer, reads the manifest list, and prunes entire manifests with min/max column bounds — files are resolved from Avro metadata, never from directory listings. Readers stay on consistent snapshots while writers swap the pointer atomically.

How It Works Under the Hood

A raw data lake on S3 has no transactional truth: Hive maps partitions to directories, so planners issue thousands of paginated LIST calls and failed ETL jobs leave orphaned files readers can stumble into. Apache Iceberg replaces directories with an immutable metadata tree: the catalog pointer resolves a snapshot, a manifest list points to manifests carrying per-file column min/max bounds, and the planner prunes data files from metadata alone. Commits swap the pointer atomically with optimistic concurrency, giving snapshot isolation and time travel over open Parquet files.

Core Architectural Principles

  • S3 LIST pagination costs roughly one round trip per 1,000 keys before any scan starts.
  • Manifest files store exact Parquet URIs plus min/max bounds, enabling metadata-level file pruning.
  • OCC writers stage immutable files, then race a catalog Compare-And-Swap, retrying on conflict.
Interview Round Script

For petabyte analytics designs, propose Iceberg or Delta over raw Hive and explain the mechanism: manifests eliminate LIST storms and provide atomic multi-file commits. Mention optimistic concurrency at the catalog pointer, time travel via historical snapshots, and multi-engine interoperability where Spark writes and Trino or Snowflake query the same Parquet without copying data.

Key Trade-Offs

Open ACID tables add metadata maintenance and commit-retry latency versus raw files or closed warehouse formats.

Related Curriculum Chapter

Data Lakes & Lakehouses: Apache Iceberg, Delta Lake, & Parquet

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs