Distributed File System Block Lab (Interactive)
Feed millions of tiny files into HDFS and watch NameNode RAM exhaust while big blocks amortize seeks. Compute block descriptors, NameNode heap pressure, 3x-replicated disk totals, and mechanical seek overhead for different file sizes and 64/128/256 MB block choices.
HDFS NameNode Memory & Block Sizing
Watch 150-byte metadata records and disk-seek amortization react to file shape.
How It Works Under the Hood
GFS and HDFS assume commodity failure and gigantic sequential files, so they standardize on 128-256 MB blocks. One 10 ms arm seek preceding a 128 MB sequential transfer at 150 MB/s costs under 1.2% overhead, while the same seek for a 4 KB block wastes the entire disk. The NameNode keeps every file and block record, roughly 150 bytes each, in DRAM: a cluster can hold hundreds of petabytes on disks yet die from a hundred million one-kilobyte files, the infamous Small Files Problem.
Core Architectural Principles
- Each file and block object consumes about 150 bytes of NameNode RAM regardless of size.
- A 1 TB file is 8,192 block descriptors at 128 MB versus 268 million at 4 KB.
- Rack-aware placement puts replicas on two same-rack nodes plus one remote rack, streamed pipelined.
Explain control-flow versus data-flow separation: clients fetch metadata from the NameNode but stream bytes directly to DataNodes. Justify the 128 MB block with seek-amortization math, then diagnose the Small Files Problem and remediation: Parquet compaction, SequenceFiles, HAR archives, or moving metadata scale onto cloud object stores.
Huge blocks maximize streaming throughput and shrink master state but forbid low-latency random in-place edits.