MapReduce Shuffle Lab (Interactive)
Tune mapper and reducer counts, toggle combiners, and inject a hot key to expose the straggler bottleneck. Model the all-to-all shuffle of a word-count job: hash partitioning, combiner pre-aggregation savings, and how one viral key stalls the whole cluster on a single reducer NIC.
MapReduce Shuffle & Skew Lab
Tune M/R task counts, combiners, and a hot key to expose the network all-to-all bottleneck.
Uniform hash distribution: every reducer pulls an equal slice.
How It Works Under the Hood
MapReduce splits computation into map emits and reduce aggregations joined by the shuffle: a hash partitioner routes every identical key to one reducer, and each reducer fetches its slice from all mappers over the network after map outputs spill to disk at the stage barrier. That disk-and-network exchange is the most expensive phase in distributed computing. A combiner collapses duplicate keys locally per mapper, cutting shuffle bytes, while a hot key such as NULL concentrates most traffic on one reducer, leaving the rest of the fleet idle and stretching job completion time.
Core Architectural Principles
- hash(key) mod R routes identical keys to a single reducer, serializing hot-key workloads.
- Combiners act as mini-reducers on each mapper, shrinking spill bytes before the network phase.
- Job completion is bounded by the slowest reducer, not the average reducer.
State that shuffle is the expensive phase and that optimizing big-data jobs means minimizing wide dependencies. When skew appears, propose salting hot keys with random prefixes and two-stage aggregation. Contrast Hadoop disk barriers with Spark in-memory DAGs and lineage recovery to show you understand why the engine changed but partitioning physics did not.
Disk-spilled stage barriers buy fault tolerance and bounded memory at extreme I/O and latency cost.