Home/Labs/MapReduce Shuffle & Skew
All 280 Labs
INTERACTIVE LAB🗺️

MapReduce Shuffle Lab (Interactive)

Tune mapper and reducer counts, toggle combiners, and inject a hot key to expose the straggler bottleneck. Model the all-to-all shuffle of a word-count job: hash partitioning, combiner pre-aggregation savings, and how one viral key stalls the whole cluster on a single reducer NIC.

MapReduce Shuffle & Skew Lab

Tune M/R task counts, combiners, and a hot key to expose the network all-to-all bottleneck.

Reducer network fetch (share of total shuffle bytes)

Uniform hash distribution: every reducer pulls an equal slice.

Map output
19.53 GB
before local aggregation
Shuffle across network
2.79 GB
combiner saved 86%
Slowest reducer load
0.09 GB
1.0x the average of 0.09 GB
Estimated job time
4 s
dominated by bottleneck reducer NIC
Classic Hadoop writes every map output to local disk at the 80% spill threshold and re-reads it during the reducer's all-to-all fetch — the stage boundary is a disk barrier. Spark collapses narrow-dependence stages into RAM pipelines, but its hash exchange still shuffles; skew mitigation (salting, adaptive partitioning) applies to both engines.

How It Works Under the Hood

MapReduce splits computation into map emits and reduce aggregations joined by the shuffle: a hash partitioner routes every identical key to one reducer, and each reducer fetches its slice from all mappers over the network after map outputs spill to disk at the stage barrier. That disk-and-network exchange is the most expensive phase in distributed computing. A combiner collapses duplicate keys locally per mapper, cutting shuffle bytes, while a hot key such as NULL concentrates most traffic on one reducer, leaving the rest of the fleet idle and stretching job completion time.

Core Architectural Principles

  • hash(key) mod R routes identical keys to a single reducer, serializing hot-key workloads.
  • Combiners act as mini-reducers on each mapper, shrinking spill bytes before the network phase.
  • Job completion is bounded by the slowest reducer, not the average reducer.
Interview Round Script

State that shuffle is the expensive phase and that optimizing big-data jobs means minimizing wide dependencies. When skew appears, propose salting hot keys with random prefixes and two-stage aggregation. Contrast Hadoop disk barriers with Spark in-memory DAGs and lineage recovery to show you understand why the engine changed but partitioning physics did not.

Key Trade-Offs

Disk-spilled stage barriers buy fault tolerance and bounded memory at extreme I/O and latency cost.

Related Curriculum Chapter

The MapReduce Paradigm: Distributed Data Processing Foundations

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs