Transcoding Farm Throughput Lab (Interactive)
Size render wall-time from uploads, ladder rungs, and CPU vs GPU speedup. Model transcoding wall time from upload volume, ladder rungs, and fleet size, then contrast CPU against GPU throughput and live egress.
YouTube-style Transcoding Farm & ABR Ladder
Chop an upload into 6-second segments across a bitrate ladder and time the parallel DAG worker fleet.
› probe/decrypt → scene-detect → parallel chunk encode (each segment × each rendition is an independent DAG task) → mux → thumbnail → watermark → CDN push.
› Player switches renditions at segment boundaries from the manifest.m3u8/.mpd, targeting the newest buffer with ≤ 2s of bandwidth headroom.
› 720,000 hours/day uploads (YouTube) make transcoding a capacity fleet problem, not a request/response API.
The ABR ladder exists because networks fluctuate mid-playback: pre-encode every rung into 6-second chunks and let the client self-tune. Transcoding is embarrassingly parallel per chunk, so publish latency falls almost linearly with fleet size — until a serial stage (audio loudness or packaging) caps you. Netflix Open Connect appliances inside ISPs absorb the 250Gbps-per-million-viewer egress locally.
How It Works Under the Hood
Video platforms transcode every upload into an adaptive bitrate ladder, many renditions from 240p to 4K, so HLS and DASH can stream fragmented segments. The pipeline is embarrassingly parallel but throughput-bound: wall time equals minutes x rungs / (fleet x per-machine speed), and a GPU farm runs about twelve times faster than CPU's roughly three-times realtime. Segmenting into six-second chunks enables adaptive streaming and byte-range seeks. Delivery is the other cliff: concurrent viewers x their rendition bitrate is raw egress, easily reaching terabits/sec, which is why you ship it to a CDN rather than origin servers. Set upload volume, ladder depth, and fleet to watch render time and egress diverge.
Core Architectural Principles
- Wall time = upload_minutes x ladder_rungs / (fleet x speedup), where speedup is about 3x CPU or 12x GPU.
- Segmenting into 6-second chunks enables adaptive bitrate HLS/DASH and cheap byte-range caching.
- Live egress = concurrent_viewers x rendition_bitrate, scaling to terabits per second at the CDN edge.
Break it into ingest, transcode ladder, storage, delivery. Stress that transcoding is compute-bound and parallelizable with a queue plus elastic GPU workers, and that ladder depth is a quality and storage tradeoff. Then hit delivery economics: egress bandwidth dwarfs everything, so CDN with edge caching and segment-level hashing is mandatory. Mention resumable uploads and HLS segment layout.
Deeper bitrate ladders cut viewer bandwidth and stalls but multiply transcode compute and storage; GPUs speed render yet cost more per node.