SFU Media Topology Lab (Interactive)
Move participants and see who chokes: mesh uplink, MCU CPU, or SFU downlink. Compare mesh, MCU, and SFU media routing on client uplink, downlink, and server CPU, adding simulcast layer selection.
Zoom/Meet Media Topology: Mesh vs MCU vs SFU
Move the participant slider and watch who chokes — the client uplink, the server CPU, or the edge bandwidth.
Downlink O(N) at the client + edge SFU bandwidth — but server CPU stays near zero
glass-to-glass latency < 150 ms (UDP/SRTP, no HOL)
P2P mesh is elegant for 1:1 but each client must upload a distinct copy to every peer — at 25 people that is 1.5 Mbps up per laptop, why real group calls centralize. MCU saves client bandwidth (1 up, 1 composite down) but taxes the server to decode and re-encode every stream, spiking cost and adding ~500 ms of latency. The SFU wins by routing packets without touching them: one upload per client, near-zero server CPU, and simulcast lets it forward the speaker in 1080p while downgrading 49 thumbnails to 180p — cutting downstream bandwidth ~90% while staying under the sub-150 ms budget because UDP drops a stale frame instead of stalling on head-of-line retransmit.
How It Works Under the Hood
Multi-party video has three routing topologies, each moving the bottleneck. In a P2P mesh every client uploads a distinct copy to every peer, so uplink scales O(N) and busts home bandwidth past a handful of people. An MCU centralizes decode-and-re-encode into one composite, sparing clients but burning server CPU and adding roughly 500 ms. The SFU forwards packets untouched — one upload per client, near-zero CPU — but shifts cost to client downlink and edge bandwidth, which simulcast fixes by uploading layered 1080p, 720p, and 180p streams so the SFU sends thumbnails to non-speakers, cutting downstream about 90% while UDP keeps glass-to-glass under 150 ms.
Core Architectural Principles
- Mesh per-client uplink = (N-1) x bitrate, blowing past the roughly 2.5 Mbps home budget fast.
- SFU: one upload per client, near-zero server CPU, and download = speaker at high plus thumbnails at low with simulcast.
- UDP/SRTP drops a stale frame instead of head-of-line retransmit, holding latency under 150 ms.
Compare all three topologies explicitly and justify the SFU as the industry winner: no re-encode, so cost and latency stay low, and simulcast handles heterogeneous client bandwidth. Explain simulcast versus SVC layering and active-speaker detection via voice activity. Mention TURN for NAT traversal and inter-SFU cascading so a large multi-region call crosses the ocean once per speaker, not once per listener.
SFU minimizes server CPU and scales to hundreds, but every client still downloads N-1 streams and decodes many thumbnails.