Home/Labs/Managed vs Self-Hosted TCO
All 280 Labs
INTERACTIVE LAB☁️

Managed Cloud Primitives TCO Lab (Interactive)

Balance managed markup against SRE salaries and downtime loss to find where self-hosting actually pays off. Model total cost of ownership for Aurora, MSK or S3 versus self-hosting, across state size, request rate and markup.

Managed vs Self-Hosted: Full TCO Ledger

TCO = cloud invoices + SRE salaries + downtime revenue loss. Move the scale sliders and find the crossover.

Self-hosted: PostgreSQL on EC2 (Patroni + pgBackRest)

raw infra $1,900/mo

SRE/DBA pager team (2) $27,500/mo

downtime loss @99.5% DIY SLA $4,380/mo

TCO $33,780/mo · 7 undifferentiated eng-hours/wk

Managed: Amazon Aurora

infra incl. 35% markup $2,565/mo

pager team $0 (provider on-calls)

downtime loss @99.99% multi-AZ $88/mo

TCO $2,653/mo · 0 ops hours

Managed Amazon Aurora is cheaper than self-hosting by $31,127/mo — the markup buys 24/7 on-call, multi-AZ replication and automated failover for free.

The pragmatic stack the industry converged on: keep stateless compute in portable OCI containers (Kubernetes/ECS) so the business logic never locks in, and delegate the stateful hard parts — relational db — to managed primitives. Aurora replicates 6 copies of your WAL across 3 AZs with sub-10s failover; that is undifferentiated heavy lifting you bought, not built. The lock-in you pay is real but rare (cloud migrations happen every 7-12 years), while the velocity tax of cloud-agnostic lowest-common-denominator abstractions is paid every single day.

How It Works Under the Hood

The build-versus-buy calculus rests on total cost of ownership, not the raw compute invoice: TCO equals cloud bills plus SRE and DBA salaries plus downtime revenue loss. Managed primitives like Aurora, DynamoDB, SQS and S3 absorb undifferentiated heavy lifting, the multi-AZ replication, automated failover, patching and backup drills that would otherwise demand a 24/7 pager team. Self-hosting on raw EC2 looks cheaper per unit but is only justified at hyperscale, the way Dropbox Magic Pocket or Netflix Open Connect beat public-cloud margins. The pragmatic pattern keeps stateless compute portable in containers while delegating stateful engines to managed services.

Core Architectural Principles

  • TCO = infrastructure + human SRE/DBA salaries + downtime revenue loss, not compute alone.
  • Managed services deliver 99.99% multi-AZ availability and automated failover you would otherwise build.
  • Self-hosting beats the markup only at extreme petabyte/QPS scale like Dropbox and Netflix.
Interview Round Script

In system design interviews, reach freely for managed primitives (S3, SQS, Aurora, DynamoDB) rather than reinventing replication, and justify it with the TCO equation: add salaries and downtime to the invoice. Invoke undifferentiated heavy lifting to explain why engineering time is better spent on business features. Close with the pragmatic split: stateless containerized compute for portability, managed stateful services for reliability, and acknowledge self-hosting only at hyperscale.

Key Trade-Offs

Managed primitives trade higher hourly markup and vendor lock-in for eliminated ops burden and out-of-the-box high availability.

Related Curriculum Chapter

Managed Cloud Primitives as Architectural Building Blocks

Read Full Chapter Blueprint

Explore More Interactive Labs

View All 280 Labs