Managed Cloud Primitives TCO Lab (Interactive)
Balance managed markup against SRE salaries and downtime loss to find where self-hosting actually pays off. Model total cost of ownership for Aurora, MSK or S3 versus self-hosting, across state size, request rate and markup.
Managed vs Self-Hosted: Full TCO Ledger
TCO = cloud invoices + SRE salaries + downtime revenue loss. Move the scale sliders and find the crossover.
Self-hosted: PostgreSQL on EC2 (Patroni + pgBackRest)
raw infra $1,900/mo
SRE/DBA pager team (2) $27,500/mo
downtime loss @99.5% DIY SLA $4,380/mo
TCO $33,780/mo · 7 undifferentiated eng-hours/wk
Managed: Amazon Aurora
infra incl. 35% markup $2,565/mo
pager team $0 (provider on-calls)
downtime loss @99.99% multi-AZ $88/mo
TCO $2,653/mo · 0 ops hours
Managed Amazon Aurora is cheaper than self-hosting by $31,127/mo — the markup buys 24/7 on-call, multi-AZ replication and automated failover for free.
The pragmatic stack the industry converged on: keep stateless compute in portable OCI containers (Kubernetes/ECS) so the business logic never locks in, and delegate the stateful hard parts — relational db — to managed primitives. Aurora replicates 6 copies of your WAL across 3 AZs with sub-10s failover; that is undifferentiated heavy lifting you bought, not built. The lock-in you pay is real but rare (cloud migrations happen every 7-12 years), while the velocity tax of cloud-agnostic lowest-common-denominator abstractions is paid every single day.
How It Works Under the Hood
The build-versus-buy calculus rests on total cost of ownership, not the raw compute invoice: TCO equals cloud bills plus SRE and DBA salaries plus downtime revenue loss. Managed primitives like Aurora, DynamoDB, SQS and S3 absorb undifferentiated heavy lifting, the multi-AZ replication, automated failover, patching and backup drills that would otherwise demand a 24/7 pager team. Self-hosting on raw EC2 looks cheaper per unit but is only justified at hyperscale, the way Dropbox Magic Pocket or Netflix Open Connect beat public-cloud margins. The pragmatic pattern keeps stateless compute portable in containers while delegating stateful engines to managed services.
Core Architectural Principles
- TCO = infrastructure + human SRE/DBA salaries + downtime revenue loss, not compute alone.
- Managed services deliver 99.99% multi-AZ availability and automated failover you would otherwise build.
- Self-hosting beats the markup only at extreme petabyte/QPS scale like Dropbox and Netflix.
In system design interviews, reach freely for managed primitives (S3, SQS, Aurora, DynamoDB) rather than reinventing replication, and justify it with the TCO equation: add salaries and downtime to the invoice. Invoke undifferentiated heavy lifting to explain why engineering time is better spent on business features. Close with the pragmatic split: stateless containerized compute for portability, managed stateful services for reliability, and acknowledge self-hosting only at hyperscale.
Managed primitives trade higher hourly markup and vendor lock-in for eliminated ops burden and out-of-the-box high availability.