DeepSeek V3 & R1: Open Research at Frontier Cost-Curve
DeepSeek, a Chinese hedge-fund-born lab, published two bombs: V3 — a 671B-parameter MoE (only 37B awake per token) claimed to cost about $5.6M to train, MIT-licensed — and R1, a reasoning model whose long "aha" chains of thought emerged from pure reinforcement learning against verifiable answers. Together they reset the industry's price/perception of what frontier-adjacent AI must cost.
01.The Problem: Frontier AI Looked Like a Billion-Dollar Club
In 2024 the accepted story was simple:
- world-class models require hundreds of millions of dollars of training compute,
- the labs that afford them keep the models closed,
- everyone else rents access through an API at the lab's price.
So the question practitioners asked was
Is frontier quality fundamentally expensive — or are we just watching one inefficient way of building it?
In December 2024, DeepSeek — a lab funded by a Chinese hedge fund, unknown outside China — answered by publishing everything: the model weights, the full technical paper, and the training cost. The weights were MIT-licensed (the most "do anything" license there is). The claimed number in the paper shook the market.
DeepSeek lineage and the ideas that mattered
DeepSeek lineage and the ideas that mattered
Each release paired a paper-worthy architecture innovation with aggressive cost efficiency, culminating in sparse-attention long-context serving economics.
Unlock Topic #308: DeepSeek V3 & R1: Open Research at Frontier Cost-Curve
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?