Key ideas
- The one idea — DeepSeek-R1 proved reasoning emerges from RL on verifiable rewards — open weights, a famously lean price tag with honest caveats, and a one-day market shock measured in hundreds of billions.
- How it is shown — R1-Zero's chains lengthening as reward climbs; the "aha moment" from the paper; the Nvidia chart cliff; the caveat ledger on the $5.6M figure.
- The trap to avoid — Both mythologies — "$6M killed the scaling story" (the figure is one final run's compute, atop years of program cost) and "nothing happened" (efficiency + open release genuinely repriced assumptions).
- What it sets up — Efficiency innovations (MoE, MLA).
The famous six-million-dollar AI model didn't cost six million. That was one final run's compute bill — the research, failures, and team sat outside it. The honest number still shocked.
The one idea
DeepSeek-R1 proved reasoning emerges from RL on verifiable rewards — open weights, a famously lean price tag with honest caveats, and a one-day market shock measured in hundreds of billions.
In January twenty-twenty-five, one open model's release coincided with the largest single-stock value drop in market history — roughly six hundred billion dollars off the leading chipmaker in a day. The cause sat on a public shelf, free to download. Tonight: what DeepSeek-R1 actually proved, priced honestly. Start with the science — the part that lasts. The team took a strong base model and gave it nothing but the verifier and reinforcement — no worked examples of reasoning, no demonstrations, no taste-judge. Just: attempt, check, reward. And reasoning emerged anyway. Chains lengthened on their own; accuracy climbed with them. The purest version of this act's story, public: incentive alone grew thought. The paper documents the arrival as "the aha moment." Mid-training, R1-Zero began flagging its own steps — "wait" — re-evaluating, rebuilding. Nobody demonstrated the move — the researchers' surprise reads through the academic prose. You know why it happened: chains that doubt themselves pass verifiers more often.
How it works — the demo
R1-Zero's chains lengthening as reward climbs; the "aha moment" from the paper; the Nvidia chart cliff; the caveat ledger on the $5.6M figure.
On the record, in the open — the era's signature result. Now the number, priced honestly. The famous five-point-six million was the final pretraining run's compute for the V3 base model — a real, published figure. It was never the program's cost — research, failed runs, infrastructure, and a world-class team sit outside it, as the team noted. Honest reading: dramatically lean engineering, built on clever architecture the model-zoo act covers — a final-run receipt, not a total invoice. Why did markets convulse? A valuation had assumed frontier reasoning belonged to colossal closed compute — and here was o-one-class reasoning, lean bill, open weights, permissive license. The repricing was violent: the largest single-stock daily loss on record. Whether it overcorrected is a market question; that the assumption needed repricing wasn't. And what it changed outlasts the chart. The reasoning recipe — verifier, reinforcement, emergence — went public and replicable; labs worldwide reproduced it within months. The team distilled R1's ability into small models of every size, seeding an ecosystem you'll meet in two episodes.
The trap to avoid
Both mythologies — "$6M killed the scaling story" (the figure is one final run's compute, atop years of program cost) and "nothing happened" (efficiency + open release genuinely repriced assumptions).
Why it matters — and what’s next
Efficiency innovations (MoE, MLA).
And open weights forced their way into the frontier conversation permanently. The walled garden acquired a contested commons next door — that's the durable event. The trap is both mythologies. "Six million killed scaling" — no: the receipt wasn't the invoice, and compute still rules the frontier. "Nothing happened" — also no: recipes went public, small models inherited reasoning, an assumption got marked to market. Efficiency bent the cost curve. It did not repeal it. Hold what survives the noise: incentive grows thought, the recipe is public, efficiency is a real axis. Two threads lead out — thinking itself scales, and giants teach children. Both are next.
This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.
Full transcript 3:00 of narration
In January twenty-twenty-five, one open model's release coincided with the largest single-stock value drop in market history — roughly six hundred billion dollars off the leading chipmaker in a day. The cause sat on a public shelf, free to download. Tonight: what DeepSeek-R1 actually proved, priced honestly.
Start with the science — the part that lasts. The team took a strong base model and gave it nothing but the verifier and reinforcement — no worked examples of reasoning, no demonstrations, no taste-judge. Just: attempt, check, reward. And reasoning emerged anyway. Chains lengthened on their own; accuracy climbed with them. The purest version of this act's story, public: incentive alone grew thought.
The paper documents the arrival as "the aha moment." Mid-training, R1-Zero began flagging its own steps — "wait" — re-evaluating, rebuilding. Nobody demonstrated the move — the researchers' surprise reads through the academic prose. You know why it happened: chains that doubt themselves pass verifiers more often. On the record, in the open — the era's signature result.
Now the number, priced honestly. The famous five-point-six million was the final pretraining run's compute for the V3 base model — a real, published figure. It was never the program's cost — research, failed runs, infrastructure, and a world-class team sit outside it, as the team noted. Honest reading: dramatically lean engineering, built on clever architecture the model-zoo act covers — a final-run receipt, not a total invoice.
Why did markets convulse? A valuation had assumed frontier reasoning belonged to colossal closed compute — and here was o-one-class reasoning, lean bill, open weights, permissive license. The repricing was violent: the largest single-stock daily loss on record. Whether it overcorrected is a market question; that the assumption needed repricing wasn't.
And what it changed outlasts the chart. The reasoning recipe — verifier, reinforcement, emergence — went public and replicable; labs worldwide reproduced it within months. The team distilled R1's ability into small models of every size, seeding an ecosystem you'll meet in two episodes. And open weights forced their way into the frontier conversation permanently. The walled garden acquired a contested commons next door — that's the durable event.
The trap is both mythologies. "Six million killed scaling" — no: the receipt wasn't the invoice, and compute still rules the frontier. "Nothing happened" — also no: recipes went public, small models inherited reasoning, an assumption got marked to market. Efficiency bent the cost curve. It did not repeal it.
Hold what survives the noise: incentive grows thought, the recipe is public, efficiency is a real axis. Two threads lead out — thinking itself scales, and giants teach children. Both are next.