Key ideas
- The one idea — Test-time compute is a second scaling axis: accuracy climbs with thinking budget — shifting spend from training-time to answer-time.
- How it is shown — The curve: same model, same question set, accuracy rising as the thinking budget sweeps up; the two-axes map redrawn.
- The trap to avoid — The curve saturates and can invert — overthinking is real; budget is a dial to tune per task, not a virtue to max.
- What it sets up — Small models + thinking budgets vs giants.
Give a modest model enough time to think, and on the right problems it can beat a giant answering instantly. Intelligence now has a second dial — and a meter.
The one idea
Test-time compute is a second scaling axis: accuracy climbs with thinking budget — shifting spend from training-time to answer-time.
For a decade the race was bigger brains — more parameters, more pretraining, act three's whole economy. Then reasoning models revealed a second axis, perpendicular to the first: same brain, longer thoughts. And accuracy climbs that axis too. The measurement that reorganized roadmaps: hold the model fixed, sweep the thinking budget, and accuracy climbs — steeply at first, then grinding, across orders of magnitude of tokens. Plottable, predictable enough to exploit: the family resemblance to episode seventy-one's scaling law is unmistakable. Except this curve is bought at answer-time, not training-time. Why the shift matters: it moves the money. The old economy paid once, colossally, at training — then answers were nearly free. The new axis bills per hard question: intelligence with a meter on it, spent exactly where difficulty lives.
How it works — the demo
The curve: same model, same question set, accuracy rising as the thinking budget sweeps up; the two-axes map redrawn.
That restructures unit economics, pricing, and what datacenters are even for. And the axes compose. Budget-matched, a modest model thinking long can beat a giant answering instantly — on the structured problems where thinking pays. Which resurrects Chinchilla's lesson one level up: the question is never just how big. It's how to split the spend — this time between brain and thought, per query. Allocation beats accumulation, again. Now the honest limits, because the curve is not a religion. It saturates — past a point, more tokens buy noise. And it can invert: overthinking is measured, models talking themselves out of correct answers, doubt spiraling past diligence into damage.
The trap to avoid
The curve saturates and can invert — overthinking is real; budget is a dial to tune per task, not a virtue to max.
Why it matters — and what’s next
Small models + thinking budgets vs giants.
The budget dial has a sweet spot per task, not a virtue at maximum. Longer thoughts, like most good things, can be overdone. The trap: lane-change thinking — "training scaling is over, it's all test-time now." Both axes are live; frontier labs push both, hard, and trade them per product. The second axis didn't replace the first. It turned a line into a plane — and strategy into a portfolio problem across it. And the new plane has a favorite creature: the small model that genuinely reasons — cheap every day of its life, potent when it thinks. Which begs the question the next episode answers: where does small competence come from? Mostly, it's taught — by giants. Next: distillation.
This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.
Full transcript 2:30 of narration
For a decade the race was bigger brains — more parameters, more pretraining, act three's whole economy. Then reasoning models revealed a second axis, perpendicular to the first: same brain, longer thoughts. And accuracy climbs that axis too.
The measurement that reorganized roadmaps: hold the model fixed, sweep the thinking budget, and accuracy climbs — steeply at first, then grinding, across orders of magnitude of tokens. Plottable, predictable enough to exploit: the family resemblance to episode seventy-one's scaling law is unmistakable. Except this curve is bought at answer-time, not training-time.
Why the shift matters: it moves the money. The old economy paid once, colossally, at training — then answers were nearly free. The new axis bills per hard question: intelligence with a meter on it, spent exactly where difficulty lives. That restructures unit economics, pricing, and what datacenters are even for.
And the axes compose. Budget-matched, a modest model thinking long can beat a giant answering instantly — on the structured problems where thinking pays. Which resurrects Chinchilla's lesson one level up: the question is never just how big. It's how to split the spend — this time between brain and thought, per query. Allocation beats accumulation, again.
Now the honest limits, because the curve is not a religion. It saturates — past a point, more tokens buy noise. And it can invert: overthinking is measured, models talking themselves out of correct answers, doubt spiraling past diligence into damage. The budget dial has a sweet spot per task, not a virtue at maximum. Longer thoughts, like most good things, can be overdone.
The trap: lane-change thinking — "training scaling is over, it's all test-time now." Both axes are live; frontier labs push both, hard, and trade them per product. The second axis didn't replace the first. It turned a line into a plane — and strategy into a portfolio problem across it.
And the new plane has a favorite creature: the small model that genuinely reasons — cheap every day of its life, potent when it thinks. Which begs the question the next episode answers: where does small competence come from? Mostly, it's taught — by giants. Next: distillation.