Act 03 · How it learns 2:30 Think longer, distill smaller, tune lighter

The second scaling axis: think longer.

Test-time compute is a second scaling axis: accuracy climbs with thinking budget — shifting spend from training-time to answer-time.

Video rendering soonThe cinematic render for this episode is being generated. The article, transcript and key ideas are all here now.

Key ideas

  • The one idea — Test-time compute is a second scaling axis: accuracy climbs with thinking budget — shifting spend from training-time to answer-time.
  • How it is shown — The curve: same model, same question set, accuracy rising as the thinking budget sweeps up; the two-axes map redrawn.
  • The trap to avoid — The curve saturates and can invert — overthinking is real; budget is a dial to tune per task, not a virtue to max.
  • What it sets up — Small models + thinking budgets vs giants.

Give a modest model enough time to think, and on the right problems it can beat a giant answering instantly. Intelligence now has a second dial — and a meter.

The one idea

Test-time compute is a second scaling axis: accuracy climbs with thinking budget — shifting spend from training-time to answer-time.

For a decade the race was bigger brains — more parameters, more pretraining, act three's whole economy. Then reasoning models revealed a second axis, perpendicular to the first: same brain, longer thoughts. And accuracy climbs that axis too. The measurement that reorganized roadmaps: hold the model fixed, sweep the thinking budget, and accuracy climbs — steeply at first, then grinding, across orders of magnitude of tokens. Plottable, predictable enough to exploit: the family resemblance to episode seventy-one's scaling law is unmistakable. Except this curve is bought at answer-time, not training-time. Why the shift matters: it moves the money. The old economy paid once, colossally, at training — then answers were nearly free. The new axis bills per hard question: intelligence with a meter on it, spent exactly where difficulty lives.

How it works — the demo

The curve: same model, same question set, accuracy rising as the thinking budget sweeps up; the two-axes map redrawn.

That restructures unit economics, pricing, and what datacenters are even for. And the axes compose. Budget-matched, a modest model thinking long can beat a giant answering instantly — on the structured problems where thinking pays. Which resurrects Chinchilla's lesson one level up: the question is never just how big. It's how to split the spend — this time between brain and thought, per query. Allocation beats accumulation, again. Now the honest limits, because the curve is not a religion. It saturates — past a point, more tokens buy noise. And it can invert: overthinking is measured, models talking themselves out of correct answers, doubt spiraling past diligence into damage.

The trap to avoid

The curve saturates and can invert — overthinking is real; budget is a dial to tune per task, not a virtue to max.

Why it matters — and what’s next

Small models + thinking budgets vs giants.

The budget dial has a sweet spot per task, not a virtue at maximum. Longer thoughts, like most good things, can be overdone. The trap: lane-change thinking — "training scaling is over, it's all test-time now." Both axes are live; frontier labs push both, hard, and trade them per product. The second axis didn't replace the first. It turned a line into a plane — and strategy into a portfolio problem across it. And the new plane has a favorite creature: the small model that genuinely reasons — cheap every day of its life, potent when it thinks. Which begs the question the next episode answers: where does small competence come from? Mostly, it's taught — by giants. Next: distillation.

This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.

Full transcript 2:30 of narration

For a decade the race was bigger brains — more parameters, more pretraining, act three's whole economy. Then reasoning models revealed a second axis, perpendicular to the first: same brain, longer thoughts. And accuracy climbs that axis too.

The measurement that reorganized roadmaps: hold the model fixed, sweep the thinking budget, and accuracy climbs — steeply at first, then grinding, across orders of magnitude of tokens. Plottable, predictable enough to exploit: the family resemblance to episode seventy-one's scaling law is unmistakable. Except this curve is bought at answer-time, not training-time.

Why the shift matters: it moves the money. The old economy paid once, colossally, at training — then answers were nearly free. The new axis bills per hard question: intelligence with a meter on it, spent exactly where difficulty lives. That restructures unit economics, pricing, and what datacenters are even for.

And the axes compose. Budget-matched, a modest model thinking long can beat a giant answering instantly — on the structured problems where thinking pays. Which resurrects Chinchilla's lesson one level up: the question is never just how big. It's how to split the spend — this time between brain and thought, per query. Allocation beats accumulation, again.

Now the honest limits, because the curve is not a religion. It saturates — past a point, more tokens buy noise. And it can invert: overthinking is measured, models talking themselves out of correct answers, doubt spiraling past diligence into damage. The budget dial has a sweet spot per task, not a virtue at maximum. Longer thoughts, like most good things, can be overdone.

The trap: lane-change thinking — "training scaling is over, it's all test-time now." Both axes are live; frontier labs push both, hard, and trade them per product. The second axis didn't replace the first. It turned a line into a plane — and strategy into a portfolio problem across it.

And the new plane has a favorite creature: the small model that genuinely reasons — cheap every day of its life, potent when it thinks. Which begs the question the next episode answers: where does small competence come from? Mostly, it's taught — by giants. Next: distillation.

Test-Time ScalingCapex to OpexOverthinking Limits