Act 03 · How it learns 2:30 Reasoning models and the $6M shock

Reasoning models vs. instant models.

Test-time thinking trades latency and money for accuracy on hard, checkable problems — a routing decision, not a ranking.

Video rendering soonThe cinematic render for this episode is being generated. The article, transcript and key ideas are all here now.

Key ideas

  • The one idea — Test-time thinking trades latency and money for accuracy on hard, checkable problems — a routing decision, not a ranking.
  • How it is shown — Same hard problem: instant answer confidently wrong; thinking answer slower, metered, right. Same easy question: thinking mode burning tokens to conclude the obvious.
  • The trap to avoid — "Reasoning model = better model" — on lookups and chat it's slower, pricier, and no righter; route per task, and products increasingly route for you.
  • What it sets up — The accuracy-vs-thinking-budget curve.

A reasoning model isn't a better model. On lookups and chat it's slower, pricier, and no righter — the wins live only where problems have structure. Routing is the skill.

The one idea

Test-time thinking trades latency and money for accuracy on hard, checkable problems — a routing decision, not a ranking.

Same hard problem, two doors. The fast door answers instantly — confident, wrong. The deep door churns, bills you by the token, and lands it. Neither door is better. They're modes — and routing is the skill. Under the hood: one machinery, two spending patterns. Instant mode is the loop you've known since act two: prompt in, answer streaming out. Thinking mode grants the same machine a private budget first — hundreds or thousands of working tokens, each a full pass through the tower, each billed — before the visible answer starts. You're not buying a different brain.

How it works — the demo

Same hard problem: instant answer confidently wrong; thinking answer slower, metered, right. Same easy question: thinking mode burning tokens to conclude the obvious.

You're buying it time. Where thinking pays: problems with structure. Multi-step math, where one slip poisons everything after. Hard debugging, planning, constraint puzzles — anywhere step two depends on step one being right, the private working space turns compounding errors into caught ones. The benchmark gaps here aren't subtle — they're chasms. And where it's waste: lookups, chat, anything the weights simply know or don't. Deliberation doesn't improve recall — the capital of France is in the vaults or it isn't, and a thousand thinking tokens spent confirming it are money on fire with a serious expression. The industry's answer is routing. Modern products sit a dispatcher ahead of the doors: lookups go fast, structure goes deep, budgets scale with difficulty.

The trap to avoid

"Reasoning model = better model" — on lookups and chat it's slower, pricier, and no righter; route per task, and products increasingly route for you.

Why it matters — and what’s next

The accuracy-vs-thinking-budget curve.

Many frontier models are hybrids with thinking as a dial rather than an identity. When your assistant sometimes pauses, sometimes fires back — that's the router. The trap: treating "reasoning model" as a rank instead of a tool. On easy work the deep door is slower, pricier, and no righter — and at production volume, that bill compounds ruinously. The practitioners' skill is routing: match mode to structure; let neither door become an identity. And that thinking dial hides a discovery big enough to reorganize the industry: turn it up, and accuracy climbs along a curve — a second scaling axis, as real as the first. That story is two episodes away. First: the model that proved all of this could be done in the open, for a fraction of the assumed price — and shook the market doing it.

This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.

Full transcript 2:30 of narration

Same hard problem, two doors. The fast door answers instantly — confident, wrong. The deep door churns, bills you by the token, and lands it. Neither door is better. They're modes — and routing is the skill.

Under the hood: one machinery, two spending patterns. Instant mode is the loop you've known since act two: prompt in, answer streaming out. Thinking mode grants the same machine a private budget first — hundreds or thousands of working tokens, each a full pass through the tower, each billed — before the visible answer starts. You're not buying a different brain. You're buying it time.

Where thinking pays: problems with structure. Multi-step math, where one slip poisons everything after. Hard debugging, planning, constraint puzzles — anywhere step two depends on step one being right, the private working space turns compounding errors into caught ones. The benchmark gaps here aren't subtle — they're chasms.

And where it's waste: lookups, chat, anything the weights simply know or don't. Deliberation doesn't improve recall — the capital of France is in the vaults or it isn't, and a thousand thinking tokens spent confirming it are money on fire with a serious expression.

The industry's answer is routing. Modern products sit a dispatcher ahead of the doors: lookups go fast, structure goes deep, budgets scale with difficulty. Many frontier models are hybrids with thinking as a dial rather than an identity. When your assistant sometimes pauses, sometimes fires back — that's the router.

The trap: treating "reasoning model" as a rank instead of a tool. On easy work the deep door is slower, pricier, and no righter — and at production volume, that bill compounds ruinously. The practitioners' skill is routing: match mode to structure; let neither door become an identity.

And that thinking dial hides a discovery big enough to reorganize the industry: turn it up, and accuracy climbs along a curve — a second scaling axis, as real as the first. That story is two episodes away. First: the model that proved all of this could be done in the open, for a fraction of the assumed price — and shook the market doing it.

Reasoning vs InstantRoutingToken Economics