Act 02 · The prediction engine 2:15 Depth, width, and the context window

Depth vs. width, and what each buys.

Deeper = more abstraction steps; wider = more capacity per step — and the bill scales differently: width is priced quadratically.

Video rendering soonThe cinematic render for this episode is being generated. The article, transcript and key ideas are all here now.

Key ideas

  • The one idea — Deeper = more abstraction steps; wider = more capacity per step — and the bill scales differently: width is priced quadratically.
  • How it is shown — Two towers, tall-thin vs short-wide, same parameter mass, different characters.
  • The trap to avoid — "Deeper is smarter" — balance is empirical; both are capacity axes, and architects tune the ratio, not a maximum.
  • What it sets up — The desk (context) is a third, separate axis.

Making a model deeper costs a flat fee per floor. Making it wider? Double the width and every floor's machinery quadruples. That square is why nobody maxes out width.

The one idea

Deeper = more abstraction steps; wider = more capacity per step — and the bill scales differently: width is priced quadratically.

Should a model be tall, or fat? Same total weight, two different animals. This is a real tradeoff, with real money attached. Depth buys steps. Each floor re-describes the floor below, so a hundred floors is a hundred chances to build on prior work — resolve the grammar, then the references, then the plot. Structure that needs stages wants a tall building. Width buys room. Width is the stream's bandwidth — the thousands of dimensions from episode twenty-five — plus proportionally bigger vaults and lenses at every floor. A wide floor holds many distinctions at once; a narrow one crowds them into interference.

How it works — the demo

Two towers, tall-thin vs short-wide, same parameter mass, different characters.

Nuance wants a wide building. Now the bill. Parameters scale roughly as twelve times the floor count times the width squared. Read the exponents: depth is priced linearly — another floor, another fixed charge. Width is priced quadratically — double the boulevard and every floor's machinery quadruples. That square is why nobody just maxes out width. So what do architects do? Both, in proportion. Across model families, height-to-width ratios cluster in a surprisingly narrow band — found by experiment, not derived from theory.

The trap to avoid

"Deeper is smarter" — balance is empirical; both are capacity axes, and architects tune the ratio, not a maximum.

Why it matters — and what’s next

The desk (context) is a third, separate axis.

When labs scale a model up, they grow the tower in both dimensions together, keeping the family silhouette. The trap: slogan architecture. "Deeper models reason better" and "wider models know more" both oversimplify — depth and width are two axes of the same capacity, and the craft is the ratio. When a model card lists layers and hidden size, you're reading a proportion decision, not a philosophy. And there's a third axis, independent of both: not how tall the tower or how wide the boulevard — but how many tokens fit on the desk at its entrance. The context window. Next episode, the most misunderstood number on every model card.

'Deeper models reason better.' 'Wider models know more.' Both slogans oversimplify. When a model card lists layers and hidden size, you're reading a proportion decision, not a philosophy.

This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.

Full transcript 2:15 of narration

Should a model be tall, or fat? Same total weight, two different animals. This is a real tradeoff, with real money attached.

Depth buys steps. Each floor re-describes the floor below, so a hundred floors is a hundred chances to build on prior work — resolve the grammar, then the references, then the plot. Structure that needs stages wants a tall building.

Width buys room. Width is the stream's bandwidth — the thousands of dimensions from episode twenty-five — plus proportionally bigger vaults and lenses at every floor. A wide floor holds many distinctions at once; a narrow one crowds them into interference. Nuance wants a wide building.

Now the bill. Parameters scale roughly as twelve times the floor count times the width squared. Read the exponents: depth is priced linearly — another floor, another fixed charge. Width is priced quadratically — double the boulevard and every floor's machinery quadruples. That square is why nobody just maxes out width.

So what do architects do? Both, in proportion. Across model families, height-to-width ratios cluster in a surprisingly narrow band — found by experiment, not derived from theory. When labs scale a model up, they grow the tower in both dimensions together, keeping the family silhouette.

The trap: slogan architecture. "Deeper models reason better" and "wider models know more" both oversimplify — depth and width are two axes of the same capacity, and the craft is the ratio. When a model card lists layers and hidden size, you're reading a proportion decision, not a philosophy.

And there's a third axis, independent of both: not how tall the tower or how wide the boulevard — but how many tokens fit on the desk at its entrance. The context window. Next episode, the most misunderstood number on every model card.

Depth vs WidthParameter EconomicsArchitecture Ratios