Act 04 · The physical machine 2:30 Chips, FP4, and the cloud-vs-self-host math

A100 → H100 → B200: why the chip matters.

Each GPU generation adds memory, bandwidth, and — crucially — new native number formats that unlock new models and prices.

Video rendering soonThe cinematic render for this episode is being generated. The article, transcript and key ideas are all here now.

Key ideas

  • The one idea — Each GPU generation adds memory, bandwidth, and — crucially — new native number formats that unlock new models and prices.
  • How it is shown — Three chips on podiums: the same workload migrating across them, gaining a format each hop; the sparsity-asterisk on the marketing numbers.
  • The trap to avoid — Reading spec sheets at face value — headline TFLOPS often carry a sparsity asterisk and format fine print.
  • What it sets up — What happens when the format outruns your silicon.

That giant TFLOPS number in the GPU announcement? Check the footnote. It often assumes sparsity and a specific format — the honest number can be a fraction of the headline.

The one idea

Each GPU generation adds memory, bandwidth, and — crucially — new native number formats that unlock new models and prices.

Why does a two-generation-old GPU struggle with this year's models? Not age — vocabulary. Each generation adds three things: more room, wider roads, and — the underrated one — new formats spoken natively. Three chips, three eras, one pattern. Learn it and you can read any announcement. Generation one: the A100, twenty-twenty — the workhorse an era grew up on. Forty then eighty gigabytes, around two terabytes a second, and the feature that mattered most: native brain-float sixteen — episode ninety-seven's range-first dialect. A generation of famous models ran on this vocabulary. Generation two: the H100, twenty-twenty-two — the chip of the current boom. Eighty gigs, three-point-three terabytes a second, thicker spans between chips — and the signature feature: native FP8. Half-width crates at double the rate — the format frontier labs now train serious fractions of runs in.

How it works — the demo

Three chips on podiums: the same workload migrating across them, gaining a format each hop; the sparsity-asterisk on the marketing numbers.

New vocabulary, new economics, new frontier. Generation three: Blackwell — the B200, twenty-twenty-four. Two dies fused into one giant, nearly two hundred gigabytes, roughly eight terabytes a second — and the headline vocabulary: native four-bit matrix math. Episode ninety-six's staircase, descended another step in silicon. The pattern is visible now: every generation, three gifts — room, roads, a smaller tongue. Formats are why the chip matters strategically: models increasingly ship tuned to one — FP8 today, FP4 arriving. On silicon that speaks it, the math runs native. On older silicon, a translator steps in — crates unpacked before computing — and part of the promise evaporates. "Can my GPU run it" is becoming "does it speak it." That trap is the next episode. The trap: spec-sheet séance. Headline numbers often assume sparsity — zeroing half the values, doubling the figure — and quote the slimmest supported format, not yours.

The trap to avoid

Reading spec sheets at face value — headline TFLOPS often carry a sparsity asterisk and format fine print.

Why it matters — and what’s next

What happens when the format outruns your silicon.

The honest question when a banner shouts: dense or sparse? Which format? At what memory? Shake the number until the asterisks fall out. They always do. Now the fine print that costs real money: take a four-bit model and load it on last generation's hardware. It runs. Something is even better. Something else quietly isn't. Which benefits survive translation, and which evaporate? Next: why FP4 needs Blackwell.

This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.

Full transcript 2:30 of narration

Why does a two-generation-old GPU struggle with this year's models? Not age — vocabulary. Each generation adds three things: more room, wider roads, and — the underrated one — new formats spoken natively. Three chips, three eras, one pattern. Learn it and you can read any announcement.

Generation one: the A100, twenty-twenty — the workhorse an era grew up on. Forty then eighty gigabytes, around two terabytes a second, and the feature that mattered most: native brain-float sixteen — episode ninety-seven's range-first dialect. A generation of famous models ran on this vocabulary.

Generation two: the H100, twenty-twenty-two — the chip of the current boom. Eighty gigs, three-point-three terabytes a second, thicker spans between chips — and the signature feature: native FP8. Half-width crates at double the rate — the format frontier labs now train serious fractions of runs in. New vocabulary, new economics, new frontier.

Generation three: Blackwell — the B200, twenty-twenty-four. Two dies fused into one giant, nearly two hundred gigabytes, roughly eight terabytes a second — and the headline vocabulary: native four-bit matrix math. Episode ninety-six's staircase, descended another step in silicon. The pattern is visible now: every generation, three gifts — room, roads, a smaller tongue.

Formats are why the chip matters strategically: models increasingly ship tuned to one — FP8 today, FP4 arriving. On silicon that speaks it, the math runs native. On older silicon, a translator steps in — crates unpacked before computing — and part of the promise evaporates. "Can my GPU run it" is becoming "does it speak it." That trap is the next episode.

The trap: spec-sheet séance. Headline numbers often assume sparsity — zeroing half the values, doubling the figure — and quote the slimmest supported format, not yours. The honest question when a banner shouts: dense or sparse? Which format? At what memory? Shake the number until the asterisks fall out. They always do.

Now the fine print that costs real money: take a four-bit model and load it on last generation's hardware. It runs. Something is even better. Something else quietly isn't. Which benefits survive translation, and which evaporate? Next: why FP4 needs Blackwell.

Hardware & InferenceHardwarePrecision