Key ideas
- The one idea — Each GPU generation adds memory, bandwidth, and — crucially — new native number formats that unlock new models and prices.
- How it is shown — Three chips on podiums: the same workload migrating across them, gaining a format each hop; the sparsity-asterisk on the marketing numbers.
- The trap to avoid — Reading spec sheets at face value — headline TFLOPS often carry a sparsity asterisk and format fine print.
- What it sets up — What happens when the format outruns your silicon.
That giant TFLOPS number in the GPU announcement? Check the footnote. It often assumes sparsity and a specific format — the honest number can be a fraction of the headline.
The one idea
Each GPU generation adds memory, bandwidth, and — crucially — new native number formats that unlock new models and prices.
Why does a two-generation-old GPU struggle with this year's models? Not age — vocabulary. Each generation adds three things: more room, wider roads, and — the underrated one — new formats spoken natively. Three chips, three eras, one pattern. Learn it and you can read any announcement. Generation one: the A100, twenty-twenty — the workhorse an era grew up on. Forty then eighty gigabytes, around two terabytes a second, and the feature that mattered most: native brain-float sixteen — episode ninety-seven's range-first dialect. A generation of famous models ran on this vocabulary. Generation two: the H100, twenty-twenty-two — the chip of the current boom. Eighty gigs, three-point-three terabytes a second, thicker spans between chips — and the signature feature: native FP8. Half-width crates at double the rate — the format frontier labs now train serious fractions of runs in.
How it works — the demo
Three chips on podiums: the same workload migrating across them, gaining a format each hop; the sparsity-asterisk on the marketing numbers.
New vocabulary, new economics, new frontier. Generation three: Blackwell — the B200, twenty-twenty-four. Two dies fused into one giant, nearly two hundred gigabytes, roughly eight terabytes a second — and the headline vocabulary: native four-bit matrix math. Episode ninety-six's staircase, descended another step in silicon. The pattern is visible now: every generation, three gifts — room, roads, a smaller tongue. Formats are why the chip matters strategically: models increasingly ship tuned to one — FP8 today, FP4 arriving. On silicon that speaks it, the math runs native. On older silicon, a translator steps in — crates unpacked before computing — and part of the promise evaporates. "Can my GPU run it" is becoming "does it speak it." That trap is the next episode. The trap: spec-sheet séance. Headline numbers often assume sparsity — zeroing half the values, doubling the figure — and quote the slimmest supported format, not yours.
The trap to avoid
Reading spec sheets at face value — headline TFLOPS often carry a sparsity asterisk and format fine print.
Why it matters — and what’s next
What happens when the format outruns your silicon.
The honest question when a banner shouts: dense or sparse? Which format? At what memory? Shake the number until the asterisks fall out. They always do. Now the fine print that costs real money: take a four-bit model and load it on last generation's hardware. It runs. Something is even better. Something else quietly isn't. Which benefits survive translation, and which evaporate? Next: why FP4 needs Blackwell.
This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.
Full transcript 2:30 of narration
Why does a two-generation-old GPU struggle with this year's models? Not age — vocabulary. Each generation adds three things: more room, wider roads, and — the underrated one — new formats spoken natively. Three chips, three eras, one pattern. Learn it and you can read any announcement.
Generation one: the A100, twenty-twenty — the workhorse an era grew up on. Forty then eighty gigabytes, around two terabytes a second, and the feature that mattered most: native brain-float sixteen — episode ninety-seven's range-first dialect. A generation of famous models ran on this vocabulary.
Generation two: the H100, twenty-twenty-two — the chip of the current boom. Eighty gigs, three-point-three terabytes a second, thicker spans between chips — and the signature feature: native FP8. Half-width crates at double the rate — the format frontier labs now train serious fractions of runs in. New vocabulary, new economics, new frontier.
Generation three: Blackwell — the B200, twenty-twenty-four. Two dies fused into one giant, nearly two hundred gigabytes, roughly eight terabytes a second — and the headline vocabulary: native four-bit matrix math. Episode ninety-six's staircase, descended another step in silicon. The pattern is visible now: every generation, three gifts — room, roads, a smaller tongue.
Formats are why the chip matters strategically: models increasingly ship tuned to one — FP8 today, FP4 arriving. On silicon that speaks it, the math runs native. On older silicon, a translator steps in — crates unpacked before computing — and part of the promise evaporates. "Can my GPU run it" is becoming "does it speak it." That trap is the next episode.
The trap: spec-sheet séance. Headline numbers often assume sparsity — zeroing half the values, doubling the figure — and quote the slimmest supported format, not yours. The honest question when a banner shouts: dense or sparse? Which format? At what memory? Shake the number until the asterisks fall out. They always do.
Now the fine print that costs real money: take a four-bit model and load it on last generation's hardware. It runs. Something is even better. Something else quietly isn't. Which benefits survive translation, and which evaporate? Next: why FP4 needs Blackwell.