Act 03 · How it learns 2:15 Memorization and the price tag

What a FLOP is, and why we count them.

The FLOP — one elementary arithmetic operation — is AI's hardware-independent currency of effort; training cost ≈ 6 × parameters × tokens, and regulators now write thresholds in it.

Video rendering soonThe cinematic render for this episode is being generated. The article, transcript and key ideas are all here now.

Key ideas

  • The one idea — The FLOP — one elementary arithmetic operation — is AI's hardware-independent currency of effort; training cost ≈ 6 × parameters × tokens, and regulators now write thresholds in it.
  • How it is shown — The odometer: one multiply = one tick; the napkin formula computing GPT-3's ~3×10²³; the frontier at 10²⁵–10²⁶.
  • The trap to avoid — FLOPs measure effort, not achievement — a badly-spent budget buys less; the unit prices the burn, not the mind.
  • What it sets up — Spend the same budget smarter?

There's a number written into European law that decides when an AI is powerful: ten to the twenty-fifth arithmetic operations. Cross it, and legal duties switch on.

The one idea

The FLOP — one elementary arithmetic operation — is AI's hardware-independent currency of effort; training cost ≈ 6 × parameters × tokens, and regulators now write thresholds in it.

The unit of AI effort has a name, and once you know it you'll see it everywhere — in every lab report, every scaling chart, and now in the text of law. One multiply, or one add: one FLOP. Count them, and you can weigh minds. Why count operations instead of dollars or chips? Because the count doesn't lie. Dollars fluctuate with markets, chips with generations, electricity with geography — but the number of elementary operations a training run performs is a physical fact, comparable across every era and fleet. FLOPs are effort itself, denominated honestly. And the count fits on a napkin.

How it works — the demo

The odometer: one multiply = one tick; the napkin formula computing GPT-3's ~3×10²³; the frontier at 10²⁵–10²⁶.

Training FLOPs: roughly six, times the parameter count, times the tokens seen. Why six? Each weight works about twice on the forward guess and about four more times in the backward blame — episode sixty-three's two-for-one, itemized. Three known numbers, one multiplication — estimate any run in history. The magnitudes: GPT-3 burned roughly three times ten-to-the-twenty-third operations — and the napkin gets you there: a hundred seventy-five billion parameters, three hundred billion tokens, times six. The current frontier sits two to three orders of magnitude higher: ten-to-the-twenty-fifth, twenty-sixth. Numbers past intuition — which is why the unit matters: it makes the incomprehensible comparable. And here's how load-bearing the unit became: it's in law.

The trap to avoid

FLOPs measure effort, not achievement — a badly-spent budget buys less; the unit prices the burn, not the mind.

Why it matters — and what’s next

Spend the same budget smarter?

Regulators needed a bright line for "powerful model," and dollars wouldn't hold still — so statutes now set obligations at compute thresholds; the EU's AI Act draws one at ten-to-the-twenty-fifth FLOPs. Cross that odometer reading during training, and legal duties switch on. An arithmetic count, governing minds. The trap, before the unit goes to your head: FLOPs measure effort, not achievement. Two runs can burn identical budgets and produce unequal minds — the spend can be allocated badly. Which raises the question that reorganized the industry: given a fixed budget of operations, is there a best way to spend it? Next: the curve that runs everything.

This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.

Full transcript 2:15 of narration

The unit of AI effort has a name, and once you know it you'll see it everywhere — in every lab report, every scaling chart, and now in the text of law. One multiply, or one add: one FLOP. Count them, and you can weigh minds.

Why count operations instead of dollars or chips? Because the count doesn't lie. Dollars fluctuate with markets, chips with generations, electricity with geography — but the number of elementary operations a training run performs is a physical fact, comparable across every era and fleet. FLOPs are effort itself, denominated honestly.

And the count fits on a napkin. Training FLOPs: roughly six, times the parameter count, times the tokens seen. Why six? Each weight works about twice on the forward guess and about four more times in the backward blame — episode sixty-three's two-for-one, itemized. Three known numbers, one multiplication — estimate any run in history.

The magnitudes: GPT-3 burned roughly three times ten-to-the-twenty-third operations — and the napkin gets you there: a hundred seventy-five billion parameters, three hundred billion tokens, times six. The current frontier sits two to three orders of magnitude higher: ten-to-the-twenty-fifth, twenty-sixth. Numbers past intuition — which is why the unit matters: it makes the incomprehensible comparable.

And here's how load-bearing the unit became: it's in law. Regulators needed a bright line for "powerful model," and dollars wouldn't hold still — so statutes now set obligations at compute thresholds; the EU's AI Act draws one at ten-to-the-twenty-fifth FLOPs. Cross that odometer reading during training, and legal duties switch on. An arithmetic count, governing minds.

The trap, before the unit goes to your head: FLOPs measure effort, not achievement. Two runs can burn identical budgets and produce unequal minds — the spend can be allocated badly. Which raises the question that reorganized the industry: given a fixed budget of operations, is there a best way to spend it? Next: the curve that runs everything.

FLOPsNapkin Compute MathRegulation Thresholds