Key ideas
- The one idea — Training sets the billions of numbers — they start as pure random noise, and everything "known" is the carving between noise and final file.
- How it is shown — The untrained model producing gibberish; the same machine mid-training, organizing; the delta as the knowledge.
- The trap to avoid — Imagining knowledge was installed — nothing was written in; everything was adjusted toward, trillions of times.
- What it sets up — The three furnace ingredients — data, compute, the adjustment rule — become the act's map.
Nobody wrote anything into AI. Nobody could — no human knows which of the billions of numbers should hold what. So how does knowledge get in there?
The one idea
Training sets the billions of numbers — they start as pure random noise, and everything "known" is the carving between noise and final file.
Before training, the model is billions of random numbers. Every stencil, every lens, every drawer — static. Feed it a prompt and the machinery runs flawlessly… and produces confetti. Same architecture as GPT; knows nothing. This act is what happens next. Hold the two versions side by side: the static machine, and the one you spent an act inside. The difference between them — billions of tiny numeric shifts — is everything the model knows. Nothing was installed. No facts were written in. Every value was adjusted toward, nudge by nudge. Knowledge, here, is a delta. Welcome to the furnace. Three ingredients, and only three. A mountain of data — text, the teacher.
How it works — the demo
The untrained model producing gibberish; the same machine mid-training, organizing; the delta as the knowledge.
A mountain of compute — the racks that will run for months. And one adjustment rule, small enough to sketch on a napkin, that turns errors into nudges. Everything in this act is one of those three, examined. One training cycle: show the machine some text. Let it guess the next word — badly, at first. Compare the guess to the truth, which is sitting right there in the text. Send a correction backward, nudging every weight a hair toward having-been-less-wrong. That's the entire ritual. There is no other ritual. Now repeat it trillions of times. Months of continuous burn. No single nudge teaches anything — each shifts a weight by less than a rounding error. But trillions of nudges, each slightly biased toward less wrong, and the static organizes. Grammar condenses.
The trap to avoid
Imagining knowledge was installed — nothing was written in; everything was adjusted toward, trillions of times.
Why it matters — and what’s next
The three furnace ingredients — data, compute, the adjustment rule — become the act's map.
Geography settles into the drawers. The map of meaning assembles itself. Structure, from repetition. The trap — the deepest folk error about AI: imagining engineers writing the knowledge in. Nobody writes anything in. Nobody could — no human knows which of the billions of numbers should hold what. Engineers build the furnace, choose the fuel, and tend the gauges. The data does the teaching. That inversion explains most of AI's strangeness. But wait — of all the lessons a machine could drill for months, why "guess the next word"? Who picked that game, and why did it work when grander plans failed? Next episode: the core training game, and the quiet genius inside it.
This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.
Full transcript 2:30 of narration
Before training, the model is billions of random numbers. Every stencil, every lens, every drawer — static. Feed it a prompt and the machinery runs flawlessly… and produces confetti. Same architecture as GPT; knows nothing. This act is what happens next.
Hold the two versions side by side: the static machine, and the one you spent an act inside. The difference between them — billions of tiny numeric shifts — is everything the model knows. Nothing was installed. No facts were written in. Every value was adjusted toward, nudge by nudge. Knowledge, here, is a delta.
Welcome to the furnace. Three ingredients, and only three. A mountain of data — text, the teacher. A mountain of compute — the racks that will run for months. And one adjustment rule, small enough to sketch on a napkin, that turns errors into nudges. Everything in this act is one of those three, examined.
One training cycle: show the machine some text. Let it guess the next word — badly, at first. Compare the guess to the truth, which is sitting right there in the text. Send a correction backward, nudging every weight a hair toward having-been-less-wrong. That's the entire ritual. There is no other ritual.
Now repeat it trillions of times. Months of continuous burn. No single nudge teaches anything — each shifts a weight by less than a rounding error. But trillions of nudges, each slightly biased toward less wrong, and the static organizes. Grammar condenses. Geography settles into the drawers. The map of meaning assembles itself. Structure, from repetition.
The trap — the deepest folk error about AI: imagining engineers writing the knowledge in. Nobody writes anything in. Nobody could — no human knows which of the billions of numbers should hold what. Engineers build the furnace, choose the fuel, and tend the gauges. The data does the teaching. That inversion explains most of AI's strangeness.
But wait — of all the lessons a machine could drill for months, why "guess the next word"? Who picked that game, and why did it work when grander plans failed? Next episode: the core training game, and the quiet genius inside it.