Key ideas
- The one idea — "Trained on the internet" means a brutal funnel — web crawl, code, books, papers — deduplicated, filtered, and blended into trillions of tokens, with most of the raw pile discarded.
- How it is shown — The funnel: the raw mountain collapsing stage by stage into the refined river; the mixture-valves; what never entered at all.
- The trap to avoid — "It read everything" — it read a heavily edited version; absences (paywalled, private, post-cutoff) shape it as much as inclusions.
- What it sets up — The highest-grade human text is running out.
Before a model trains, most of the pile gets thrown away — deduplicated, graded, discarded in sheets. What survives: a river of trillions of tokens. Here's the funnel.
The one idea
"Trained on the internet" means a brutal funnel — web crawl, code, books, papers — deduplicated, filtered, and blended into trillions of tokens, with most of the raw pile discarded.
It didn't read "the internet." It read a heavily edited version — the survivors of a brutal funnel. Tonight: what goes in the top, what dies on the way down, and what never entered at all. The strata: web crawl is the bulk — billions of pages, mostly junk. Code repositories, which matter far beyond coding. Books. Encyclopedias. Papers. Forums, with their strange conversational gold. The raw mountain is the written exhaust of civilization, unsorted. Now the funnel. Language identification sorts the torrents. Deduplication — the biggest kill — collapses the internet's staggering self-repetition into single copies, and the mountain visibly shrinks.
How it works — the demo
The funnel: the raw mountain collapsing stage by stage into the refined river; the mixture-valves; what never entered at all.
Quality classifiers grade what's left and discard in sheets. Scrubbers spark on private data. What survives: a river of trillions of tokens. Most of the pile never became fuel. Then the recipe. How much code versus prose? How many passes over the books? Upweight the papers, or the forums? Ratios measurably shift the model's character — heavier code famously sharpens reasoning — and frontier mixtures are guarded. Open models publish recipes; frontier ones don't. And weigh the absences, because they shape the model as much as the fuel. Paywalled and licensed text — subject of ongoing lawsuits.
The trap to avoid
"It read everything" — it read a heavily edited version; absences (paywalled, private, post-cutoff) shape it as much as inclusions.
Why it matters — and what’s next
The highest-grade human text is running out.
Private messages, never crawled. Most of the world's non-English writing, thinly represented. And everything after the cutoff date — a clean cliff where the mountain simply ends. The model's blind spots are a map of what the funnel never carried. The trap: "it read everything." It read one pile, chosen by someone, filtered by one pipeline, frozen at one date. Its knowledge is a portrait of those choices — including their biases, their gaps, and their era. When a model seems to "know the internet," you're really seeing the funnel's editorial judgment, speaking fluently. One more thing is visible from up here: the best strata — the books, the papers, the curated prose — are close to fully quarried. Frontier appetites have nearly consumed the good mountain. And so a strange new tributary has begun to flow into the funnel: text written by the machines themselves. Next: models trained on models — the synthetic turn, its promise, and its documented dangers.
This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.
Full transcript 2:30 of narration
It didn't read "the internet." It read a heavily edited version — the survivors of a brutal funnel. Tonight: what goes in the top, what dies on the way down, and what never entered at all.
The strata: web crawl is the bulk — billions of pages, mostly junk. Code repositories, which matter far beyond coding. Books. Encyclopedias. Papers. Forums, with their strange conversational gold. The raw mountain is the written exhaust of civilization, unsorted.
Now the funnel. Language identification sorts the torrents. Deduplication — the biggest kill — collapses the internet's staggering self-repetition into single copies, and the mountain visibly shrinks. Quality classifiers grade what's left and discard in sheets. Scrubbers spark on private data. What survives: a river of trillions of tokens. Most of the pile never became fuel.
Then the recipe. How much code versus prose? How many passes over the books? Upweight the papers, or the forums? Ratios measurably shift the model's character — heavier code famously sharpens reasoning — and frontier mixtures are guarded. Open models publish recipes; frontier ones don't.
And weigh the absences, because they shape the model as much as the fuel. Paywalled and licensed text — subject of ongoing lawsuits. Private messages, never crawled. Most of the world's non-English writing, thinly represented. And everything after the cutoff date — a clean cliff where the mountain simply ends. The model's blind spots are a map of what the funnel never carried.
The trap: "it read everything." It read one pile, chosen by someone, filtered by one pipeline, frozen at one date. Its knowledge is a portrait of those choices — including their biases, their gaps, and their era. When a model seems to "know the internet," you're really seeing the funnel's editorial judgment, speaking fluently.
One more thing is visible from up here: the best strata — the books, the papers, the curated prose — are close to fully quarried. Frontier appetites have nearly consumed the good mountain. And so a strange new tributary has begun to flow into the funnel: text written by the machines themselves. Next: models trained on models — the synthetic turn, its promise, and its documented dangers.