Act 03 · How it learns 2:30 Think longer, distill smaller, tune lighter

Big model teaches small model (distillation).

A large teacher's outputs become a small student's curriculum — ability transfers that the student could never have grown alone.

Video rendering soonThe cinematic render for this episode is being generated. The article, transcript and key ideas are all here now.

Key ideas

  • The one idea — A large teacher's outputs become a small student's curriculum — ability transfers that the student could never have grown alone.
  • How it is shown — The teacher writing worked examples at scale; the student fine-tuned on them; R1's reasoning appearing in pocket-size models.
  • The trap to avoid — The student inherits the teacher's ceiling and its flaws — copies of copies need the gauntlet (ep 67's law returns); and cross-vendor distillation lives in legal gray.
  • What it sets up — Distillation closes EP 067's teacher-student vignette · small-model serving economics.

The fast, cheap model in your favorite AI family? Widely reported to be substantially a distillation — the flagship taught it. Giants teaching children is the industry's quiet pipeline.

The one idea

A large teacher's outputs become a small student's curriculum — ability transfers that the student could never have grown alone.

When a pocket-sized model reasons like something ten times its mass, you're watching inheritance. It didn't grow that ability — it was taught, by a giant. Distillation: the quiet pipeline behind every "small model shocks benchmarks" headline. Two things you own, composed. The teacher generates — thousands of worked problems with full reasoning, a curriculum no human team could author at that scale or consistency. The student fine-tunes on it: episode seventy-five's ritual, with a binder written by a giant. Distillation is synthetic data aimed downward — episode sixty-seven's vignette, industrialized. Why does copying beat studying? Curriculum quality.

How it works — the demo

The teacher writing worked examples at scale; the student fine-tuned on them; R1's reasoning appearing in pocket-size models.

The internet taught the teacher through a maze of noise; the teacher hands the student a shortcut — every page worked, complete, clean, pitched at the right difficulty. The student learns destinations without wandering the maze that found them. Small capacity, spent only on distilled lessons, goes shockingly far. The proof went public with R1: its reasoning distilled into open children of every size, pocket models suddenly beating far larger peers on math and code. And the practice runs everywhere, unlabeled — the fast sibling in most frontier families is widely reported to be substantially a distillation of the flagship. Frontier teaches product. The fine print of inheritance. The student learns the teacher's answers, so the teacher's ceiling becomes the student's sky — surpassing your curriculum is rare. The flaws transfer too: quirks, blind spots, biases, copied at industrial fidelity.

The trap to avoid

The student inherits the teacher's ceiling and its flaws — copies of copies need the gauntlet (ep 67's law returns); and cross-vendor distillation lives in legal gray.

Why it matters — and what’s next

Distillation closes EP 067's teacher-student vignette · small-model serving economics.

And chains of copies-of-copies decay exactly as episode sixty-seven warned — distillation without the verification gauntlet is how ecosystems go soft. The trap: assuming the pipeline is all licensed and cozy. Distilling your own flagship into your own products — standard. Training on a rival's outputs breaches most terms of service — and is loudly alleged in every direction, including at the labs doing the accusing. Elegant technique; unsettled ground. Walk it knowingly. So giants teach children, and the children staff the products. Which leaves the question every builder actually faces — not how to train a model, but how to change one: three ways, three prices, and a decision tree that saves real money. Next: prompting, LoRA, or the full furnace.

This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.

Full transcript 2:30 of narration

When a pocket-sized model reasons like something ten times its mass, you're watching inheritance. It didn't grow that ability — it was taught, by a giant. Distillation: the quiet pipeline behind every "small model shocks benchmarks" headline.

Two things you own, composed. The teacher generates — thousands of worked problems with full reasoning, a curriculum no human team could author at that scale or consistency. The student fine-tunes on it: episode seventy-five's ritual, with a binder written by a giant. Distillation is synthetic data aimed downward — episode sixty-seven's vignette, industrialized.

Why does copying beat studying? Curriculum quality. The internet taught the teacher through a maze of noise; the teacher hands the student a shortcut — every page worked, complete, clean, pitched at the right difficulty. The student learns destinations without wandering the maze that found them. Small capacity, spent only on distilled lessons, goes shockingly far.

The proof went public with R1: its reasoning distilled into open children of every size, pocket models suddenly beating far larger peers on math and code. And the practice runs everywhere, unlabeled — the fast sibling in most frontier families is widely reported to be substantially a distillation of the flagship. Frontier teaches product.

The fine print of inheritance. The student learns the teacher's answers, so the teacher's ceiling becomes the student's sky — surpassing your curriculum is rare. The flaws transfer too: quirks, blind spots, biases, copied at industrial fidelity. And chains of copies-of-copies decay exactly as episode sixty-seven warned — distillation without the verification gauntlet is how ecosystems go soft.

The trap: assuming the pipeline is all licensed and cozy. Distilling your own flagship into your own products — standard. Training on a rival's outputs breaches most terms of service — and is loudly alleged in every direction, including at the labs doing the accusing. Elegant technique; unsettled ground. Walk it knowingly.

So giants teach children, and the children staff the products. Which leaves the question every builder actually faces — not how to train a model, but how to change one: three ways, three prices, and a decision tree that saves real money. Next: prompting, LoRA, or the full furnace.

DistillationTeacher-Student PipelineInherited Ceilings