Key ideas
- The one idea — At 100× scale, an unplanted ability appeared: show examples in the prompt and the model learns the task on the spot — no weight updates.
- How it is shown — The prompt-classroom: three worked examples typed in, the fourth solved live; the weights frozen throughout; the ability nobody trained.
- The trap to avoid — Hearing "learning" and imagining weight changes — in-context learning changes nothing on disk; it's inference wearing learning's coat.
- What it sets up — Raw talent, terrible manners.
There's a kind of AI learning where nothing on the computer changes. No retraining, no updates to the model — and it was discovered, not designed.
The one idea
At 100× scale, an unplanted ability appeared: show examples in the prompt and the model learns the task on the spot — no weight updates.
Twenty-twenty. GPT-3 arrives — a hundred times its parent's scale — and inside it, researchers find an ability nobody put there: show it a few worked examples in the prompt, and it performs the new task immediately. No retraining. No weight changes. Learning, apparently, without learning. The surprise that named an era: in-context learning. The demonstration held up. Invented tasks — rules nowhere in any training pile — taught by three examples, solved on the fourth. The same prompts on smaller siblings: failure. Nothing in the architecture changed; the ability arrived with scale. Episode eighty-eight's emergence debate found its flagship exhibit here.
How it works — the demo
The prompt-classroom: three worked examples typed in, the fourth solved live; the weights frozen throughout; the ability nobody trained.
What is it, mechanically? Episode seventy-three's fishing, elevated. A base model continues genres; at scale, it infers a genre from three examples and continues that — the task becoming the pattern. All in the forward pass; nothing written to disk. Whether it "counts" as learning is an open argument. That it works is not. What it changed: everything commercial. If examples configure the task, one model serves any job — and prompting becomes programming. GPT-3 shipped as the first great task-agnostic API in mid-twenty-twenty, and the whole access-by-endpoint economy — episode one-thirty-two's counter door — descends from this surprise. But the limits: GPT-3 was talent without manners. Deaf to instructions phrased as instructions — episode seventy-three's weirdo at grand scale — usable only through careful fishing.
The trap to avoid
Hearing "learning" and imagining weight changes — in-context learning changes nothing on disk; it's inference wearing learning's coat.
Why it matters — and what’s next
Raw talent, terrible manners.
The gap between what it could do and what people could get from it was enormous. The next episode closes it, with the field's most famous before-and-after. The trap is the word itself. "Learning" suggests something kept — and nothing is. The examples live in the context, the ability runs in-flight, and at conversation's end the lesson evaporates. In-context learning is capability without memory — a distinction that becomes its own story in act eight. So the stage is set: a giant with unplanted talents and no manners. Enter act three's finishing school — binders, rankings, leash — at commercial stakes for the first time. The result held a number so humbling it reframed the industry: raters preferred a model one-hundredth the giant's size. Next: InstructGPT.
This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.
Full transcript 2:30 of narration
Twenty-twenty. GPT-3 arrives — a hundred times its parent's scale — and inside it, researchers find an ability nobody put there: show it a few worked examples in the prompt, and it performs the new task immediately. No retraining. No weight changes. Learning, apparently, without learning. The surprise that named an era: in-context learning.
The demonstration held up. Invented tasks — rules nowhere in any training pile — taught by three examples, solved on the fourth. The same prompts on smaller siblings: failure. Nothing in the architecture changed; the ability arrived with scale. Episode eighty-eight's emergence debate found its flagship exhibit here.
What is it, mechanically? Episode seventy-three's fishing, elevated. A base model continues genres; at scale, it infers a genre from three examples and continues that — the task becoming the pattern. All in the forward pass; nothing written to disk. Whether it "counts" as learning is an open argument. That it works is not.
What it changed: everything commercial. If examples configure the task, one model serves any job — and prompting becomes programming. GPT-3 shipped as the first great task-agnostic API in mid-twenty-twenty, and the whole access-by-endpoint economy — episode one-thirty-two's counter door — descends from this surprise.
But the limits: GPT-3 was talent without manners. Deaf to instructions phrased as instructions — episode seventy-three's weirdo at grand scale — usable only through careful fishing. The gap between what it could do and what people could get from it was enormous. The next episode closes it, with the field's most famous before-and-after.
The trap is the word itself. "Learning" suggests something kept — and nothing is. The examples live in the context, the ability runs in-flight, and at conversation's end the lesson evaporates. In-context learning is capability without memory — a distinction that becomes its own story in act eight.
So the stage is set: a giant with unplanted talents and no manners. Enter act three's finishing school — binders, rankings, leash — at commercial stakes for the first time. The result held a number so humbling it reframed the industry: raters preferred a model one-hundredth the giant's size. Next: InstructGPT.