Act 05 · Architectures & the model zoo 2:30 GPT-3, InstructGPT, and the open frontier

InstructGPT: the assistant is born.

RLHF's first commercial landing: a 1.3B model tuned with human feedback was PREFERRED over the raw 175B giant — alignment beat scale, and the recipe became ChatGPT.

Video rendering soonThe cinematic render for this episode is being generated. The article, transcript and key ideas are all here now.

Key ideas

  • The one idea — RLHF's first commercial landing: a 1.3B model tuned with human feedback was PREFERRED over the raw 175B giant — alignment beat scale, and the recipe became ChatGPT.
  • How it is shown — The coronation: base ramble vs instruct answer side by side; the preference vote crowning the hundredfold-smaller model; ChatGPT igniting from the recipe.
  • The trap to avoid — Reading "preferred" as "smarter" — helpfulness and capability are different axes; the school shaped behavior, not knowledge.
  • What it sets up — The recipe escaped the palace.

ChatGPT didn't start with a bigger brain. It started with a 2022 paper where the tiny model beat the giant — on being useful.

The one idea

RLHF's first commercial landing: a 1.3B model tuned with human feedback was PREFERRED over the raw 175B giant — alignment beat scale, and the recipe became ChatGPT.

Twenty-twenty-two. A paper reports a result so humbling it reframes the industry: human raters preferred a one-point-three-billion-parameter model over the raw hundred-seventy-five-billion giant — a model one-hundredth the size, winning on usefulness. The difference wasn't scale. It was school. This is InstructGPT: the day the assistant was born. Remember the before. Ask raw GPT-3 a question and it might answer — or continue it into a quiz, a tangent, an essay. Episode seventy-three's weirdo at scale: brilliance locked behind fishing-craft only practitioners held. The capability existed; the access didn't. What came next was access engineering, and it changed who AI was for. The school you know was applied — episodes seventy-five and seventy-six, at commercial stakes for the first time.

How it works — the demo

The coronation: base ramble vs instruct answer side by side; the preference vote crowning the hundredfold-smaller model; ChatGPT igniting from the recipe.

Binders taught answer-shape; preference-pointing built the judge; reinforcement, leashed, tuned toward it. Out came something new: models that heard an instruction as an instruction. The mechanics weren't novel to you. The consequence is history. Then the number. Blind pairwise preference: outputs side by side, judges choosing. The schooled one-point-three-billion beat the raw hundred-seventy-five-billion — not on knowledge, on usefulness: the question answered, in the shape needed, without the weirdness. A hundredfold of scale, outweighed by finishing. The ignition followed. The same recipe on a stronger sibling, wrapped in a chat window, shipped as ChatGPT in November twenty-twenty-two — the fastest consumer adoption of its era. What actually shipped that day: not a new brain.

The trap to avoid

Reading "preferred" as "smarter" — helpfulness and capability are different axes; the school shaped behavior, not knowledge.

Why it matters — and what’s next

The recipe escaped the palace.

A finishing school, productized. The assistant you've talked to since — episode seventy-four's mask — was cast here. The trap: reading "preferred" as "smarter." Two axes. Capability — what it can do — still tracked scale; the giant knew more. Helpfulness — how accessibly it serves — is what the school built. The upset was the second axis winning human hearts. A charming model isn't necessarily a capable one, and act nine lives in that gap. And the recipe didn't stay in the palace. Once its mechanics were published and open weights arrived, anyone could graduate an assistant — and the open canopy learned fast. Which brings the family story to its most consequential modern branch: the downloadable frontier. Next: Llama, Qwen, DeepSeek.

This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.

Full transcript 2:30 of narration

Twenty-twenty-two. A paper reports a result so humbling it reframes the industry: human raters preferred a one-point-three-billion-parameter model over the raw hundred-seventy-five-billion giant — a model one-hundredth the size, winning on usefulness. The difference wasn't scale. It was school. This is InstructGPT: the day the assistant was born.

Remember the before. Ask raw GPT-3 a question and it might answer — or continue it into a quiz, a tangent, an essay. Episode seventy-three's weirdo at scale: brilliance locked behind fishing-craft only practitioners held. The capability existed; the access didn't. What came next was access engineering, and it changed who AI was for.

The school you know was applied — episodes seventy-five and seventy-six, at commercial stakes for the first time. Binders taught answer-shape; preference-pointing built the judge; reinforcement, leashed, tuned toward it. Out came something new: models that heard an instruction as an instruction. The mechanics weren't novel to you. The consequence is history.

Then the number. Blind pairwise preference: outputs side by side, judges choosing. The schooled one-point-three-billion beat the raw hundred-seventy-five-billion — not on knowledge, on usefulness: the question answered, in the shape needed, without the weirdness. A hundredfold of scale, outweighed by finishing.

The ignition followed. The same recipe on a stronger sibling, wrapped in a chat window, shipped as ChatGPT in November twenty-twenty-two — the fastest consumer adoption of its era. What actually shipped that day: not a new brain. A finishing school, productized. The assistant you've talked to since — episode seventy-four's mask — was cast here.

The trap: reading "preferred" as "smarter." Two axes. Capability — what it can do — still tracked scale; the giant knew more. Helpfulness — how accessibly it serves — is what the school built. The upset was the second axis winning human hearts. A charming model isn't necessarily a capable one, and act nine lives in that gap.

And the recipe didn't stay in the palace. Once its mechanics were published and open weights arrived, anyone could graduate an assistant — and the open canopy learned fast. Which brings the family story to its most consequential modern branch: the downloadable frontier. Next: Llama, Qwen, DeepSeek.

Model FamiliesAI HistoryAlignment