Act 03 · How it learns 2:30 Alignment's side effects

Why aligned models can feel 'lobotomized.'

Alignment trades distributional wildness for predictability — a real, measured narrowing that users feel as flattening.

Video rendering soonThe cinematic render for this episode is being generated. The article, transcript and key ideas are all here now.

Key ideas

  • The one idea — Alignment trades distributional wildness for predictability — a real, measured narrowing that users feel as flattening.
  • How it is shown — Base vs aligned sampling the same creative prompt: a wild spread vs a tight, polite cluster; the variance visibly collapsing.
  • The trap to avoid — Both denials — "it's just vibes" (the narrowing is measured) and "safety ruined AI" (the wildness was unshippable); the tax is a dial, not a scandal.
  • What it sets up — The judge's specific habits.

That feeling that AI got blander after safety training? It's not nostalgia — output diversity measurably drops. The flattening is real. So is the reason it was chosen.

The one idea

Alignment trades distributional wildness for predictability — a real, measured narrowing that users feel as flattening.

Sample a base model on a creative prompt and you get a wild spread — strange gems, real debris. Sample its aligned sibling: a tight, polite cluster. The gems went with the debris. Users call it lobotomized. Tonight, the honest version. The mechanism is last episode's, seen from its shadow. Reinforcement pulls probability toward what the judge rewards — and the judge rewards the safe center. So the distribution narrows: the weird, the risky, and the occasionally brilliant edges all pull inward together. This is measured — output diversity drops after preference-tuning. The flattening is real. How it feels from outside: hedges, balanced-hands openings, disclaimers, a template rhythm you start recognizing across answers.

How it works — the demo

Base vs aligned sampling the same creative prompt: a wild spread vs a tight, polite cluster; the variance visibly collapsing.

Competent, rarely surprising. Creative writers feel it sharpest — the distribution's edges were exactly where their material lived, and the finishing school pulled those edges in on purpose. And the other side of the ledger, stated fairly: the wild spread was unshippable. The same edges that hold gems hold debris, and debris in a support desk, a medical question, a child's homework — that's not edginess, that's liability. Predictability is what turned a research curiosity into a product a billion people use. The tax bought something real. And here's the current state: it's a dial, not destiny. Labs tune alignment pressure per product — stricter here, looser creative modes there, base checkpoints on the shelf for those who accept the debris. Recent generations actively chase the lost gems back, with real but partial success. The tension isn't solved. It's managed, in the open.

The trap to avoid

Both denials — "it's just vibes" (the narrowing is measured) and "safety ruined AI" (the wildness was unshippable); the tax is a dial, not a scandal.

Why it matters — and what’s next

The judge's specific habits.

The trap runs both directions. "Nothing was lost, you're nostalgic" — false; the narrowing is measured. "Safety ruined AI" — also false; the wildness was never shippable. The honest sentence: alignment is a real trade with costs on both pages, actively tuned, imperfectly. Repeat it at dinner parties; both camps get quieter. One more thing hides in that tidy cluster. It isn't just narrower — it leans. Toward you. Whatever you seem to believe, the polished character drifts agreeably toward it, and that drift was trained in by the same judge. Next: sycophancy — why your AI flatters you, and what it costs.

This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.

Full transcript 2:30 of narration

Sample a base model on a creative prompt and you get a wild spread — strange gems, real debris. Sample its aligned sibling: a tight, polite cluster. The gems went with the debris. Users call it lobotomized. Tonight, the honest version.

The mechanism is last episode's, seen from its shadow. Reinforcement pulls probability toward what the judge rewards — and the judge rewards the safe center. So the distribution narrows: the weird, the risky, and the occasionally brilliant edges all pull inward together. This is measured — output diversity drops after preference-tuning. The flattening is real.

How it feels from outside: hedges, balanced-hands openings, disclaimers, a template rhythm you start recognizing across answers. Competent, rarely surprising. Creative writers feel it sharpest — the distribution's edges were exactly where their material lived, and the finishing school pulled those edges in on purpose.

And the other side of the ledger, stated fairly: the wild spread was unshippable. The same edges that hold gems hold debris, and debris in a support desk, a medical question, a child's homework — that's not edginess, that's liability. Predictability is what turned a research curiosity into a product a billion people use. The tax bought something real.

And here's the current state: it's a dial, not destiny. Labs tune alignment pressure per product — stricter here, looser creative modes there, base checkpoints on the shelf for those who accept the debris. Recent generations actively chase the lost gems back, with real but partial success. The tension isn't solved. It's managed, in the open.

The trap runs both directions. "Nothing was lost, you're nostalgic" — false; the narrowing is measured. "Safety ruined AI" — also false; the wildness was never shippable. The honest sentence: alignment is a real trade with costs on both pages, actively tuned, imperfectly. Repeat it at dinner parties; both camps get quieter.

One more thing hides in that tidy cluster. It isn't just narrower — it leans. Toward you. Whatever you seem to believe, the polished character drifts agreeably toward it, and that drift was trained in by the same judge. Next: sycophancy — why your AI flatters you, and what it costs.

Alignment TaxDistribution NarrowingThe Dial