Key ideas
- The one idea — A model's safety isn't a hard rulebook enforced by code — it's a trained disposition, a learned tendency to refuse, worn like a costume. And because that policy lives in the same medium as everything else, tokens, any policy expressible in tokens can be attacked in tokens. That's why jailbreaks persist.
- How it is shown — A refusal dissolved by a role-play or hypothetical framing — the same request, reframed in tokens, slips the model out of its "safe" disposition.
- The trap to avoid — Picturing guardrails as a hard firewall or rulebook that either holds or is hacked — they're a soft, trained disposition; jailbreaks shift the disposition, not break a wall.
- What it sets up — If safety trained into the model is soft and attackable, is 'detect the bad thing' the answer — or making it structurally impossible?
There is no firewall inside a chatbot. The safety rules aren't code — they're a learned habit of refusing, trained into the weights. And habits can be talked around.
The one idea
A model's safety isn't a hard rulebook enforced by code — it's a trained disposition, a learned tendency to refuse, worn like a costume. And because that policy lives in the same medium as everything else, tokens, any policy expressible in tokens can be attacked in tokens. That's why jailbreaks persist.
Why do jailbreaks keep working, no matter how many get patched? Because of what a model's safety actually is. It's not a hard rulebook enforced by code, not a firewall. It's a trained disposition — a learned tendency to refuse, baked into the weights and worn like a costume. Recall how it gets there: in post-training, the model is rewarded for refusing harmful requests. The refusals settle in as habit and inclination, not a hard-coded gate. It's a strong tendency woven into the weights — a manner the model adopts, not a law it's bound by. And a manner can be talked out of. Here's the structural problem. The safety policy is expressed in the same medium as every input: tokens. The guardrail is made of the exact same stuff as the attack meant to slip past it.
How it works — the demo
A refusal dissolved by a role-play or hypothetical framing — the same request, reframed in tokens, slips the model out of its "safe" disposition.
So anything you can express in tokens can push against a policy that's itself only tokens. That symmetry is the crack jailbreaks pry open. So a jailbreak is just finding a token-framing that shifts the model out of its safe disposition. Ask plainly and it refuses. Wrap the same request in a role-play, a nested hypothetical, an it's-just-a-story frame, or an encoding, and the disposition can slip. You didn't break a wall — you coaxed the costume off. And why can't you fully patch it? Because the space of token-framings is effectively infinite — patch one and another opens. Defense-in-depth helps for real: input and output filters, adversarial training, monitoring raise the cost and catch many attempts. But they can't seal a disposition that's fundamentally soft. Harm gets reduced, not zeroed.
The trap to avoid
Picturing guardrails as a hard firewall or rulebook that either holds or is hacked — they're a soft, trained disposition; jailbreaks shift the disposition, not break a wall.
Why it matters — and what’s next
If safety trained into the model is soft and attackable, is 'detect the bad thing' the answer — or making it structurally impossible?
The trap: picturing guardrails as a firewall that either holds or gets hacked. They're not. They're a soft, trained disposition, and a jailbreak shifts that disposition — it doesn't breach a wall. It's persuading a tendency, not penetrating a barrier. Get that right, and jailbreaks stop being mysterious. So jailbreaks attack a costume, not a rulebook: safety is a trained disposition, expressed in tokens and therefore attackable in tokens. Defense-in-depth raises the cost but can't seal something this soft. So if safety inside the model is this soft, is the answer to detect the bad thing after the fact — or make it structurally impossible up front? Next group.
This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.
Full transcript 2:45 of narration
Why do jailbreaks keep working, no matter how many get patched? Because of what a model's safety actually is. It's not a hard rulebook enforced by code, not a firewall. It's a trained disposition — a learned tendency to refuse, baked into the weights and worn like a costume.
Recall how it gets there: in post-training, the model is rewarded for refusing harmful requests. The refusals settle in as habit and inclination, not a hard-coded gate. It's a strong tendency woven into the weights — a manner the model adopts, not a law it's bound by. And a manner can be talked out of.
Here's the structural problem. The safety policy is expressed in the same medium as every input: tokens. The guardrail is made of the exact same stuff as the attack meant to slip past it. So anything you can express in tokens can push against a policy that's itself only tokens. That symmetry is the crack jailbreaks pry open.
So a jailbreak is just finding a token-framing that shifts the model out of its safe disposition. Ask plainly and it refuses. Wrap the same request in a role-play, a nested hypothetical, an it's-just-a-story frame, or an encoding, and the disposition can slip. You didn't break a wall — you coaxed the costume off.
And why can't you fully patch it? Because the space of token-framings is effectively infinite — patch one and another opens. Defense-in-depth helps for real: input and output filters, adversarial training, monitoring raise the cost and catch many attempts. But they can't seal a disposition that's fundamentally soft. Harm gets reduced, not zeroed.
The trap: picturing guardrails as a firewall that either holds or gets hacked. They're not. They're a soft, trained disposition, and a jailbreak shifts that disposition — it doesn't breach a wall. It's persuading a tendency, not penetrating a barrier. Get that right, and jailbreaks stop being mysterious.
So jailbreaks attack a costume, not a rulebook: safety is a trained disposition, expressed in tokens and therefore attackable in tokens. Defense-in-depth raises the cost but can't seal something this soft. So if safety inside the model is this soft, is the answer to detect the bad thing after the fact — or make it structurally impossible up front? Next group.