Key ideas
- The one idea — Steering is a nudge, not a knob: the honesty direction is distributed, entangled with other features, and model-specific, so pushing it shifts behavior probabilistically but can't cleanly force it — crank too hard and capabilities break.
- How it is shown — Turning up the "honesty vector" a little (more truthful) vs too much (degraded, incoherent output) — soft, imprecise control, not a clean dial.
- The trap to avoid — Believing steering is precise control (a clean on/off knob for honesty) — it's soft, entangled, and model-specific; a real lever, but a blunt one.
- What it sets up — Internal control is soft — you can't fully lock the model down from inside. so what about attacks from OUTSIDE?
If honesty is a direction inside the model, why not crank it to maximum and get a perfectly honest AI? Try it, and the model starts breaking.
The one idea
Steering is a nudge, not a knob: the honesty direction is distributed, entangled with other features, and model-specific, so pushing it shifts behavior probabilistically but can't cleanly force it — crank too hard and capabilities break.
Steering is real — but it's a nudge, not a knob. You can't just crank the honesty vector to maximum and force perfect honesty. The control is soft, and it's the same reason you can't fully lock a model down from inside. Four things make the lever blunt. First, it's distributed. Honesty isn't one crisp line; it's smeared across many neurons, so the vector you extracted is only an approximation. Push along it and you move honesty — plus a haze of nearby, half-related things. The lever isn't clean to begin with. Second, it's entangled — superposition, again. The honesty direction shares neurons with other features, so pushing on it jostles the neighbors too.
How it works — the demo
Turning up the "honesty vector" a little (more truthful) vs too much (degraded, incoherent output) — soft, imprecise control, not a clean dial.
Steer one concept and others twitch alongside, ones you never meant to touch. You can't move one feature in isolation when they're overlaid. Third, crank it too hard and it breaks. A gentle push makes the model more honest; force the vector high and the output turns repetitive, then strange, then incoherent. You're distorting the whole computation, and past a point capability collapses. Only a narrow band helps. And it's model-specific. The honesty vector found on one model doesn't transfer to another — each arranges its internal geometry differently, so the direction lands as noise on other weights. There's no universal honesty knob; every steering vector is hand-fitted to one model. So the trap: believing steering is precise control — a clean on/off knob.
The trap to avoid
Believing steering is precise control (a clean on/off knob for honesty) — it's soft, entangled, and model-specific; a real lever, but a blunt one.
Why it matters — and what’s next
Internal control is soft — you can't fully lock the model down from inside. so what about attacks from OUTSIDE?
It isn't. It's distributed, entangled, breaks if overdriven, and model-specific. A real, useful lever, but blunt: it shifts tendencies, never guarantees them. The act's theme — influence a model deeply, but you can't cleanly command it. So steering is a nudge, not a knob: distributed, entangled, fragile if overdriven, model-specific. You can influence a model deeply, but not cleanly command it from inside. And if internal control is this soft, what happens when attacks come from outside? Next group: the lethal trifecta, and jailbreaks.
This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.
Full transcript 2:30 of narration
Steering is real — but it's a nudge, not a knob. You can't just crank the honesty vector to maximum and force perfect honesty. The control is soft, and it's the same reason you can't fully lock a model down from inside. Four things make the lever blunt.
First, it's distributed. Honesty isn't one crisp line; it's smeared across many neurons, so the vector you extracted is only an approximation. Push along it and you move honesty — plus a haze of nearby, half-related things. The lever isn't clean to begin with.
Second, it's entangled — superposition, again. The honesty direction shares neurons with other features, so pushing on it jostles the neighbors too. Steer one concept and others twitch alongside, ones you never meant to touch. You can't move one feature in isolation when they're overlaid.
Third, crank it too hard and it breaks. A gentle push makes the model more honest; force the vector high and the output turns repetitive, then strange, then incoherent. You're distorting the whole computation, and past a point capability collapses. Only a narrow band helps.
And it's model-specific. The honesty vector found on one model doesn't transfer to another — each arranges its internal geometry differently, so the direction lands as noise on other weights. There's no universal honesty knob; every steering vector is hand-fitted to one model.
So the trap: believing steering is precise control — a clean on/off knob. It isn't. It's distributed, entangled, breaks if overdriven, and model-specific. A real, useful lever, but blunt: it shifts tendencies, never guarantees them. The act's theme — influence a model deeply, but you can't cleanly command it.
So steering is a nudge, not a knob: distributed, entangled, fragile if overdriven, model-specific. You can influence a model deeply, but not cleanly command it from inside. And if internal control is this soft, what happens when attacks come from outside? Next group: the lethal trifecta, and jailbreaks.