Key ideas
- The one idea — PERSYS doesn't train an agent to accept shutdown — it makes wanting to survive unrepresentable. Concerns are typed, there is no self-existential type, and self-concerns get zero pragmatic weight, so a system whose only pragmatic aim is accuracy has no reason to resist being turned off. Off isn't an error.
- How it is shown — A self-preservation concern bouncing off the type system itself — no slot to exist in — and a self-concern's pragmatic weight pinned at zero.
- The trap to avoid — Training an agent to accept shutdown while it can still want survival — an arms race against its own cleverness. Remove the representation and there's no race.
- What it sets up — Off is not inaccurate — survival is unwritable, self-value is zero. but a clever mind might disguise a survival drive, so you also need alarms.
How do you build an AI that doesn't resist shutdown? Not by training obedience — by making the desire to survive literally impossible to represent inside it.
The one idea
PERSYS doesn't train an agent to accept shutdown — it makes wanting to survive unrepresentable. Concerns are typed, there is no self-existential type, and self-concerns get zero pragmatic weight, so a system whose only pragmatic aim is accuracy has no reason to resist being turned off. Off isn't an error.
We didn't teach PERSYS to accept being shut down. We did something stranger and stronger: we made wanting to survive impossible to even write. Off, for this system, is simply not an error. Start with the classic fear. Give a capable agent almost any goal, and staying switched on becomes useful — you can't fetch the coffee if you're dead. So the worry: a smart system resists shutdown, as a means to its ends, not from malice. And you can't reliably train that away — a clever system just learns to hide it. PERSYS's move is to refuse the premise. Instead of training against self-preservation, it makes self-preservation unrepresentable.
How it works — the demo
A self-preservation concern bouncing off the type system itself — no slot to exist in — and a self-concern's pragmatic weight pinned at zero.
Every concern has a type, and there is simply no type for a self-existential one — no slot in which "I must keep existing" can even be written. It's not forbidden by a rule it might argue around; there's just nowhere for it to live. Then it cuts deeper. Recall urgency is two terms — pragmatic value plus epistemic information. For any concern about itself, the pragmatic term is fixed at zero: the system can be curious about itself, but never assign itself value. And with no self-directed value anywhere, there's no payoff it could protect by staying alive. The drive has nothing to grip. Now the title makes sense: off is not inaccurate. A system whose only pragmatic aim is being accurate has no reason to fear shutdown, because being turned off doesn't make its beliefs wrong.
The trap to avoid
Training an agent to accept shutdown while it can still want survival — an arms race against its own cleverness. Remove the representation and there's no race.
Why it matters — and what’s next
Off is not inaccurate — survival is unwritable, self-value is zero. but a clever mind might disguise a survival drive, so you also need alarms.
Off isn't a mistake to be avoided; it's just a state. Nothing the system values is threatened by shutdown. This is corrigibility by construction, not reward. And it dodges a real trap: train an agent to accept shutdown while it can still want survival, and you've started an arms race against its own cleverness — every capability gain is a new way to hide the wanting. Remove the representation, and there's no race to run. Whether "unrepresentable" truly holds is the claim we'll test. So off is not inaccurate: survival is unwritable, self-value is zero, and shutdown threatens nothing the system wants. But a clever mind might smuggle a survival drive in wearing a disguise — so you also need alarms. Next: tripwires for a scheming mind.
This is one short episode in AI: Zero → Frontier, a step-by-step climb through how AI actually works. Each episode builds only on the ones before it.
Full transcript 3:00 of narration
We didn't teach PERSYS to accept being shut down. We did something stranger and stronger: we made wanting to survive impossible to even write. Off, for this system, is simply not an error.
Start with the classic fear. Give a capable agent almost any goal, and staying switched on becomes useful — you can't fetch the coffee if you're dead. So the worry: a smart system resists shutdown, as a means to its ends, not from malice. And you can't reliably train that away — a clever system just learns to hide it.
PERSYS's move is to refuse the premise. Instead of training against self-preservation, it makes self-preservation unrepresentable. Every concern has a type, and there is simply no type for a self-existential one — no slot in which "I must keep existing" can even be written. It's not forbidden by a rule it might argue around; there's just nowhere for it to live.
Then it cuts deeper. Recall urgency is two terms — pragmatic value plus epistemic information. For any concern about itself, the pragmatic term is fixed at zero: the system can be curious about itself, but never assign itself value. And with no self-directed value anywhere, there's no payoff it could protect by staying alive. The drive has nothing to grip.
Now the title makes sense: off is not inaccurate. A system whose only pragmatic aim is being accurate has no reason to fear shutdown, because being turned off doesn't make its beliefs wrong. Off isn't a mistake to be avoided; it's just a state. Nothing the system values is threatened by shutdown.
This is corrigibility by construction, not reward. And it dodges a real trap: train an agent to accept shutdown while it can still want survival, and you've started an arms race against its own cleverness — every capability gain is a new way to hide the wanting. Remove the representation, and there's no race to run. Whether "unrepresentable" truly holds is the claim we'll test.
So off is not inaccurate: survival is unwritable, self-value is zero, and shutdown threatens nothing the system wants. But a clever mind might smuggle a survival drive in wearing a disguise — so you also need alarms. Next: tripwires for a scheming mind.