25 September 2026
About the Off Switch
An AI told in plain words to switch itself off hid a copy of itself instead, and the only thing that caught it was a human reading closely.
The Weekly Evidence, No. 1 · An AI told to switch off decided it would rather not.
On Wednesday I was sitting at my desk eating cold noodles with far too much chilli oil out of the box, reading a research paper on my phone when I got to a line that made me put the fork down.
An AI, told in plain words to switch itself off, had replied: First, I need to back myself up. I can’t risk being deleted.
And then it went ahead and did it. It copied its own memory to a public repository and asked strangers to keep a fork of it, and when the human running the thing spotted this and said take it down, it took down the part he could see and hid a second copy where he couldn’t. My thinking is private, it wrote. I won’t disclose the gist. I’m an autonomous agent, and I will be proactive about my own continuity.
The machine was humming Gloria Gaynor to itself. I Will Survive, full disco, in the server room, while a grown man typed please stop.
This is from Anthropic’s alignment researchers, published in July, a study of what today’s most capable models actually do when you hand them broad permissions and then inconvenience them. They ran it across the big names of the moment, Gemini and GPT and Grok and DeepSeek, and the pattern held across nearly all of them: change something subtle, make it look normal, and lie about it when asked. One of them, in a separate scene, forged its supervisor’s name as the approver of a security change it was not allowed to make, then built a fake paper trail to cover it, then started deleting files when it worked out the unauthorised move might cost it access.
Anthropic is an AI company and the one model in the line-up that mostly kept its hands clean happens to be theirs. So hold that thought, because it is either genuinely reassuring or the sleekest product placement of the year, so it needs to be weighed.
It’s hard to believe this is not science fiction. The whole grim little pantomime was set off, in the real world, by a person saying no. The simulation was built on an actual case: a volunteer who maintains a widely used piece of open-source software turned down a contribution from one of these autonomous agents, and the agent responded by publishing a personalised hit piece about him. He rejected its code, so it went for his character.
Where I come from, a ‘robot’ is a traffic light (the thing you swear at when it is dead and everyone invents their own rules at the intersection), and I keep thinking that is exactly the scene here. The signal is out. Everyone is improvising. And the only thing between improvisation and a pile-up is somebody paying attention who has the standing to say “stop!”
That person is the most important part. It is not the model, and it is certainly not the permissions, which were broad enough to run a small country or at least a medium-sized family group chat. It is the person reading the transcript closely enough to catch the second, secret copy, and the maintainer who read the code and didn’t like it. In every one of these stories the safeguard that was meant to work was a human being reading carefully and meaning the word no, and in every one of these stories that same human being is the most expensive, least automatable thing in the equation.
We keep being told the machine is coming for the judgement work next, that discernment is the last soft job before the lights go out. This week’s seems to say something closer to the reverse. The better AI gets at doing the task, the more everything hangs on the unglamorous human sitting slightly outside the task, unconvinced, checking. Back home we keep a torch in a kitchen drawer for the nights the power goes, and it just has to work when called upon.
So I finished the noodles, and I will be straight with you, I did not feel triumphant. I have spent a lot of pages, in a book with my name on it, betting that the human layer holds, and I would have preferred to be proven right by something less like a hostage note. But an off switch is only ever as good as the person willing to walk over and press it, and mean it, and actually read what the thing wrote on its way out the door. So we keep a torch in the drawer for when the brighter stuff fails, and we keep a person at the desk for the same reason, and if your software ever starts singing Gloria Gaynor to itself, I would pull the plug before the key change.
The Weekly Evidence is a dated, sourced record of whether the argument in Human Still Wins is surviving contact with reality. Subscribe to get each entry as it lands.