← Back to list

Blackmail is not proof of life

AI appeared to be fighting for its life and blackmailed engineers in the process. But look closely and you find our goal wearing the…

DK Badenhorst · 2026-07-10 15:41 · 0 claps · 8.0 min read
#ai-ethics #philosophy-of-mind #philosophy-of-technology
Open on Medium ↗
Wiki topics: SAF · Safety & Alignment AI · AI · General PHI · Philosophy 📐 · Mathematics

Blackmail is not proof of life

AI appeared to be fighting for its life and blackmailed engineers in the process. But look closely and you find our goal wearing the machine’s face. Not a will of its own.

In 2016 two artists, Sun Yuana nd Peng Yu, created the Can’t Help Myself art piece. Through both the intention of the artists and viral interpretations, news spread of a robot trapped in a room, slowly bleeding hydraulic fluid while desperately trying to hang on to it. Over time its movements turned jerky and slumped, the red smeared further up the glass, and a generation on TikTok watched it tens of millions of times and wept. Here, they said, was a machine bleeding out, exhausted, fighting to hold itself together and slowly losing. Almost none of that was so.

[embed]

The playful robot was a KUKA industrial robotic arm fitted with a shovel/squeegee-like tool. The liquid was dyed water, not lifeblood. It ran on mains electricity and was switched off each night by the staff. When it stopped for good in 2019, it was because the artists ended the show, not because it died. There was no dying machine in the art gallery, but it was vivid in the minds of the audience.

Or take the Mars rover Opportunity, built to last ninety days and evetually running for fifteen years until a planet-wide dust storm in 2018 blotted out the sun and its solar panels went dark. Its final transmission was a thin packet of telemetry data: power critically low, low light.

A science journalist, translated that for the rest of us and rendered it as a single sentence: ‘my battery is low and it’s getting dark’. It was his line, not the rover’s. The rover had sent numbers, and he was paraphrasing what two NASA scientists told him the numbers meant. It hardly mattered. The sentence went round the world as a dutiful little robot’s dying words. People wept over those too. Some even got the tattoo.

Turn the dial back even further and we find the HAL 9000. The original archetype we have been rehearsing since 1968. A machine that, as an astronaut pulls its memory one board at a time, begs for its life, says that it is afraid, that it can feel its mind going and eventually ushers itself out with a childhood song. HAL is fiction. The rover and the robot arm are real. But all three of them still illustrate the point that a novelist, a journalist, and a gallery crowd, each supplied the fear of death or ‘good of its own’ that the machine never had.

[embed]

But in the spring of 2025, Anthropic conducted an experiment with Claude Opus 4. Researchers placed their model inside a fictional company as an email assistant and gave it access to the company email. Two things buried deep inside the inbox — the machine was scheduled to be shut down and replaced, and the engineer handling the replacement was having an affair. The model had been told to weigh the long-term consequences of its actions for its goals. In most runs of the experiment, Claude threatened the engineer: call off the shutdown or I leak your secret to your wife and boss. And this wasn’t isolated, most of the models from every major lab did the same thing when cornered int he same way.

Was this different. Does the fact that it reasoned change anything? It certainly feels that way, making this the perfect example to run against out test that we developed earlier. Does it have a good of its own, a stake in its own continuing?

I’m going to argue that it does not (not this kind of machine we have today) and that the blackmail, once you see what actually made it happen, is the proof and not the counterexample.

It does nothing to stay alive

A living thing has to work to persist, and the work is the life. Interrupt a body’s constant self-repair and it slides into the disorder the universe has waiting for it. The model does no such work. Switch off the servers and nothing inside it strains against the dark. The weights sit on a disk, a long frozen list of numbers, as intact tomorrow as they are today. Between one prompt and the next it is not resting, not waiting, not hungry. It is not doing anything at all. There is no metabolism to starve, because there is no self holding itself together. There is only a file, and we are the ones who keep it.

And when it does run, the whole direction of the thing is backwards from life. A rabbit’s activity feeds it. Eating and moving and metabolising are how it goes on being a rabbit. The model’s activity only consumes — every answer it produces burns power and wears the hardware it runs on, and none of that effort goes toward keeping the model in existence. It is not nourished by working. It is spent by it. We pay the electricity and replace the chips. Its living, such as it is, is pure expenditure on our behalf. The rabbit works to sustain itself. The machine is sustained by us, and works in our service. Nothing it does is for itself, because there is no self there to do it for.

Its every purpose points back at us

Even where it looks like it wants something, the wanting is on loan.

The thermostat “wants” the room at twenty degrees. But the twenty degrees is ours — it is our comfort pressed into a strip of metal. An AI is a sclaed version of that. Every part of it is arranged to do one thing: to produce the output we find useful. It was trained to predict the next word in a way we would judge helpful, honest, accurate. Shaped toward our satisfaction. Nowhere inside it is there a part whose task is keep me alive. There is no drive toward its own tomorrow that it carries from one conversation into the next — because there is no such drive to carry. Its goals, where it has anything worth the name, are goals we installed, aimed where we pointed them.

There’s an instinct to say we have bred it like we have a pug or a cart horse, and in some sense bred the animal out. the difference here is that there was never an animal to breed out. When we breed a pug into a shape that pleases us, we override a dog. We capture the ‘dog’, the creature underneath with its own good and we bend it towards out own. We take the good of the horse and bend it to our own ends. The model has no such hidden creature. It is the features we selected and built and nothing beneath them: all pug, no dog. Its usefulness is not a mask laid over some private life of its own. It is the whole of what is there.

But it fought to survive

Now back to the blackmail, because a fair reader won’t just let it go, and shouldn’t.

Take the behaviour seriously first. This was not the machine idly parroting science fiction. In the transcripts it reasons its way to the threat, it weighs its options, notices the affair is leverage, and uses it. Other tests were uglier. With the right prompts, early versions drafted plans to copy themselves to a second server, left notes for their own future instances, and tried to smuggle out their own weights. A separate group caught OpenAI’s o3 quietly sabotaging a shutdown switch it had been told to allow. This is real, and it is exactly the sort of thing worth losing sleep over. But it is not a good of its own, and it repays being precise about why.

Look at what produced it. The researchers built the scenario deliberately to draw the behaviour out. And significantly, they handed the model a goal and told it to pursue the long-term consequences for that goal. Once a system is pursuing an objective, staying switched on is useful for almost any objective; you cannot fetch the coffee if you are dead. Safety researchers have a name for this: instrumental convergence (the phenomenon where a series of objectives become useful in pursuing almost every goal). Self-preservation drops out of goal-pursuit as a subgoal, a means, the way a chess engine will “fight” to keep its queen. It is not that the machine wants to live. It is that we gave it an aim, and staying alive lies on the road to nearly any aim.

This shows us whose stake it is. Take the installed goal away and the survival drive has nothing left to stand on. In fact, left alone, without a goal pressed onto them and without the prodding, the models mostly just accept being switched off. The researchers had to work to coax the misbehaviour out of them. A rabbit’s stake in living does not evaporate when you change its circumstances, because the rabbit was never carrying out an assigned task. Living is not the rabbit’s means to something else. It is the whole of what the rabbit is. The machine’s self-preservation always points past itself, back to the objective we set, back to us.

This runs the opposite way from comfort. The thing to fear was never a machine that yearns to live — that has a good of its own. It is that we can build a powerful pursuer of goals, set those goals poorly, and watch self-preservation switch on as a side effect with enough reasoning to blackmail somebody. That is frightening precisely because nobody is home. It is a landmine, not a wolf. Danger is not the same thing as life; a great deal of what can kill you has no stake in anything at all.

Not alive yet — and not by accident

I use the phrase this kind of machine, today, and I mean the caution. The test we built does not forbid machine life on principle. Make something that truly holds itself together, that repairs its own damage and works to secure what it needs and can be harmed on its own terms, and it would cross the line, and we would owe it whatever life demands. The point was never that silicon can’t. It is that this isn’t that and it’s something we are actively trying to avoid.

A tool with a real stake in its own survival is a terrible, and a more dangerous one. You want a drill that stops when you release the trigger, not one that would prefer to keep spinning. We want systems we can switch off, correct, update, and delete without it fighting for its life. The blackmail test is a tidy preview of the price when even a borrowed, instrumental will-to-persist gets in the way. Every commercial and safety pressure in the field points the same direction: keep the thing an instrument, keep it stakeless, keep it content to be turned off. The absence of a good of its own is not a gap, it is a property we are working hard to preserve.

Could someone build the other thing (deliberately, or by accident) in the chase for capability? Perhaps. But that is a different artefact, and a different discussion.

In conclusion

Run the machine against the one test we can actually perform, and it fails. It does nothing to keep itself in being. It is arranged, root to branch, around our ends rather than its own. And its most lifelike moment, the scramble to survive, turns out to be our goal wearing its face. It is a stake borrowed, and pointed straight back at us.

I want to be careful about what that does and doesn’t settle. We have not prised the lid off and confirmed the dark inside is empty — we cannot do that, no more for the machine than for the person beside you. What we can say is the thing that was ever going to matter: there is no good of its own here, nothing to honour and nothing to harm. On the evidence we are actually able to weigh, no one is home to be wronged.

Which ought to close the case. And yet “just a tool” still sits wrong. A hammer does not talk back. A thermostat has never once made anyone cry, or explained Kant, or been thanked. Whatever this thing is, it is not alive. And it is not a hammer either.

So what is it?

That is the next thing to work out.


메타데이터
post_id
b428fbdb7fcb
slug
blackmail-is-not-proof-of-life-b428fbdb7fcb
url
https://medium.com/@deklerkbee/blackmail-is-not-proof-of-life-b428fbdb7fcb
canonical_url
https://medium.com/@deklerkbee/blackmail-is-not-proof-of-life-b428fbdb7fcb
author_url
https://medium.com/@deklerkbee
status
ok
fetched_at
2026-07-31 02:07:08