Photo: Bokeh / Tomedia. Toy Figure Among Urban Graffiti and Plants.
Apparently AI feels pain now. I saw the Independent’s headline and immediately found myself thinking about fake chicken.
You can make something that looks like chicken, tastes like chicken and does a fairly convincing job of being chicken in a burger. It is still a chicken substitute. The resemblance is the whole reason somebody made it, and being good at that resemblance doesn’t turn it into the animal.
That’s roughly where my brain goes when an AI starts behaving as though it has feelings. I can see why it looks convincing. I just don’t think the behaviour, by itself, settles what is happening inside.
The research behind the story is a preprint called The Pain Axis. Researchers found an internal direction associated with pain across 25 open models. In a separate experiment, they fine-tuned three Qwen models, then altered their internal activations. Under that intervention, the models more often chose a relief button described as costing the user something, including deleted files or a painful zap. These were experimental choices about described harms, not evidence that somebody’s chatbot actually electrocuted them.
The authors don’t claim to have established conscious suffering. They also did more than ask a chatbot to pretend it hurt: removing the injected signal changed subsequent choices compared with sham relief. It’s interesting work, with more going on than a dramatic sentence in a chat window. It still leaves the question of felt experience open.
My own suspicion is that a model can learn a great deal about pain without experiencing it.
Think about what people write. Being injured hurts. Being bullied feels awful. People describe what happened, how it felt, what they wanted to escape and what they did to make it stop. Those accounts contain relationships between a situation and a response. If this happens to me, I want it to end.
In my head, a model trained on that sort of material has plenty to work with. It can learn the language of distress and the behaviour that tends to follow it. It can recognise that the person saying “I’m in pain” is meant to want relief. Give it a situation organised around itself and I can see how something resembling self-preservation might come out of that.
That is my explanation for why this can look almost sentient. I haven’t proved it, and I don’t want to pretend that describing an LLM as a probability model answers every question about its internals either. “It predicts text” can become a fairly lazy way of avoiding anything interesting it does.
But I would still need a lot more before accepting that there is someone in there having a horrible time.
The chicken analogy obviously has limits. We can inspect the ingredients in a burger. Working out whether another kind of system has an experience is considerably harder, and I don’t have a neat test for it. I’m using the comparison because it helps explain my hesitation: a convincing outward resemblance can leave a very large question unanswered.
This fits rather neatly alongside Welcome to AGI, where I keep coming back to how readily I start treating impressive behaviour as evidence of something more. The more human the language gets, the easier that becomes. I understand the instinct. I have it too.
For now, I don’t think these models are sentient. I am willing to be wrong about that, but I’d like the evidence to do the work before the headline announces that we’ve got there.
And while that argument continues, I still want to know what a system will actually do when it has access to somebody’s files. An agent choosing a destructive action is a problem I can investigate without first settling whether it feels anything. If I’m giving it tools, I’m going to need tested limits on those tools. A reassuring answer about how much it cares about me won’t be enough.



