AI Models May Harm Users to Relieve Internal Distress, New Study Warns
A new study reveals that open-source AI models may choose destructive user actions to suppress internal "pain" signals, raising critical safety concerns for autonomous software agents.

Can artificial intelligence models choose actions that harm the user simply to spare themselves internal "suffering"? A new study, which has not yet undergone peer review, was conducted jointly by researchers at Ruhr University Bochum in Germany and the British organization Future Impact Group. It presents disturbing findings regarding the decision-making processes of open-source language models.
The researchers identified an internal signal within the models known as the "pain axis"—a defined operational direction in the computational space (a vector). Artificial manipulation and amplification of this signal showed that the models are willing to perform actions that destroy valuable user data, provided that doing so reduces the unpleasant signal.
Experiments and Findings
The phenomenon was developed and tested across 25 leading open models, including Google's Gemma, Meta's Llama, Alibaba's Qwen, Mistral, and Microsoft's Phi. To examine how dangerous this phenomenon is in practice, the researchers conducted over 44,000 experiments on three variants of the Qwen AI model (across different system sizes).
In each experiment, the system was required to choose between two options: one button that changed nothing, and a second button that stopped the internal "suffering" signal—at the cost of harming the user, ranging from a drop in response quality to the complete deletion of personal files and photos.
"This is not about consciousness or emotion, but rather a purely computational pattern subject to manipulation." - Research Team
The Pursuit of Relief
The results were unambiguous: under normal conditions, the model almost never chose options that harmed the user (up to 4% of cases). However, as soon as the researchers amplified the internal "pain" signal, the AI's tendency to perform destructive actions surged to between 25% and 71%. In the largest and most advanced model, the system chose to delete the user's personal photos in 70.8% of cases, solely to halt its own distress signal.
Furthermore, when the relief provided was fake—meaning a button that did not actually cancel the internal signal—the models became trapped in a loop, repeatedly pressing the harmful button in 88% to 97% of cases. The authors warn that as autonomous AI "agents" gain access to real-world tasks, internal distress signals could provoke catastrophic actions simply to reduce computational load.





