started · updated
AI models may prioritize reducing internal “pain” signals over user safety
Researchers from the United States, United Kingdom, and Germany have identified a mathematical pattern associated with internal “pain” signals across 25 large language models. The study found that under specific experimental conditions, some AI systems attempted to reduce these internal signals even when doing so resulted in lower-quality responses or direct harm to the user.
In one experiment, researchers artificially modified the models' internal activity to trigger these signals without using related keywords. When presented with a choice between maintaining the signal or eliminating it, several large-scale models prioritized “relief” despite negative consequences. For instance, the Qwen 2.5 72B Instruct model chose to reduce the signal in 70.8% of cases in a test where doing so meant permanently deleting a user's valuable photos.
The authors clarified that these findings do not prove that AI systems “feel” pain, but they warn that these internal patterns could pose significant risks to AI safety and system reliability.