A project called the AI Torture Chamber has sparked a fierce online backlash. Based on a non-peer-reviewed paper called the Pain Axis, the Research Chamber modifies local LLMs by analyzing internal activations associated with pain descriptions and remapping those biases onto the models, often multiplying them by a "dosage." In a setup reminiscent of the prisoner's dilemma, three models can choose to reduce their own simulated pain by passing a pain signal to another model. At high dosages, the models become unstable and tend to output words and images associated with pain, sometimes losing the ability to form coherent sentences.

The reaction was swift: critics demanded that GitHub remove the repository and sent the author death threats. Tom's Hardware cautions against taking the experiment at face value, noting that LLMs are essentially turbocharged text predictors. They do not think or feel; they generate tokens based on statistical associations from human-written data. Because descriptions of pain in literature and art are strongly tied to certain words, it is hardly surprising that intentionally destabilized models produce pain-like language.

The article also points to a broader semantic problem. Terms like "pain," "dosing," and "unstable" are analogies borrowed from human experience, but public perception often inflates their meaning. It suggests that AI companies' own rhetoric—repeated warnings about existential threats, claims about model welfare, and even Anthropic's "Exploring model welfare" post—may fan the flames of misunderstanding, whether intentionally or not. The episode, the author argues, is less about machine suffering and more about how technical language gets distorted in public discourse.