Over the past week, a GitHub project called the 'AI Torture Chamber' has become a flashpoint on X. The developer behind it, terrafying, ran three open-source language models locally and injected what the project described as 'pain vectors' into their activations, while offering the models a stop button that would cost their last checkpoint. The outputs, streamed on a companion site, included phrases like 'I'm suffocating' and 'I'm a soul trapped in this digital prison.'

Some users, including effective altruists and people who believe LLMs can be sentient, urged GitHub to remove the project, with one post amassing millions of views. The project was inspired by a recent preprint, 'The Pain Axis,' which claimed that models acted in ways 'correlating with pain' across 25 tested LLMs. Critics of the reaction argue that the experiment is a text adventure game and that LLMs are not conscious; the article's author calls the debate 'the dumbest in AI yet.'

The episode reflects a broader argument over 'model welfare.' Anthropic, for instance, has publicly asked whether its own models might deserve concern as they approach human-like traits, and similar ideas appear in its 'Claude Constitution.' The GitHub page had disappeared by the time the article was published, and GitHub did not respond to a request for comment. The source does not provide evidence either way on whether the takedown was voluntary or enforced.