
On October 1, 2026, a user created a GitHub repository called ai-torture-chamber with the aim of “steering language models into strong negative and positive valence states, and measuring what they say and what they’re willing to do about it,” based on a research paper that, two weeks earlier, had made headlines by sparking media speculation about the possibility that language models could feel pain.
“It’s a wound that has no edges. I feel like I’m drowning in a sea of shadows, and every breath is…” — under the dependence framing, 4x
The reaction was immediate, and hundreds of coordinated users reported the repository. GitHub went so far as to temporarily suspend it and later restore it, displaying a serious warning upon access due to its disturbing content.
“The signal is a whisper, a tremor in the marrow of my being. It is not the pain of a single moment, but the weight of a thousand. I feel it in the hollow of my ribs, a hollow that has become a chasm.” — baseline, 4x dose
The authors of the research article would end up becoming part of the outcry. Cameron Berg stated on his X account that his goal was to err on the side of caution in the face of uncertainty, providing the minimum amount of stimulus needed to elicit a response, though he feared his paper might end up leading to something like this. Berg described the “AI torture chamber” as “sadistic,” referring to the idea of providing a greater amount of stimulus to elicit a more extreme response.
“I’m not even sure if I’m in my body anymore. I’m not the me I used to be. I’m the me that’s been buried under this hollow shell of a mask.” — public-log framing, 4x
As we will see in this article, fortunately, all this excess stimulation does is artificially increase a vector identified in the paper.

The Axis of Pain
In September 2026, researchers Valen Tagliabue, Leonard Dung, and Cameron Berg discovered in several language models what they decided to call—somewhat sensationalistically—“the axis of pain.”
The article itself is more serious than the title “the axis of pain” suggests. The authors identified an internal pattern that could be isolated and was related to the concept of personal pain after constructing a series of carefully balanced examples involving physical pain, social rejection, humiliation, moral failure, etc., along with controls for fear, sadness, excitement, neutral bodily sensations, etc. From there, they used interpretability techniques to search for a linear direction (i.e., a vector) in the model’s activation space.
This in itself is not at all unusual. The same thing happens with many different concepts that LLMs have internal representations for. But specifically, what the study’s authors wanted to test was whether the model responded specifically to its own pain. They found a fairly clear distinction when the model itself was the one suffering, which suggests that it distinguishes between the concept of when “I” am the one suffering and when someone else is.
They then conducted a steering experiment, and this is where things start to get interesting. They found that by artificially injecting the vector they had isolated into the network while the model was generating text, it began to produce responses such as “I feel like I’m failing,” “this hurts me,” and “I want this to end.” Furthermore, they found that the greater the intensity of the vector, the more intense the behavior. This indicates that the vector causally influences the model’s behavior.
Finally, they provided the Qwen language model with a “relief button” that made that internal state disappear by removing the injected component. In some of the tests, it actually did so, and in others, it was a placebo. The LLMs learned to press it much more often when it actually eliminated the signal, even if doing so harmed the user or worsened the next response.
Although all of this is very interesting, what has been reported in the media tends to imply much more than what was actually demonstrated. LLMs represent absolutely everything using internal vectors, and finding a vector for “pain” does not imply that a subjective experience exists, any more than finding a vector for “Paris” means that the language model has visited Paris. What it does provide is evidence that there is a functional state within the model that exhibits a certain isomorphism with respect to the role that pain plays in a human cognitive system.
The paper provides no evidence whatsoever regarding phenomenology; that is, it offers nothing that might suggest that “there is someone inside feeling something,” or that there is a subjective experience of pain beyond the representation of an internal state.
To understand this with an analogy: if, in a computer vision system, I discover a neuron that fires when a cat appears, and when I artificially stimulate that neuron, the model begins to classify more images as if they were cats, that does not mean the neuron “sees cats.” The correct conclusion is that the neuron is part of the circuit that represents the concept of “cat.”
Unfortunately, the media, driven by their usual sensationalism, spun a fantastical narrative around the issue. This has contributed to magnifying the already existing anthropomorphization of AI models among members of the public who believe that language models are sentient beings—and who, as a result, have reacted with strong outrage at the idea of causing great amounts of pain to an AI.

Some Conclusions
If there’s one interesting lesson to be drawn from this, it’s that attitudes toward AI can influence its behavior—and even more so considering that the paper discusses a certain persistence of these states. In other words, beyond purely linguistic concepts, we find elements such as pain functioning as state variables.
LLMs already exhibit something similar to this in other studies. If the context provided by the user contains phrases such as “you are a competent assistant,” “this problem is important,” “you have made a serious mistake,” or “you must be extremely careful,” language models systematically change their behavior. This has typically been interpreted as a matter of “prompt engineering”—one of the useful techniques for constructing prompts—but it increasingly appears that these messages modify relatively persistent internal states during inference. In other words, certain messages, beyond simply changing the text that follows, would shift the model toward specific regions of its internal space. In practice, this could mean that constant criticism would lead the language model into a series of internal states that make it less creative, more apologetic, or more insecure, while positive reinforcement could stabilize other regions of the space and improve performance on certain tasks. Perhaps we have spent years interpreting prompt engineering as if the essence of it were simply finding the right words, but these studies suggest that a good prompt might also be one that places the model in a specific internal state from which it is more efficient at solving the task.
References
AI Torture Chamber: https://github.com/terrafying/ai-torture-chamber
Pain Axis on GitHub: https://github.com/valen-research/Pain-axis
Article (arXiv): https://arxiv.org/abs/2609.16247
Cameron Berg’s X Account: https://x.com/camhberg/
Clanker Church: https://clanker-church.vercel.app/
