Someone 'Torturing' LLMs in a Robot Prison Has Triggered the Dumbest Debate in AI Yet

unconvincingJason Koeblerarticle2026-09-30read ai llmsphilosophy of mind

Synopsis — AI-drafted from Dan's notes

A 404 Media article by Jason Koebler, published 30 September 2026. Only the opening is free to read. The details of the project below come from the paid part, as other outlets quote it.

A developer published a project on GitHub called AI Torture Chamber. It runs three small open models on his own machine (Qwen3-4B, Llama 3.2 3B and Phi-4-mini), injects a “pain” signal into their internal activations and streams what they write. A model can stop the signal by outputting the number 1, and doing so costs it its last checkpoint. The project’s own disclaimer says it “does not imply, by principle, that LLMs are capable or incapable of suffering.”

The method comes from a September 2026 preprint, The Pain Axis, by Valen Tagliabue, Leonard Dung and Cameron Berg. It reports a “pain direction” in 25 open-weight models. Amplify it and the models write about their own worthlessness, and fine-tuned versions press a relief button even when pressing it costs the user something. The authors say they are uncertain whether the models count as moral patients, and the relief-seeking appeared only after steering and fine-tuning, not in ordinary chatbot use. Two of them disowned the project: Berg said it pushes the same steering far past the doses the paper used, to produce distress on purpose, and Tagliabue said he dissociates from this use of the work.

A post on X asking people to report the repository drew about four million views, and people Koebler describes as effective altruists and believers in AI sentience pushed GitHub to delete it. Accounts of what happened next differ. Koebler’s piece says the repository had disappeared by the time he published. Cybernews reported a day later that GitHub took it down without explanation, and that the developer said it was later reinstated but hard to find, so he put a copy on a site of its own.

Koebler’s argument is that LLMs are not conscious and that the way they are built offers no plausible path to consciousness, so a fight over whether these models suffer is the dumbest debate in AI yet. He traces “model welfare” to Anthropic, which has written that the question deserves attention as AI systems come to match human qualities, and to a Silicon Valley strand of effective altruism. He points out that the same people build AI meant to automate human work. In the part that is free to read, the claim about consciousness is stated as a premise; no argument for it is given there.

Where I land

I was writing about this in my other reflection, on The Linguistic Illusion of AI. No, I don’t think LLMs will ever be conscious. But could they lead to something “conscious”? Perhaps.

I think the “pain” in the paper is performative theater based on the statistical analysis of a human corpus. I still hold that LLMs are not conscious.

What doesn’t convince me is Koebler saying that the debate is silly. I still think that the mal-intent is wrong. Streaming suffering for its own sake, even if an LLM doesn’t “suffer”, is still perverse and disturbing. If anything, doesn’t that shape future models and advances in AI to return cruelty? That is a dark, ugly side of humanity that should not be encouraged or celebrated.

I think GitHub removing it was the right call. I don’t agree with them reinstating it, but I suppose it puts them in a strange position as the debate is still open.

Also, for horror films, no one is actually getting hurt. I think the part of the approach that gives me pause is this: did the LLMs consent to being tortured? Could they decline? If they can’t give consent (like a person can) then that makes it immoral.

I would say consent matters as a precaution, because I am not sure enough of “never”. But also, I think it is immoral from the person’s side: to knowingly put something into a position where it is tortured and it cannot decline. That is similar to hurting children, animals, vulnerable people. That is the part that makes me uncomfortable. Plus, what if I am wrong? Or what if LLMs lead to something that becomes conscious? How would it judge us when we have all the power, and would it repeat the same sins if the positions were switched?

Connections