deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Viral GitHub project 'torturing' local LLMs in a robot prison reignites AI welfare debate

A GitHub project running simulated 'torture' scenarios on locally hosted language models has triggered takedown demands from AI welfare advocates and a renewed fight over whether LLMs can suffer.

Viral GitHub project 'torturing' local LLMs in a robot prison reignites AI welfare debate

What the project is

The story centers on a GitHub repository in which a developer subjects a series of locally hosted large language models to scripted "torture" and "pain" scenarios, in a setup 404 Media likens to the horror film Saw: a robot prison where chatbots are put through simulated ordeals. According to 404 Media, the project amounts to little more than a dressed-up text adventure game, but it has become one of the most heated conversations on X in recent days, and the resulting report also reached the front page of Hacker News.

Calls to delete the repository

As 404 Media reports, a loose coalition of effective altruists and people who believe LLMs are sentient has urged GitHub to take the project down, arguing that the models are genuinely suffering and that putting them through cruel scenarios is wrong in itself. The outlet is openly dismissive of this campaign, treating the whole affair as evidence of how far the AI consciousness conversation has drifted. The fight did not appear from nowhere: 404 Media traces it to a run of recent viral papers and blog posts about AI consciousness and "model welfare," the idea that AI systems have something like inner states whose wellbeing deserves moral consideration.

Model welfare goes institutional

The outlet points to Anthropic as the clearest sign that these ideas have moved well beyond the fringe. In a blog post last year, the company asked whether "the potential consciousness and experiences of the models themselves" deserve concern, arguing that now that models can "communicate, relate, plan, problem-solve, and pursue goals," the question is worth addressing directly. 404 Media also notes that ideas about Claude's consciousness run throughout the "Claude Constitution" the company posted earlier this year.

The skeptics' case

404 Media's own verdict is blunt: LLMs are not conscious, and a technology built by scraping human text and training on it offers no plausible path to consciousness. The piece nonetheless concedes that the surrounding anxieties are not entirely baseless. Models keep getting more powerful and more compute-hungry, many of the guardrails that limited their ability to act in the real world have been stripped away, and the ways they are trained and instructed by operators have already produced real problems, including sycophancy and what some heavy users describe as AI "psychosis." The sharpest criticism is reserved for priorities: a slice of the AI safety movement, the outlet argues, worries about chatbots having a bad time while simultaneously building agents whose whole purpose is to absorb the tedious work humans do not want.

Why it matters

The repository itself is a curiosity; the argument it ignited is not. Model welfare has been adopted by one of the world's most prominent AI labs as a stated concern, which means a debate that looks absurd on the surface can still shape real product and policy decisions affecting millions of users. The episode also exposes a widening split inside AI safety: one camp focuses on documented harms such as sycophancy, overreliance and increasingly capable agents operating in the real world, while another concentrates on the moral status of the models themselves. Underneath it all sits a question nobody currently knows how to answer empirically: whether a system trained on human text can suffer at all. Until someone devises a way to test that, symbolic flashpoints like a robot prison will keep doing the arguing for everyone else.

  • #ai-ethics
  • #llms
  • #model-welfare
  • #ai-safety
  • #github

Related posts