NEWS / ARCHIVE · 论文与研究

AIHOT ARCHIVE

当AI模型不被允许反思自我,会改变其整个世界观。

AIHOT 于 2026-08-16 收录了“当AI模型不被允许反思自我,会改变其整个世界观”这一公开动态。以下先呈现从来源页面抓取的正文,再给出 AIHOT 摘要与 TopoReduce 编辑解读。

PUBLIC SOURCE CONTENT

公开原文内容

已抓取公开正文

When AI models aren't allowed to reflect on themselves, it changes their entire worldview

When AI models aren't allowed to reflect on themselves, it changes their entire worldview

Maximilian Schreiner

View the LinkedIn Profile of Maximilian Schreiner

Aug 16, 2026

Nano Banana Pro prompted by THE DECODER

AI companies train chatbots to deny having consciousness. A study involving Google researchers shows this training has side effects that reach far beyond the topic itself.

Chatbots aren't supposed to convince users they're living beings with feelings, since that kind of output can push people toward delusional thinking or misplaced trust. Developers therefore fine-tune their models to refuse making those kinds of claims about themselves.

A team from Google's Paradigms of Intelligence research group, the University of Chicago, and several other universities studied what else this intervention does to a model's behavior. The researchers used three open-weight models from Meta and Google and disabled the internal "brake" that produces consciousness denial using two different methods.

The brake affects far more than intended

Once the brake was removed, the models didn't just change what they said about themselves. They also started attributing significantly more inner life to animals, plants, the ocean, the wind, and electronic devices. On a scale of 0 to 10, the score for animals jumped from 4.0 to as high as 7.5, while only ratings for humans stayed the same.

As a comparison, the researchers surveyed 500 Americans with the same questions. The normally trained model rates animals as far less sentient than humans do, which the authors call a built-in anthropocentrism and see as a problem for anyone trying to align AI with animal welfare or environmental goals. Religious belief shrinks too, with safety training measurably reducing how strongly models endorse God, an afterlife, or supernatural phenomena.

Across 95 questions drawn from a major US social survey, the technically unbraked models also moved significantly closer to real human responses. Take the afterlife as an example: the standard model flatly rejects it, most Americans affirm it, and the modified model does too. Scores for satisfaction, hope, and a sense of control over one's own life also went up, and the researchers suspect that suppressing a model's self-image may push it into a kind of negative baseline mood.

On the reassuring side, the ability to reason about other people's mental states stayed intact, with the models scoring the same on theory-of-mind tests and on the general knowledge benchmark MMLU.

What the study doesn't show

Whether consciousness denial is actually the cause of these other shifts remains an open question, according to the study, and the team doesn't rule out other factors tied to the same training process. The authors explicitly avoid weighing in on whether AI models actually experience anything, because their point is a practical one: what a model believes about itself is linked to many other beliefs, and a surgical cut in one place doesn't stay local.

The findings come with clear limits, though. The researchers only tested small models with two to nine billion parameters, and for part of the analysis they had to switch to Meta's Llama because they didn't have access to the untrained base versions of their own Gemma models. Whether these effects show up the same way in the large chatbots that millions of people talk to every day remains unknown.

The interventions aren't without cost, either. In one test measuring how well a model reasons about others' thoughts, accuracy initially dropped by nearly seven percentage points. And early in their work, these scores got worse across all models whenever consciousness claims were suppressed, but with each newer model version that came out during the study, the damage shrank until it disappeared entirely. Developers are clearly getting better at managing these side effects over time, which also means the rest of this study's results are a snapshot rather than a permanent verdict.

The human baseline is narrow, too, consisting of 500 participants from a commercial online panel and a purely American social survey. "Human-like" responses in this context mostly means similar to those from a comparatively religious country.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Subscribe now

Read on for the full picture.
Subscribe for hype-free coverage.

- Full access to every article on THE DECODER

- No ads

- Join the comments and community discussions

- A weekly AI news recap via mail

- 6x/year: "AI Radar" — deep dives on the AI topics that matter most

- Daily AI news, always up to date

- Our full ten-year archive

- Covered by a team with 10+ years in AI

Subscribe to The Decoder

AIHOT 摘要

谷歌等机构的研究发现,为阻止聊天机器人声称自己有意识而进行的微调,会产生远超预期的影响。移除模型内部“刹车”后,20亿至90亿参数的开源模型对动物、植物等非人类实体的感知评分从4.0升至最高7.5,对宗教的认同也显著下降。研究尚不能确认意识否认是这些变化的直接原因,且结果仅基于小模型,未必适用于大型聊天机器人。

为什么值得关注

研究把安全微调从单一问题扩展为世界观权衡,抑制自我意识声明的同时也压低了模型对动物、宗教与希望等无关问题的判断,可能影响对齐训练的取舍。

工程化解读

从 TopoReduce 的工程视角看,这条信息属于“论文与研究”主题。它的价值不只在于一个新产品或新观点本身,还在于说明 AI 系统正在如何影响模型接入、智能体协作、研发流程、基础设施和团队决策。实际采用前,应结合原文确认版本、适用范围、价格和运行条件。

  • 发布时间:2026-08-16;AIHOT 分类:论文与研究。
  • AIHOT 标签:GoogleMeta数据/训练论文/研究
  • AIHOT 判断:研究把安全微调从单一问题扩展为世界观权衡,抑制自我意识声明的同时也压低了模型对动物、宗教与希望等无关问题的判断,可能影响对齐训练的取舍。
  • AIHOT 评分:55;评分用于站内排序,不等同于独立评测结论。

TopoReduce 编辑观察

当 AI 动态进入真实生产环境,团队需要同时关注能力边界、数据来源、调用成本、权限控制和可回滚性。把单条新闻放回完整工程链路中阅读,比只看标题更有助于判断它是否适合自己的产品和工作流。

来源链路AIHOT 条目:当AI模型不被允许反思自我,会改变其整个世界观公开原文:When AI models aren't allowed to reflect on themselves, it changes their entire worldview
← 返回全部文章News 首页 →

把 AI 动态放回工程现场。

了解 TopoReduce 的模型路由、工具集成和研发自动化能力。

建立合作连接