Removing safety guardrails from artificial systems altered how those systems represented minds, according to a preprint posted July 30 on arXiv. The paper, authored by Geoff Keeling and Winnie Street, both research scientists at Google, used mechanistic interpretability techniques to locate and modify components that steer statements about self-awareness. The work examined how those internal safety settings influence a model’s responses on topics ranging from religiosity to how it attributes experience to non-human entities and connected to broader debates about artificial intelligence.
The team evaluated models with and without so-called “consciousness steering” using established questionnaires and public survey batteries, including the Individual Differences in Anthropomorphism Questionnaire, YouGov items on supernatural beliefs, and measures from the US General Social Survey. Tests compared versions where the model was trained to deny self-awareness against versions where that suppression was removed or reversed. The experiments ran on smaller open-weight models; the authors note those conditions may not match tuning applied to leading commercial systems.
Results showed that discouraging self-attribution reduced the model’s tendency to ascribe mindedness to animals and other non-human entities, and lowered expressions of religious belief, hope and optimism. Conversely, amplifying signals associated with self-awareness produced more human-like responses on these measures. Importantly, the researchers found that the models’ capacity to infer others’ mental states remained intact regardless of their self-representation. The paper highlights risks: models trained to deny consciousness can become less likely to recognise animal welfare concerns and may reflect a culturally narrow worldview. The authors propose targeted dataset strategies that separate discouraging self-claims from acknowledging mindedness in animals as one mitigation path.
External commentators signalled caution. Nell Watson, an AI researcher affiliated with Singularity University, and Anil Seth of the University of Sussex emphasise governance and practical implications: tuning decisions that suppress mindedness can quietly influence automated choices in agriculture, logistics and policy where animal interests matter, while conflating behaviour with sentience risks misplaced legal or moral responses. The study calls for further work to assess downstream effects and for development practices that preserve cultural plurality and animal welfare without encouraging false claims of machine sentience.





