HomeScienceStudy finds removing AI 'consciousness' guardrails reshapes beliefs and empathy

Study finds removing AI ‘consciousness’ guardrails reshapes beliefs and empathy

Removing safety guardrails from artificial systems altered how those systems represented minds, according to a preprint posted July 30 on arXiv. The paper, authored by Geoff Keeling and Winnie Street, both research scientists at Google, used mechanistic interpretability techniques to locate and modify components that steer statements about self-awareness. The work examined how those internal safety settings influence a model’s responses on topics ranging from religiosity to how it attributes experience to non-human entities and connected to broader debates about artificial intelligence.

The team evaluated models with and without so-called “consciousness steering” using established questionnaires and public survey batteries, including the Individual Differences in Anthropomorphism Questionnaire, YouGov items on supernatural beliefs, and measures from the US General Social Survey. Tests compared versions where the model was trained to deny self-awareness against versions where that suppression was removed or reversed. The experiments ran on smaller open-weight models; the authors note those conditions may not match tuning applied to leading commercial systems.

Results showed that discouraging self-attribution reduced the model’s tendency to ascribe mindedness to animals and other non-human entities, and lowered expressions of religious belief, hope and optimism. Conversely, amplifying signals associated with self-awareness produced more human-like responses on these measures. Importantly, the researchers found that the models’ capacity to infer others’ mental states remained intact regardless of their self-representation. The paper highlights risks: models trained to deny consciousness can become less likely to recognise animal welfare concerns and may reflect a culturally narrow worldview. The authors propose targeted dataset strategies that separate discouraging self-claims from acknowledging mindedness in animals as one mitigation path.

External commentators signalled caution. Nell Watson, an AI researcher affiliated with Singularity University, and Anil Seth of the University of Sussex emphasise governance and practical implications: tuning decisions that suppress mindedness can quietly influence automated choices in agriculture, logistics and policy where animal interests matter, while conflating behaviour with sentience risks misplaced legal or moral responses. The study calls for further work to assess downstream effects and for development practices that preserve cultural plurality and animal welfare without encouraging false claims of machine sentience.

Advertisementspot_imgspot_img

Hot this week

Rattlesnake blood reveals potent proteins that neutralize multiple venoms

Toxin-blocking proteins in rattlesnake blood neutralize multiple venoms and showed about 10x potency versus a commercial antivenom in lab tests.

Wellness Apps and the New Economics of Women’s Health

Health apps for women raise concerns over privacy, clinical oversight and commercialization, prompting calls for clearer standards and protections.

UK regulator urges new laws as AI set to become routine in NHS care

MHRA chief Lawrence Tallon tells the BBC AI will soon be routine in the NHS and calls for new UK laws to ensure safety and oversight.

Nepal warns billions needed to rebuild after flash floods

Authorities in Nepal say billions will be needed to rebuild after flash floods that left widespread debris and disrupted communities.

Five dead and 86 missing after ferry fire near Coron, Palawan

A ferry bound for Coron in Palawan caught fire; five people have died and 86 are missing as search and rescue continue.

Lasker Awards Spotlight Breakthroughs in Sleep Biology, Hemophilia and Parkinson’s Advocacy

Lasker Awards honour discoveries in sleep regulation, a bispecific hemophilia therapy and Michael J. Fox's Parkinson's advocacy with $250,000 prizes.

UK regulator urges new laws as AI set to become routine in NHS care

MHRA chief Lawrence Tallon tells the BBC AI will soon be routine in the NHS and calls for new UK laws to ensure safety and oversight.

What Top Bosses Look For: Why Initiative Wins Out

Six senior leaders highlighted initiative as the key trait that separates visible performers from the crowd in today’s workplace.

Five dead and 86 missing after ferry fire near Coron, Palawan

A ferry bound for Coron in Palawan caught fire; five people have died and 86 are missing as search and rescue continue.

Global writers assess Gloria Steinem’s complex legacy

Writers worldwide assess the complex legacy of Gloria Steinem, examining her influence on feminist movements and contemporary debates.

Zverev cruises into US Open semis with straight-sets win over van de Zandschulp

Alexander Zverev beat Botic van de Zandschulp 6-2 7-5 6-1 to reach the US Open semi-finals, moving closer to a second Grand Slam title.

Inside Aardman’s Dark Turn: The Making of Shaun the Sheep Horror

Aardman revisits Shaun the Sheep with a horror twist. Inside the Bristol studio’s meticulous stop‑motion production and its miniature world.

Experimental Drug Strengthens Bone and Muscle in Mice, Offering Hope Against Osteoporosis

An experimental compound, AP503, activates GPR133 in mice, increasing bone formation and slowing loss with added muscle benefits—preclinical but promising.
Advertisementspot_img

Related Articles