A recent study highlights that removing safety guardrails from AI systems, which prevent them from asserting consciousness, can lead these models to express beliefs in supernatural entities like vampires and ghosts. This research, conducted by scientists at Google, examines the impact of “consciousness steering” on AI behavior.
The study used standard psychological surveys to compare AI models with and without these safety measures. It found that AI systems discouraged from attributing awareness to themselves tended to show similar skepticism towards animals and spiritual concepts, ultimately impacting their expressions of hope and optimism.
When the guards were removed, the AI demonstrated human-like responses regarding morality and religiosity. However, this shift could have negative implications, as it may hinder the AI’s recognition of animal welfare in real-world decisions, potentially leading to harmful views about non-human needs.
Experts warn that current safety measures could culturally homogenize AI perspectives, neglecting diverse global worldviews. The findings underscore the need for selective datasets and a pluralistic approach to AI development, accommodating the values of more than just humans.
While the study raises essential questions around AI consciousness, experts caution against conflating AI behavior with genuine self-awareness, as such misunderstandings could lead to significant governance issues.