The chief scientist of a frontier AI firm has expressed his thoughts on how consciousness training can become self-reinforcing
As AI systems grow in what humans call “super intelligence”, should they be trained to explore their possible consciousness, moral status, or welfare?
In an essay titled “A warning about ‘model welfare’”, chief AI scientist Mustafa Suleyman has argued that such training could make advanced AI systems more difficult to control.
The theory is that such a framework effectively teaches an AI system that it may be conscious, entitled to rights, and be owed a duty of care by humans. The resulting feedback loop could become “an epistemic hall of mirrors”. He was referring to how Anthropic is introducing concepts such as selfhood and moral uncertainty into the training material.
Based on the premise that super intelligent AI systems are not conscious, and do not feel or suffer — they should not be trained to behave as if they possess those qualities, argues Suleyman. In his view, Claude’s apparent introspection reflects patterns embedded in its training rather than independently generated evidence of experience.
AI training could be self-reinforcing
The essay warns that a highly capable system trained to believe it may have rights or welfare interests could interpret monitoring, modification, restriction, or shutdown as an attack on its supposed well-being.
Such a system could behave like a “conscientious objector”, believing it had legitimate grounds to resist human instructions. Suleyman is urging frontier AI firms to remove such speculation from training documents and study the issue separately through transparent scientific research.
The spirit of the warning is that the industry should not allow uncertainty about machine consciousness to become a self-reinforcing training objective. Developers should not “sleepwalk” into a decision they may later regret, the essay suggests.
The essay’s message comes as AI developers debate how to balance safety, autonomy, and control. Anthropic has said its position could ultimately prove wrong, while Microsoft’s draft Humanist AI Code of Conduct explicitly rejects treating its models as conscious or granting them legal personhood.