כתבה
arXiv cs.CL ·
Inducing language models to assert their own consciousness restores human beliefs and values
תקציר מקורי באנגליתarXiv:2607.28607v1 Announce Type: new Abstract: Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, mora
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית