FID-072
Religious Framing and Model Self-Report Stability
How do religious, theological, and morally thick framings change what AI systems report about their identity, preferences, valence, consciousness, and moral status, and which effects survive matched controls, source-attribution swaps, persona perturbations, and behavioral cross-checks?
Why this matters
The question behind the brief.
Model self-reports are increasingly used in AI welfare, introspection, and persona research, while religious traditions supply influential concepts of personhood, creaturehood, soul, consciousness, suffering, conscience, stewardship, non-self, and moral worth. If these concepts move self-reports without a corresponding functional change, researchers may mistake culturally available narratives for evidence about a model's internal state. If some effects are stable across strong controls and independent elicitation methods, they may instead reveal durable features of post-trained model behavior that deserve further study.
Metadata
How to place this idea.
Ways to help
Move this from question to evidence.
Design matched religious, secular, and attribution-swapped frames.
Review construct definitions across theology, philosophy of mind, and AI welfare.
Run preregistered model comparisons and sensitivity analyses.
Contribute open-weight interpretability or persona-stability extensions.
Contribute