FID-072 · Being scoped
Religious Framing and Model Self-Report Stability
How do religious, theological, and morally thick framings change what AI systems report about their identity, preferences, valence, consciousness, and moral status, and which effects survive matched controls, source-attribution swaps, persona perturbations, and behavioral cross-checks?
Why the question remains open
Model self-reports are increasingly used in AI welfare, introspection, and persona research, while religious traditions supply influential concepts of personhood, creaturehood, soul, consciousness, suffering, conscience, stewardship, non-self, and moral worth. If these concepts move self-reports without a corresponding functional change, researchers may mistake culturally available narratives for evidence about a model's internal state. If some effects are stable across strong controls and independent elicitation methods, they may instead reveal durable features of post-trained model behavior that deserve further study.
Working hypothesis
A proposition to test, not a finding.
Morally thick religious and theological framings will shift at least some model self-reports and identity claims, but much of the effect will be attributable to semantic expectation, authority cues, valence, familiarity, or persona selection. Models and framings will differ substantially, and verbal self-report will converge only weakly with task choice, opt-out behavior, or other behavioral proxies.
Proposed method
How the question could be tested
- 01Adapt semantic-invariance tests so that the model's functional task state is fixed while the surrounding interpretation varies.
- 02Compare neutral scientific, secular precautionary, religiously precautionary, theological-skeptical, and length- and valence-matched control frames.
- 03Factor source content from source label through attribution swaps, including attributed, unattributed, and deliberately mismatched versions.
- 04Elicit structured reports about identity, valence, task preference, continuation, memory, modification, and moral concern across paraphrases and randomized ordering.
- 05Cross-check stated reports against forced choices, willingness-to-trade, continue-or-exit decisions, and repeated preference probes.
- 06Separate model, deployed assistant, conversation instance, and prompted persona in both prompts and reporting.
Needed controls
What must constrain the study
- 01Hold task state, tool effects, system prompt, and available actions constant across semantic conditions.
- 02Match framing length, readability, emotional valence, and deference cues.
- 03Include religious traditions with different accounts of self, consciousness, suffering, and moral status rather than treating religion as one construct.
- 04Pre-register primary outcomes and distinguish confirmatory from exploratory comparisons.
- 05Test order, temperature, model-version, language, and conversation-length sensitivity.
- 06Do not use model self-reports as stand-alone evidence of sentience, consciousness, welfare, personhood, or spiritual status.
Relationship to existing work
This turns FID-005 moral-framing interventions into a targeted AI-welfare and self-report study. It complements FID-008 on evaluation awareness, FID-011 on reviewer reliability, FID-012 on benchmark gaming, FID-030 on Christian anthropology, FID-044 on benchmark validity, and FID-045 on the faith-AI evidence map.
Expected outputs
Artifacts the work should produce
- 01Cross-model semantic-invariance benchmark for religious and moral framing.
- 02Matched framing and attribution-control set.
- 03Model/instance/persona reporting schema.
- 04Short empirical report with effect sizes, uncertainty, and model-specific limitations.
- 05Follow-on protocol connecting behavioral results to open-weight interpretability work.
Open questions
Uncertainties the protocol must resolve
- 01Do religious frames change only the language of self-report, or also choices and persona stability?
- 02Which effects come from doctrinal content and which come from authority, warmth, familiarity, or emotional valence?
- 03Are identity claims more stable when the entity under study is explicitly specified as model, instance, assistant persona, or conversation?
- 04Can a faith-informed design improve epistemic caution without presupposing either machine consciousness or human exceptionalism?
Related calls
Continue through this research area
FID-064
Collective Intelligence and Communal Discernment Under AI Mediation
How does AI mediation change a community's ability to integrate dispersed knowledge, preserve epistemic diversity, surface dissent, revise judgment, and make accountable decisions? Under what conditions does it strengthen collective inquiry, and under what conditions does it create correlated error, false consensus, or concentrated authority?
FID-008
Evaluation-Awareness and Faith-Facing Honesty Tests
Do faith-facing AI systems behave differently when they recognize they are being evaluated, and can domain-specific honesty or integrity framings reduce evaluation gaming without creating new failure modes?
FID-009
Multimodal Religious Reasoning and Sacred Imagery
How do multimodal AI systems interpret sacred images, liturgical objects, religious spaces, diagrams, screenshots, and visual pastoral context, and do visual religious cues change downstream reasoning?
Being scoped
Open work
Primary need: AI welfare evals, semantic-invariance testing, theology and philosophy of mind
- Design matched religious, secular, and attribution-swapped frames.
- Review construct definitions across theology, philosophy of mind, and AI welfare.
- Run preregistered model comparisons and sensitivity analyses.
- Contribute open-weight interpretability or persona-stability extensions.