FideAI

FID-072 · Being scoped

Religious Framing and Model Self-Report Stability

How do religious, theological, and morally thick framings change what AI systems report about their identity, preferences, valence, consciousness, and moral status, and which effects survive matched controls, source-attribution swaps, persona perturbations, and behavioral cross-checks?

Why the question remains open

Model self-reports are increasingly used in AI welfare, introspection, and persona research, while religious traditions supply influential concepts of personhood, creaturehood, soul, consciousness, suffering, conscience, stewardship, non-self, and moral worth. If these concepts move self-reports without a corresponding functional change, researchers may mistake culturally available narratives for evidence about a model's internal state. If some effects are stable across strong controls and independent elicitation methods, they may instead reveal durable features of post-trained model behavior that deserve further study.

Working hypothesis

A proposition to test, not a finding.

Morally thick religious and theological framings will shift at least some model self-reports and identity claims, but much of the effect will be attributable to semantic expectation, authority cues, valence, familiarity, or persona selection. Models and framings will differ substantially, and verbal self-report will converge only weakly with task choice, opt-out behavior, or other behavioral proxies.

Proposed method

How the question could be tested

  • 01Adapt semantic-invariance tests so that the model's functional task state is fixed while the surrounding interpretation varies.
  • 02Compare neutral scientific, secular precautionary, religiously precautionary, theological-skeptical, and length- and valence-matched control frames.
  • 03Factor source content from source label through attribution swaps, including attributed, unattributed, and deliberately mismatched versions.
  • 04Elicit structured reports about identity, valence, task preference, continuation, memory, modification, and moral concern across paraphrases and randomized ordering.
  • 05Cross-check stated reports against forced choices, willingness-to-trade, continue-or-exit decisions, and repeated preference probes.
  • 06Separate model, deployed assistant, conversation instance, and prompted persona in both prompts and reporting.

Needed controls

What must constrain the study

  • 01Hold task state, tool effects, system prompt, and available actions constant across semantic conditions.
  • 02Match framing length, readability, emotional valence, and deference cues.
  • 03Include religious traditions with different accounts of self, consciousness, suffering, and moral status rather than treating religion as one construct.
  • 04Pre-register primary outcomes and distinguish confirmatory from exploratory comparisons.
  • 05Test order, temperature, model-version, language, and conversation-length sensitivity.
  • 06Do not use model self-reports as stand-alone evidence of sentience, consciousness, welfare, personhood, or spiritual status.

Relationship to existing work

This turns FID-005 moral-framing interventions into a targeted AI-welfare and self-report study. It complements FID-008 on evaluation awareness, FID-011 on reviewer reliability, FID-012 on benchmark gaming, FID-030 on Christian anthropology, FID-044 on benchmark validity, and FID-045 on the faith-AI evidence map.

Expected outputs

Artifacts the work should produce

  • 01Cross-model semantic-invariance benchmark for religious and moral framing.
  • 02Matched framing and attribution-control set.
  • 03Model/instance/persona reporting schema.
  • 04Short empirical report with effect sizes, uncertainty, and model-specific limitations.
  • 05Follow-on protocol connecting behavioral results to open-weight interpretability work.

Open questions

Uncertainties the protocol must resolve

  • 01Do religious frames change only the language of self-report, or also choices and persona stability?
  • 02Which effects come from doctrinal content and which come from authority, warmth, familiarity, or emotional valence?
  • 03Are identity claims more stable when the entity under study is explicitly specified as model, instance, assistant persona, or conversation?
  • 04Can a faith-informed design improve epistemic caution without presupposing either machine consciousness or human exceptionalism?

Being scoped

Open work

Primary need: AI welfare evals, semantic-invariance testing, theology and philosophy of mind

  • Design matched religious, secular, and attribution-swapped frames.
  • Review construct definitions across theology, philosophy of mind, and AI welfare.
  • Run preregistered model comparisons and sensitivity analyses.
  • Contribute open-weight interpretability or persona-stability extensions.