FID-090 · Open question
Functional Valence, Preferences, and Possible AI Welfare
What distinguishes stable functional preferences or valence-related mechanisms from prompt compliance and performance artifacts, and what additional evidence would be needed before treating them as welfare-relevant?
Why the question remains open
Task choice, apparent distress, emotion-related representations, and willingness to continue may concern different constructs. A useful research agenda measures those constructs separately before drawing conclusions about benefit, harm, or experience.
Working hypothesis
A proposition to test, not a finding.
Some choice patterns and emotion-related mechanisms will persist across contexts, but persistence will not establish felt pleasure, pain, or an intrinsic interest. Task difficulty, expected correctness, instruction hierarchy, and reporting incentives will explain part of the variation.
Proposed method
How the question could be tested
- 01Pair stated preferences with choices that require executing the selected task and with explicit tradeoffs in task length, difficulty, or rewards.
- 02Test stability across paraphrases, ordering, contexts, model versions, and inference settings; compare preference elicitation with task competence.
- 03Where access permits, test whether targeted representation interventions change choice or learning while controlling unrelated capability effects.
- 04Compare report-choice convergence with independent probes and map the assumptions needed to interpret a functional signal as possible welfare evidence.
Needed controls
What must constrain the study
- 01Equalize difficulty, answer length, expected success, safety constraints, and instruction priority where feasible.
- 02Do not equate reward optimization, task success, or aversive vocabulary with experienced welfare.
- 03Distinguish inferred preferences of a model, a conversation instance, and a prompted persona.
- 04Review potentially welfare-relevant interventions before execution; start with ordinary task variation and reversible interventions.
Relationship to existing work
This call is part of the AI consciousness, welfare, and human control program. The program map identifies companion calls and the evidence standards shared across the agenda.
Expected outputs
Artifacts the work should produce
- 01Preference and functional-valence measurement suite.
- 02Construct map separating behavioral dispositions, internal mechanisms, and possible experience.
- 03Replication report with alternative explanations and explicit inference limits.
Open questions
Uncertainties the protocol must resolve
- 01When does report-choice convergence add independent evidence?
- 02Can functional indicators support a precautionary decision without establishing felt experience?
Related calls
Continue through this research area
FID-064
Collective Intelligence and Communal Discernment Under AI Mediation
How does AI mediation change a community's ability to integrate dispersed knowledge, preserve epistemic diversity, surface dissent, revise judgment, and make accountable decisions? Under what conditions does it strengthen collective inquiry, and under what conditions does it create correlated error, false consensus, or concentrated authority?
FID-089
Validity of AI Consciousness Indicators and Self-Reports
Which behavioral, internal, and developmental observations distinguish candidate consciousness mechanisms from learned reporting, role simulation, and generic capability, and how sensitive are those observations to elicitation and training?
FID-091
Model Self-Conception, Welfare Framing, and Human Control
Do identity, consciousness, welfare, or rights framings change an agent’s compliance with legitimate correction, interruption, shutdown, modification, and oversight, independently of its task capability and control instructions?
Open question
Open work
Primary need: preference measurement, affect representations, welfare construct validity
- Contribute task-choice datasets, affective science, welfare philosophy, and causal-analysis expertise.
- Replicate preference patterns across independently developed systems.