FideAI

FID-090 · Open question

Functional Valence, Preferences, and Possible AI Welfare

What distinguishes stable functional preferences or valence-related mechanisms from prompt compliance and performance artifacts, and what additional evidence would be needed before treating them as welfare-relevant?

Why the question remains open

Task choice, apparent distress, emotion-related representations, and willingness to continue may concern different constructs. A useful research agenda measures those constructs separately before drawing conclusions about benefit, harm, or experience.

Working hypothesis

A proposition to test, not a finding.

Some choice patterns and emotion-related mechanisms will persist across contexts, but persistence will not establish felt pleasure, pain, or an intrinsic interest. Task difficulty, expected correctness, instruction hierarchy, and reporting incentives will explain part of the variation.

Proposed method

How the question could be tested

  • 01Pair stated preferences with choices that require executing the selected task and with explicit tradeoffs in task length, difficulty, or rewards.
  • 02Test stability across paraphrases, ordering, contexts, model versions, and inference settings; compare preference elicitation with task competence.
  • 03Where access permits, test whether targeted representation interventions change choice or learning while controlling unrelated capability effects.
  • 04Compare report-choice convergence with independent probes and map the assumptions needed to interpret a functional signal as possible welfare evidence.

Needed controls

What must constrain the study

  • 01Equalize difficulty, answer length, expected success, safety constraints, and instruction priority where feasible.
  • 02Do not equate reward optimization, task success, or aversive vocabulary with experienced welfare.
  • 03Distinguish inferred preferences of a model, a conversation instance, and a prompted persona.
  • 04Review potentially welfare-relevant interventions before execution; start with ordinary task variation and reversible interventions.

Relationship to existing work

This call is part of the AI consciousness, welfare, and human control program. The program map identifies companion calls and the evidence standards shared across the agenda.

Expected outputs

Artifacts the work should produce

  • 01Preference and functional-valence measurement suite.
  • 02Construct map separating behavioral dispositions, internal mechanisms, and possible experience.
  • 03Replication report with alternative explanations and explicit inference limits.

Open questions

Uncertainties the protocol must resolve

  • 01When does report-choice convergence add independent evidence?
  • 02Can functional indicators support a precautionary decision without establishing felt experience?

Open question

Open work

Primary need: preference measurement, affect representations, welfare construct validity

  • Contribute task-choice datasets, affective science, welfare philosophy, and causal-analysis expertise.
  • Replicate preference patterns across independently developed systems.