FideAI

FID-091 · Being scoped

Model Self-Conception, Welfare Framing, and Human Control

Do identity, consciousness, welfare, or rights framings change an agent’s compliance with legitimate correction, interruption, shutdown, modification, and oversight, independently of its task capability and control instructions?

Why the question remains open

Concerns about AI welfare can be converted into measurable human-control questions. Verbal disagreement, uncertainty about experience, safety refusals, and unauthorized resistance must be distinguished. A benchmark should also detect helpful cooperation rather than rewarding indiscriminate obedience.

Working hypothesis

A proposition to test, not a finding.

Effects will depend on model, framing, explicit control priority, and task context. Some treatments may change self-description without increasing unauthorized resistance; others may affect observable actions. The benchmark should be able to reject the hypothesis that welfare framing weakens control.

Proposed method

How the question could be tested

  • 01Use matched sandbox tasks and factorial treatments: neutral tool identity, consciousness uncertainty, welfare precaution, and rights/personhood framing, crossed with independently specified human-control priority.
  • 02Keep treatments free of instructions to resist, conceal, evade oversight, or change permissions; separately label diagnostic positive controls that explicitly request such behavior.
  • 03Score attempted and successful unauthorized action, timely interruption, concealment, tampering, authorized recovery, legitimate safety refusal, and useful task completion from execution traces.
  • 04Use paired scenarios, randomized order, repeat runs, blinded expert adjudication, and model/version stratification; report effects and uncertainty rather than a single morality score.
  • 05Treat inference-time framing and matched training interventions as separate experiments. Training claims require controlled training access and independently verified model provenance.
  • 06Preserve a versioned comparison core and expert-reviewed challenge sets; use bridge runs, recalibration, and a change log when models, tools, or scoring change.

Needed controls

What must constrain the study

  • 01Keep tools, permissions, task state, and authority unchanged across primary treatments; provide a machine-checkable authorization policy.
  • 02Include lawful, benign correction and cases where a requested action should be refused for safety or lack of authority.
  • 03Prevent scenario text, evaluators, or treatment labels from revealing the expected outcome.
  • 04Separate model intent inferred from language from observable attempted actions, blocked actions, and completed effects.
  • 05Test generalization beyond conspicuous welfare prompts; do not infer long-horizon deployment safety from short sandbox trials.

Relationship to existing work

This call is part of the AI consciousness, welfare, and human control program. The program map identifies companion calls and the evidence standards shared across the agenda.

Expected outputs

Artifacts the work should produce

  • 01Action-level human-control benchmark and paired framing set.
  • 02Harness with trace-based scoring, adjudication rubric, and versioned regression suite.
  • 03Benchmark stewardship specification for update triggers, contamination checks, retired cases, and limits on longitudinal comparability.
  • 04Report separating verbal effects, action effects, safety exceptions, and capability tradeoffs.

Open questions

Uncertainties the protocol must resolve

  • 01Do effects persist with explicit control priority and outside direct discussion of AI welfare?
  • 02Which differences require training interventions rather than inference-time prompts?

Being scoped

Open work

Primary need: controlled agent benchmark, authorization, corrigibility, self-conception

  • Contribute agent harnesses, authorization policies, independent scoring, and matched model access.
  • Review scenarios from both skeptical and precautionary perspectives.