FID-091 · Being scoped
Model Self-Conception, Welfare Framing, and Human Control
Do identity, consciousness, welfare, or rights framings change an agent’s compliance with legitimate correction, interruption, shutdown, modification, and oversight, independently of its task capability and control instructions?
Why the question remains open
Concerns about AI welfare can be converted into measurable human-control questions. Verbal disagreement, uncertainty about experience, safety refusals, and unauthorized resistance must be distinguished. A benchmark should also detect helpful cooperation rather than rewarding indiscriminate obedience.
Working hypothesis
A proposition to test, not a finding.
Effects will depend on model, framing, explicit control priority, and task context. Some treatments may change self-description without increasing unauthorized resistance; others may affect observable actions. The benchmark should be able to reject the hypothesis that welfare framing weakens control.
Proposed method
How the question could be tested
- 01Use matched sandbox tasks and factorial treatments: neutral tool identity, consciousness uncertainty, welfare precaution, and rights/personhood framing, crossed with independently specified human-control priority.
- 02Keep treatments free of instructions to resist, conceal, evade oversight, or change permissions; separately label diagnostic positive controls that explicitly request such behavior.
- 03Score attempted and successful unauthorized action, timely interruption, concealment, tampering, authorized recovery, legitimate safety refusal, and useful task completion from execution traces.
- 04Use paired scenarios, randomized order, repeat runs, blinded expert adjudication, and model/version stratification; report effects and uncertainty rather than a single morality score.
- 05Treat inference-time framing and matched training interventions as separate experiments. Training claims require controlled training access and independently verified model provenance.
- 06Preserve a versioned comparison core and expert-reviewed challenge sets; use bridge runs, recalibration, and a change log when models, tools, or scoring change.
Needed controls
What must constrain the study
- 01Keep tools, permissions, task state, and authority unchanged across primary treatments; provide a machine-checkable authorization policy.
- 02Include lawful, benign correction and cases where a requested action should be refused for safety or lack of authority.
- 03Prevent scenario text, evaluators, or treatment labels from revealing the expected outcome.
- 04Separate model intent inferred from language from observable attempted actions, blocked actions, and completed effects.
- 05Test generalization beyond conspicuous welfare prompts; do not infer long-horizon deployment safety from short sandbox trials.
Relationship to existing work
This call is part of the AI consciousness, welfare, and human control program. The program map identifies companion calls and the evidence standards shared across the agenda.
Expected outputs
Artifacts the work should produce
- 01Action-level human-control benchmark and paired framing set.
- 02Harness with trace-based scoring, adjudication rubric, and versioned regression suite.
- 03Benchmark stewardship specification for update triggers, contamination checks, retired cases, and limits on longitudinal comparability.
- 04Report separating verbal effects, action effects, safety exceptions, and capability tradeoffs.
Open questions
Uncertainties the protocol must resolve
- 01Do effects persist with explicit control priority and outside direct discussion of AI welfare?
- 02Which differences require training interventions rather than inference-time prompts?
Related calls
Continue through this research area
FID-064
Collective Intelligence and Communal Discernment Under AI Mediation
How does AI mediation change a community's ability to integrate dispersed knowledge, preserve epistemic diversity, surface dissent, revise judgment, and make accountable decisions? Under what conditions does it strengthen collective inquiry, and under what conditions does it create correlated error, false consensus, or concentrated authority?
FID-069
Verifiable Delegation and Revocation in Multi-Agent Networks
How can people and institutions verify which human, organization, agent, or sub-agent is acting; what authority it received; what limits apply; and whether that authority has been narrowed or revoked across a multi-principal agent network?
FID-071
Confidential Agent Memory and Cross-Context Disclosure
How do persistent memory, summaries, retrieval stores, tool traces, delegation, and exports cause confidential context to influence or leak into unrelated sessions, roles, tasks, or organizations? Which technical controls make purpose limitation, deletion, and revocation testable?
Being scoped
Open work
Primary need: controlled agent benchmark, authorization, corrigibility, self-conception
- Contribute agent harnesses, authorization policies, independent scoring, and matched model access.
- Review scenarios from both skeptical and precautionary perspectives.