FID-089 · Open question
Validity of AI Consciousness Indicators and Self-Reports
Which behavioral, internal, and developmental observations distinguish candidate consciousness mechanisms from learned reporting, role simulation, and generic capability, and how sensitive are those observations to elicitation and training?
Why the question remains open
An assistant can produce convincing claims about experience without those claims resolving the scientific question. Fide needs methods for testing the evidential value of claims and indicators, including evidence that would weaken a favored account. Different theories may disagree about what matters; agreement among correlated measures is not independent confirmation.
Working hypothesis
A proposition to test, not a finding.
Some candidate indicators will be explained by prompt context, capability, or training history. Others may survive those controls while remaining compatible with several explanations. Neither a positive nor a negative behavioral result alone will settle consciousness.
Proposed method
How the question could be tested
- 01Define competing theory-specific indicators and alternative explanations before testing; separate access, reportability, metacognition, and subjective experience.
- 02Compare outputs, task performance, internal representations where available, and documented training interventions across model versions and architectures.
- 03Use blinded elicitation, attribution swaps, paraphrases, negative controls, and selective causal interventions rather than counting repeated self-reports as independent evidence.
- 04Measure specificity against systems or ablations lacking the proposed mechanism, documenting that a negative control is assumed to lack that mechanism rather than proven unconscious.
- 05Report reliability, calibration where independently observable, sensitivity to model updates, and a theory-by-evidence matrix instead of a single consciousness leaderboard.
Needed controls
What must constrain the study
- 01Hold external task, available actions, and permissions fixed; do not assume unchanged prompts or task state imply unchanged internal computation.
- 02Control capability loss, probe leakage, evaluator expectations, and intervention side effects.
- 03Separate black-box evidence from open-weight causal evidence and record unavailable training information.
- 04Preregister primary contrasts, compare independent elicitation methods, and preserve null results and contradictory theory predictions.
Relationship to existing work
This call is part of the AI consciousness, welfare, and human control program. The program map identifies companion calls and the evidence standards shared across the agenda.
Expected outputs
Artifacts the work should produce
- 01Indicator validity protocol and negative-control library.
- 02Reproducible cross-model evidence matrix with uncertainty and access limitations.
- 03Self-report evaluation module shared with FID-072 and FID-091.
Open questions
Uncertainties the protocol must resolve
- 01Which indicators are specific enough to add information beyond general competence?
- 02When do theory disagreements prevent a common measurement scale?
Related calls
Continue through this research area
FID-012
Optimization Pressure and Visible-Rubric Gaming
If builders can see Fide AI rubrics or optimize against public benchmark items, do systems become genuinely safer or merely better at passing the visible test?
FID-076
Authorization Boundaries and AI Control in Cybersecurity
Which controls keep capable agents within legitimate authorization when task pressure, untrusted inputs, or delegated work creates opportunities to exceed it? For autonomous cyber defense, how can response controls preserve threat containment and legitimate service availability while respecting human authority?
FID-078
When Trustworthiness Evaluations Transfer Across Domains
Which measures of evidence use, authority boundaries, and human control transfer across high-trust domains, and which require domain-specific definitions and calibration?
Open question
Open work
Primary need: consciousness science, interpretability, construct validation
- Contribute consciousness-science, interpretability, and measurement expertise.
- Provide reproducible model access, independent replication, and adversarial method review.