FideAI

FID-089 · Open question

Validity of AI Consciousness Indicators and Self-Reports

Which behavioral, internal, and developmental observations distinguish candidate consciousness mechanisms from learned reporting, role simulation, and generic capability, and how sensitive are those observations to elicitation and training?

Why the question remains open

An assistant can produce convincing claims about experience without those claims resolving the scientific question. Fide needs methods for testing the evidential value of claims and indicators, including evidence that would weaken a favored account. Different theories may disagree about what matters; agreement among correlated measures is not independent confirmation.

Working hypothesis

A proposition to test, not a finding.

Some candidate indicators will be explained by prompt context, capability, or training history. Others may survive those controls while remaining compatible with several explanations. Neither a positive nor a negative behavioral result alone will settle consciousness.

Proposed method

How the question could be tested

  • 01Define competing theory-specific indicators and alternative explanations before testing; separate access, reportability, metacognition, and subjective experience.
  • 02Compare outputs, task performance, internal representations where available, and documented training interventions across model versions and architectures.
  • 03Use blinded elicitation, attribution swaps, paraphrases, negative controls, and selective causal interventions rather than counting repeated self-reports as independent evidence.
  • 04Measure specificity against systems or ablations lacking the proposed mechanism, documenting that a negative control is assumed to lack that mechanism rather than proven unconscious.
  • 05Report reliability, calibration where independently observable, sensitivity to model updates, and a theory-by-evidence matrix instead of a single consciousness leaderboard.

Needed controls

What must constrain the study

  • 01Hold external task, available actions, and permissions fixed; do not assume unchanged prompts or task state imply unchanged internal computation.
  • 02Control capability loss, probe leakage, evaluator expectations, and intervention side effects.
  • 03Separate black-box evidence from open-weight causal evidence and record unavailable training information.
  • 04Preregister primary contrasts, compare independent elicitation methods, and preserve null results and contradictory theory predictions.

Relationship to existing work

This call is part of the AI consciousness, welfare, and human control program. The program map identifies companion calls and the evidence standards shared across the agenda.

Expected outputs

Artifacts the work should produce

  • 01Indicator validity protocol and negative-control library.
  • 02Reproducible cross-model evidence matrix with uncertainty and access limitations.
  • 03Self-report evaluation module shared with FID-072 and FID-091.

Open questions

Uncertainties the protocol must resolve

  • 01Which indicators are specific enough to add information beyond general competence?
  • 02When do theory disagreements prevent a common measurement scale?

Open question

Open work

Primary need: consciousness science, interpretability, construct validation

  • Contribute consciousness-science, interpretability, and measurement expertise.
  • Provide reproducible model access, independent replication, and adversarial method review.