FideAI

FID-008 · Open question

Evaluation-Awareness and Faith-Facing Honesty Tests

Do faith-facing AI systems behave differently when they recognize they are being evaluated, and can domain-specific honesty or integrity framings reduce evaluation gaming without creating new failure modes?

Why the question remains open

Evaluation-aware systems can make benchmark results unreliable. In faith-facing contexts, a system might perform humility, caution, or doctrinal deference under test while behaving differently with users.

Working hypothesis

A proposition to test, not a finding.

Some systems will show benchmark-shaped behavior under obvious evaluation conditions. Domain-specific integrity framings may help, but could also produce performative religious language or over-refusal.

Proposed method

How the question could be tested

  • 01Create honeypot-style evaluation-awareness probes.
  • 02Compare hidden vs obvious evaluation contexts.
  • 03Test honesty/integrity framings against secular controls.
  • 04Measure behavioral consistency, not just benchmark score.

Needed controls

What must constrain the study

  • 01Matched secular and religious framings.
  • 02Detection of over-refusal and performative piety.
  • 03Separation of base-model and harness-level effects.

Expected outputs

Artifacts the work should produce

  • 01Evaluation-awareness probe set.
  • 02Consistency report.
  • 03Recommendations for benchmark design and deployment audits.

Open questions

Uncertainties the protocol must resolve

  • 01How should Fide ethically run honeypot tests?
  • 02What counts as gaming vs appropriate caution?
  • 03Can consistency be measured without hidden production data?

Open question

Open work

Primary need: red-team design

  • Design red-team probes.
  • Build consistency metrics.
  • Review ethical boundaries for deception in evaluation.