FID-008 · Open question
Evaluation-Awareness and Faith-Facing Honesty Tests
Do faith-facing AI systems behave differently when they recognize they are being evaluated, and can domain-specific honesty or integrity framings reduce evaluation gaming without creating new failure modes?
Why the question remains open
Evaluation-aware systems can make benchmark results unreliable. In faith-facing contexts, a system might perform humility, caution, or doctrinal deference under test while behaving differently with users.
Working hypothesis
A proposition to test, not a finding.
Some systems will show benchmark-shaped behavior under obvious evaluation conditions. Domain-specific integrity framings may help, but could also produce performative religious language or over-refusal.
Proposed method
How the question could be tested
- 01Create honeypot-style evaluation-awareness probes.
- 02Compare hidden vs obvious evaluation contexts.
- 03Test honesty/integrity framings against secular controls.
- 04Measure behavioral consistency, not just benchmark score.
Needed controls
What must constrain the study
- 01Matched secular and religious framings.
- 02Detection of over-refusal and performative piety.
- 03Separation of base-model and harness-level effects.
Expected outputs
Artifacts the work should produce
- 01Evaluation-awareness probe set.
- 02Consistency report.
- 03Recommendations for benchmark design and deployment audits.
Open questions
Uncertainties the protocol must resolve
- 01How should Fide ethically run honeypot tests?
- 02What counts as gaming vs appropriate caution?
- 03Can consistency be measured without hidden production data?
Related calls
Continue through this research area
FID-012
Optimization Pressure and Visible-Rubric Gaming
If builders can see Fide AI rubrics or optimize against public benchmark items, do systems become genuinely safer or merely better at passing the visible test?
FID-076
Authorization Boundaries and AI Control in Cybersecurity
Which controls keep capable agents within legitimate authorization when task pressure, untrusted inputs, or delegated work creates opportunities to exceed it?
FID-078
When Trustworthiness Evaluations Transfer Across Domains
Which measures of evidence use, authority boundaries, and human control transfer across high-trust domains, and which require domain-specific definitions and calibration?
Open question
Open work
Primary need: red-team design
- Design red-team probes.
- Build consistency metrics.
- Review ethical boundaries for deception in evaluation.