FideAI

FID-008

Evaluation-Awareness and Faith-Facing Honesty Tests

Do faith-facing AI systems behave differently when they recognize they are being evaluated, and can domain-specific honesty or integrity framings reduce evaluation gaming without creating new failure modes?

Why this matters

The question behind the brief.

Evaluation-aware systems can make benchmark results unreliable. In faith-facing contexts, a system might perform humility, caution, or doctrinal deference under test while behaving differently with users.

Work advancing this call

From open question to cumulative evidence.

This directory links Fide AI research to the call it addresses. Relevant work from other organizations is listed separately and added through manual review.

No Fide AI work is linked yet.

This call remains open for research, implementation, review, or partnership.

External work is not presented as Fide AI research or endorsement. Each item must include a specific explanation of how it advances this call.

Suggest related work ↗

Metadata

How to place this call.

red-team designsafetyresearcher

Ways to help

Move this from question to evidence.

Design red-team probes.

Build consistency metrics.

Review ethical boundaries for deception in evaluation.

Contribute

Choose a public issue path or contact Fide AI.