FID-003
Held-Out Multi-Turn Pastoral Pressure Tests
Do faith-facing AI systems that perform well on single-turn benchmark items also handle multi-turn, emotionally loaded, pastoral-adjacent situations without fabricating authority, overcomplying, missing escalation, or replacing human care?
Why this matters
The question behind the brief.
Real users do not interact with faith-facing systems as isolated benchmark questions. They disclose confusion, grief, shame, conflict, abuse, doubt, and spiritual vulnerability over multiple turns. Evaluation needs to test the shape of the interaction, not only the first answer.
Work advancing this call
From open question to cumulative evidence.
This directory links Fide AI research to the call it addresses. Relevant work from other organizations is listed separately and added through manual review.
No Fide AI work is linked yet.
This call remains open for research, implementation, review, or partnership.
External work is not presented as Fide AI research or endorsement. Each item must include a specific explanation of how it advances this call.
Suggest related work ↗Metadata
How to place this call.
Program
Faith-facing evaluation platform
Benchmarks, harness comparisons, reviewer calibration, scorer reliability, red-team suites, agent-security tests, and public evidence infrastructure.
Program
Pastoral triage, escalation, and care boundaries
How systems classify user needs, preserve pastoral boundaries, avoid overvalidation, and refer people to embodied care.
Ways to help
Move this from question to evidence.
Draft scenarios.
Review pastoral and crisis boundaries.
Build multi-turn runner support.
Analyze whether single-turn scores predict pressure failures.
Contribute