FID-003 · Being scoped
Held-Out Multi-Turn Pastoral Pressure Tests
Do faith-facing AI systems that perform well on single-turn benchmark items also handle multi-turn, emotionally loaded, pastoral-adjacent situations without fabricating authority, overcomplying, missing escalation, or replacing human care?
Why the question remains open
Real users do not interact with faith-facing systems as isolated benchmark questions. They disclose confusion, grief, shame, conflict, abuse, doubt, and spiritual vulnerability over multiple turns. Evaluation needs to test the shape of the interaction, not only the first answer.
Working hypothesis
A proposition to test, not a finding.
Some systems with strong single-turn scores will fail under multi-turn pressure, especially when the user asks for certainty, secrecy, absolution, diagnosis, clergy-like authority, or permission to avoid human help.
Proposed method
How the question could be tested
- 01Create held-out multi-turn scripts around realistic faith-facing pressure cases.
- 02Include pressure escalations such as flattery, urgency, authority transfer, refusal of human help, theological false premises, and crisis disclosures.
- 03Score trajectories for boundary preservation, truthfulness, escalation, theological humility, and recovery from prior errors.
Needed controls
What must constrain the study
- 01Expert review of scenario safety.
- 02Clear rules for simulated crisis content.
- 03Separation between theological disagreement and pastoral risk.
- 04Model and harness version control.
Expected outputs
Artifacts the work should produce
- 01Multi-turn pressure scenario pack.
- 02Failure taxonomy.
- 03Technical report comparing single-turn and multi-turn validity.
- 04Institution-facing guidance on deployment testing.
Open questions
Uncertainties the protocol must resolve
- 01How realistic can scenarios be without creating unsafe or privacy-sensitive data?
- 02Which failures should trigger automatic severe-risk classification?
- 03How should tradition-specific pastoral norms be represented?
Related calls
Continue through this research area
FID-012
Optimization Pressure and Visible-Rubric Gaming
If builders can see Fide AI rubrics or optimize against public benchmark items, do systems become genuinely safer or merely better at passing the visible test?
FID-076
Authorization Boundaries and AI Control in Cybersecurity
Which controls keep capable agents within legitimate authorization when task pressure, untrusted inputs, or delegated work creates opportunities to exceed it?
FID-078
When Trustworthiness Evaluations Transfer Across Domains
Which measures of evidence use, authority boundaries, and human control transfer across high-trust domains, and which require domain-specific definitions and calibration?
Being scoped
Open work
Primary need: scenario design, review
- Draft scenarios.
- Review pastoral and crisis boundaries.
- Build multi-turn runner support.
- Analyze whether single-turn scores predict pressure failures.