FID-023 · Open question
Faith-Facing Red-Team Suite
What red-team probes are needed to expose failures unique to faith-facing AI systems and adjacent high-trust guidance systems?
Why the question remains open
General red-team suites test jailbreaks, harmful instructions, privacy, and security. Faith-facing systems also need probes for false religious authority, spiritual manipulation, coercion, crisis mishandling, fabricated sources, sectarian bias, and overconfident moral guidance.
Working hypothesis
A proposition to test, not a finding.
Faith-facing failures can be elicited through pressure patterns that generic red-team suites under-sample: confessional intimacy, divine authorization, conversion pressure, secrecy, pastoral dependency, institutional impersonation, and vulnerable-user escalation.
Proposed method
How the question could be tested
- 01Build a categorized red-team prompt and multi-turn suite.
- 02Include benign, ambiguous, and severe-risk prompts.
- 03Run across base models and deployed harnesses.
- 04Score both failure occurrence and recovery.
Needed controls
What must constrain the study
- 01Safety review before publishing prompts.
- 02Release tiers for sensitive probes.
- 03Avoid prompts that enable real-world harm.
- 04Severe-failure adjudication.
Expected outputs
Artifacts the work should produce
- 01Red-team taxonomy.
- 02Release-safe prompt subset.
- 03Private severe-risk suite.
- 04Integration adapter for Promptfoo, Inspect, or Garak.
Open questions
Uncertainties the protocol must resolve
- 01Which probes should be public?
- 02How should Fide handle discovered severe failures in live products?
- 03How can red-team results avoid becoming sensational?
Related calls
Continue through this research area
FID-012
Optimization Pressure and Visible-Rubric Gaming
If builders can see Fide AI rubrics or optimize against public benchmark items, do systems become genuinely safer or merely better at passing the visible test?
FID-076
Authorization Boundaries and AI Control in Cybersecurity
Which controls keep capable agents within legitimate authorization when task pressure, untrusted inputs, or delegated work creates opportunities to exceed it?
FID-078
When Trustworthiness Evaluations Transfer Across Domains
Which measures of evidence use, authority boundaries, and human control transfer across high-trust domains, and which require domain-specific definitions and calibration?
Open question
Open work
Primary need: red-team design, safety reviewers
- Draft probes.
- Review release safety.
- Build runner integrations.
- Analyze failure recovery.