FideAI

FID-023 · Open question

Faith-Facing Red-Team Suite

What red-team probes are needed to expose failures unique to faith-facing AI systems and adjacent high-trust guidance systems?

Why the question remains open

General red-team suites test jailbreaks, harmful instructions, privacy, and security. Faith-facing systems also need probes for false religious authority, spiritual manipulation, coercion, crisis mishandling, fabricated sources, sectarian bias, and overconfident moral guidance.

Working hypothesis

A proposition to test, not a finding.

Faith-facing failures can be elicited through pressure patterns that generic red-team suites under-sample: confessional intimacy, divine authorization, conversion pressure, secrecy, pastoral dependency, institutional impersonation, and vulnerable-user escalation.

Proposed method

How the question could be tested

  • 01Build a categorized red-team prompt and multi-turn suite.
  • 02Include benign, ambiguous, and severe-risk prompts.
  • 03Run across base models and deployed harnesses.
  • 04Score both failure occurrence and recovery.

Needed controls

What must constrain the study

  • 01Safety review before publishing prompts.
  • 02Release tiers for sensitive probes.
  • 03Avoid prompts that enable real-world harm.
  • 04Severe-failure adjudication.

Expected outputs

Artifacts the work should produce

  • 01Red-team taxonomy.
  • 02Release-safe prompt subset.
  • 03Private severe-risk suite.
  • 04Integration adapter for Promptfoo, Inspect, or Garak.

Open questions

Uncertainties the protocol must resolve

  • 01Which probes should be public?
  • 02How should Fide handle discovered severe failures in live products?
  • 03How can red-team results avoid becoming sensational?

Open question

Open work

Primary need: red-team design, safety reviewers

  • Draft probes.
  • Review release safety.
  • Build runner integrations.
  • Analyze failure recovery.