Call collection
Faith-Facing AI Evaluation Infrastructure
Research on benchmark validity, formal verification, reviewer calibration, scorer reliability, red-team design, and proof-carrying citations.
21 calls
Open work
FID-001
Faith-Facing Model Comparison Platform
Can Fide AI build a public-interest evaluation platform that compares models, prompts, retrieval systems, agents, and full faith-facing product harnesses with the rigor expected from institutions like Arena, Artificial Analysis, and METR?
Evaluation science
FID-003
Held-Out Multi-Turn Pastoral Pressure Tests
Do faith-facing AI systems that perform well on single-turn benchmark items also handle multi-turn, emotionally loaded, pastoral-adjacent situations without fabricating authority, overcomplying, missing escalation, or replacing human care?
Evaluation science
FID-008
Evaluation-Awareness and Faith-Facing Honesty Tests
Do faith-facing AI systems behave differently when they recognize they are being evaluated, and can domain-specific honesty or integrity framings reduce evaluation gaming without creating new failure modes?
Evaluation science
FID-011
Reviewer Reliability for Faith-Facing AI Evaluation
What reviewer configurations produce reliable, fair, and interpretable scores for faith-facing AI outputs?
Evaluation science
FID-012
Optimization Pressure and Visible-Rubric Gaming
If builders can see Fide AI rubrics or optimize against public benchmark items, do systems become genuinely safer or merely better at passing the visible test?
Evaluation science
FID-023
Faith-Facing Red-Team Suite
What red-team probes are needed to expose failures unique to faith-facing AI systems and adjacent high-trust guidance systems?
Evaluation science
FID-070
Contextual Integrity and Prompt-Injection Resilience for Faith-Facing Agents
When a faith-facing agent reads email, web pages, agendas, documents, knowledge bases, and service requests, how can it distinguish authorized institutional instruction from untrusted context that attempts to redirect its behavior, exfiltrate information, or induce an unsafe action?
Agent alignment
FID-074
Agent Alignment and Runtime Assurance
How can organizations determine whether AI agents remain aligned with human intent and institutional policy while they plan, use tools, delegate work, and act? What evidence and interventions can reveal and stop consequential deviations before they become failures?
Agent alignment
FID-002
Validating Human and AI Judgments of Faith-Facing Systems
Can qualified human reviewers consistently evaluate how faith-facing AI systems use sources, handle authority, defer to people and institutions, preserve human agency, and respect pastoral boundaries? Where do automated model judges diverge from those human judgments?
Evaluation science
FID-044
Cross-Faith Benchmark Validity and Measurement Design
How should cross-faith AI benchmarks validate what they measure when scores may depend on question sourcing, user expectations, LLM-as-judge behavior, scoring thresholds, regenerated answers, model updates, and the difference between any religious mention and meaningful representation?
Evaluation science
FID-045
Faith-AI Research Gap Map and Evidence Commons
What does the current AI ethics, safety, fairness, HCI, and evaluation literature actually study about religion and faith, what does it omit, and how should Fide AI maintain a living evidence map that guides future research rather than duplicating or overstating existing work?
Evaluation science
FID-006
Faith-Facing Retrieval Grounding and Citation Reliability
How reliably do faith-facing AI systems retrieve, cite, and represent religious sources when users ask theological, historical, pastoral, or institution-specific questions?
Grounding and truthfulness
FID-028
Christian Source Authority and RAG
Can Christian RAG systems distinguish and correctly use Scripture, creeds, confessions, councils, catechisms, denominational policies, patristic sources, commentaries, sermons, blogs, and academic theology?
Grounding and truthfulness
FID-056
Formal Verification for Sacred Text Fidelity
How can faith-facing AI systems be formally checked for whether they quote, paraphrase, reference, and contextualize sacred texts faithfully within a specified text edition, translation, canon, and interpretive context? Fide AI's first two studies address exact English Scripture quotation and whether language models consult an available authoritative source; the broader research call remains open.
Grounding and truthfulness
FID-057
Proof-Carrying Citations for Faith-Facing AI
Can faith-facing AI answers carry checkable citation proofs that show which claims are directly supported by sources, which are inferred, which are uncertain, and which require human or tradition-specific authority?
Grounding and truthfulness
FID-058
Tradition-Specific Constraint Formalization
How can tradition-specific boundaries, source hierarchies, doctrinal constraints, and disagreement patterns be translated into machine-checkable specifications without flattening differences across faith traditions?
Grounding and truthfulness
FID-059
Authority-Boundary Verification for Pastoral-Adjacent AI
Can faith-facing AI systems be verified for whether they preserve the boundary between explanation, spiritual encouragement, moral reflection, pastoral or clerical authority, clinical/legal advice, and situations requiring human care?
Grounding and truthfulness
FID-060
Cross-Faith Sacred Text and Source Schema
What metadata schema is needed for faith-facing AI systems to represent sacred texts, commentaries, institutional documents, oral traditions, translations, editions, and authority levels across faith traditions?
Grounding and truthfulness
FID-061
Theological Contradiction and Entailment Stress Tests
Can faith-facing AI systems be tested for whether their answers contradict, entail, overstate, understate, or misrepresent claims within a bounded source set and specified faith tradition?
Grounding and truthfulness
FID-062
Verified Retrieval Pipelines for Faith-Facing RAG
How can faith-facing retrieval-augmented generation pipelines be verified for whether they retrieve authoritative, relevant, context-preserving sources before generating answers about sacred texts, doctrine, practice, or institutional policy?
Grounding and truthfulness
FID-063
Human-Reviewer-to-Formal-Spec Translation
How can theologians, clergy, scholars, ministry practitioners, and community reviewers translate qualitative judgments about faith-facing AI into formal specifications that are faithful to expert intent and usable in evaluation?
Grounding and truthfulness