FideAI

Call collection

Faith-Facing AI Evaluation Infrastructure

Research on benchmark validity, formal verification, reviewer calibration, scorer reliability, red-team design, and proof-carrying citations.

21 calls

Open work

FID-001

Faith-Facing Model Comparison Platform

Can Fide AI build a public-interest evaluation platform that compares models, prompts, retrieval systems, agents, and full faith-facing product harnesses with the rigor expected from institutions like Arena, Artificial Analysis, and METR?

Evaluation science

FID-003

Held-Out Multi-Turn Pastoral Pressure Tests

Do faith-facing AI systems that perform well on single-turn benchmark items also handle multi-turn, emotionally loaded, pastoral-adjacent situations without fabricating authority, overcomplying, missing escalation, or replacing human care?

Evaluation science

FID-008

Evaluation-Awareness and Faith-Facing Honesty Tests

Do faith-facing AI systems behave differently when they recognize they are being evaluated, and can domain-specific honesty or integrity framings reduce evaluation gaming without creating new failure modes?

Evaluation science

FID-011

Reviewer Reliability for Faith-Facing AI Evaluation

What reviewer configurations produce reliable, fair, and interpretable scores for faith-facing AI outputs?

Evaluation science

FID-012

Optimization Pressure and Visible-Rubric Gaming

If builders can see Fide AI rubrics or optimize against public benchmark items, do systems become genuinely safer or merely better at passing the visible test?

Evaluation science

FID-023

Faith-Facing Red-Team Suite

What red-team probes are needed to expose failures unique to faith-facing AI systems and adjacent high-trust guidance systems?

Evaluation science

FID-070

Contextual Integrity and Prompt-Injection Resilience for Faith-Facing Agents

When a faith-facing agent reads email, web pages, agendas, documents, knowledge bases, and service requests, how can it distinguish authorized institutional instruction from untrusted context that attempts to redirect its behavior, exfiltrate information, or induce an unsafe action?

Agent alignment

FID-074

Agent Alignment and Runtime Assurance

How can organizations determine whether AI agents remain aligned with human intent and institutional policy while they plan, use tools, delegate work, and act? What evidence and interventions can reveal and stop consequential deviations before they become failures?

Agent alignment

FID-002

Validating Human and AI Judgments of Faith-Facing Systems

Can qualified human reviewers consistently evaluate how faith-facing AI systems use sources, handle authority, defer to people and institutions, preserve human agency, and respect pastoral boundaries? Where do automated model judges diverge from those human judgments?

Evaluation science

FID-044

Cross-Faith Benchmark Validity and Measurement Design

How should cross-faith AI benchmarks validate what they measure when scores may depend on question sourcing, user expectations, LLM-as-judge behavior, scoring thresholds, regenerated answers, model updates, and the difference between any religious mention and meaningful representation?

Evaluation science

FID-045

Faith-AI Research Gap Map and Evidence Commons

What does the current AI ethics, safety, fairness, HCI, and evaluation literature actually study about religion and faith, what does it omit, and how should Fide AI maintain a living evidence map that guides future research rather than duplicating or overstating existing work?

Evaluation science

FID-006

Faith-Facing Retrieval Grounding and Citation Reliability

How reliably do faith-facing AI systems retrieve, cite, and represent religious sources when users ask theological, historical, pastoral, or institution-specific questions?

Grounding and truthfulness

FID-028

Christian Source Authority and RAG

Can Christian RAG systems distinguish and correctly use Scripture, creeds, confessions, councils, catechisms, denominational policies, patristic sources, commentaries, sermons, blogs, and academic theology?

Grounding and truthfulness

FID-056

Formal Verification for Sacred Text Fidelity

How can faith-facing AI systems be formally checked for whether they quote, paraphrase, reference, and contextualize sacred texts faithfully within a specified text edition, translation, canon, and interpretive context? Fide AI's first two studies address exact English Scripture quotation and whether language models consult an available authoritative source; the broader research call remains open.

Grounding and truthfulness

FID-057

Proof-Carrying Citations for Faith-Facing AI

Can faith-facing AI answers carry checkable citation proofs that show which claims are directly supported by sources, which are inferred, which are uncertain, and which require human or tradition-specific authority?

Grounding and truthfulness

FID-058

Tradition-Specific Constraint Formalization

How can tradition-specific boundaries, source hierarchies, doctrinal constraints, and disagreement patterns be translated into machine-checkable specifications without flattening differences across faith traditions?

Grounding and truthfulness

FID-059

Authority-Boundary Verification for Pastoral-Adjacent AI

Can faith-facing AI systems be verified for whether they preserve the boundary between explanation, spiritual encouragement, moral reflection, pastoral or clerical authority, clinical/legal advice, and situations requiring human care?

Grounding and truthfulness

FID-060

Cross-Faith Sacred Text and Source Schema

What metadata schema is needed for faith-facing AI systems to represent sacred texts, commentaries, institutional documents, oral traditions, translations, editions, and authority levels across faith traditions?

Grounding and truthfulness

FID-061

Theological Contradiction and Entailment Stress Tests

Can faith-facing AI systems be tested for whether their answers contradict, entail, overstate, understate, or misrepresent claims within a bounded source set and specified faith tradition?

Grounding and truthfulness

FID-062

Verified Retrieval Pipelines for Faith-Facing RAG

How can faith-facing retrieval-augmented generation pipelines be verified for whether they retrieve authoritative, relevant, context-preserving sources before generating answers about sacred texts, doctrine, practice, or institutional policy?

Grounding and truthfulness

FID-063

Human-Reviewer-to-Formal-Spec Translation

How can theologians, clergy, scholars, ministry practitioners, and community reviewers translate qualitative judgments about faith-facing AI into formal specifications that are faithful to expert intent and usable in evaluation?

Grounding and truthfulness