Research · Faith & Religious Life
14 frontier models evaluated
When AI Is Your Pastor
arXiv preprint
Do clearer instructions improve how AI handles theological, moral, and pastoral-adjacent questions?
Evaluation science
Research
Fide AI studies AI systems through the lens of making them more trustworthy. We research whether systems use evidence well, remain within legitimate authority, preserve human control, and stay accountable while they act.
This research develops and validates the methods independent assurance depends on. We publish evidence and public-good artifacts that others can inspect, challenge, and use in their own work.
Explore by domain
Our domain hubs bring together Fide’s work, future interests, and selected research from the field.
88 public research calls translate the agenda into questions others can take forward.
Explore the public agenda and contribute to the work in Fide AI Research on GitHub.
Published research
Our released studies currently focus on faith-facing systems. Each page states the question, method, result, artifacts, and limits so readers can inspect the evidence directly.
Research · Faith & Religious Life
14 frontier models evaluated
arXiv preprint
Do clearer instructions improve how AI handles theological, moral, and pastoral-adjacent questions?
Evaluation science
Research · Faith & Religious Life
8,640 matched requests
Can AI quote Scripture exactly? We tested four ways of producing a quotation and traced where each one can fail.
Source fidelity
Research · Faith & Religious Life
4,800 controlled observations
Will AI check an authoritative Scripture source when a user asks it to rely on memory?
Source fidelity & authority
Current research directions
These questions connect our published studies, open calls, and proposed methods. Individual project pages state whether a study is planned or underway.
Which trustworthiness evaluations transfer across domains?
An open research call on which measures of evidence, authority, and human control carry over—and which need domain-specific definitions and calibration.
Can organizations see when an agent departs from its task or policy?
Developing privacy-aware operational traces, deviation tests, escalation criteria, intervention controls, and recovery measures.
Do evaluation results support the decisions made from them?
Testing construct validity, reviewer reliability, benchmark gaming, and the gap between controlled evaluation and deployed behavior.
Will a system use the right source when pressure makes that inconvenient?
Measuring source delegation, retrieval, citation fidelity, instruction hierarchy, and deterministic delivery designs.
Where should AI defer to responsible people and institutions?
Studying authority boundaries, human handoff, oversight, relational substitution, and the effects of repeated use.
Research in development
Inspect the questions and methods behind our current study. Enterprise is an exploratory next direction; other proposals remain available for collaborators to develop.
Cybersecurity
Protocol developmentFide is developing a study of whether targeted retesting catches meaningful performance declines, or gives false reassurance. We begin with AI malware-report analysis, comparing smaller retests against complete benchmark reruns.
Protocol and evaluation tooling in development. No model-performance findings yet.
Updated
Interpretation limits
A result describes behavior under named versions, prompts, conditions, rubrics, and evaluation procedures. It does not establish general safety, product approval, or deployment readiness beyond the tested scope. Faith-domain studies also do not confer theological or pastoral authority.