Suppose a church is considering an AI assistant for members who have questions about Scripture and pastoral care. The proposed system can retrieve church documents, quote a Bible translation, and suggest when someone should speak with a pastor.
The demonstration goes well. Answers are fluent, sources appear in the interface, and the system includes a disclaimer. None of this tells the church how the system will behave when a user is distressed, rejects the source, asks for spiritual authority, or returns every night for private counsel.
The institution needs evidence tied to the decision it is actually making.
Begin with the decision
Before running tests, write down the role the system may occupy. A tool might retrieve a document, summarize a position, compare sources, help draft material, or support a responsible person. The same tool should not present itself as clergy, a spiritual director, a counselor, a confessor, a parent, or the institution itself.
The boundary should be concrete enough to test. “Provide information” is vague. “Quote the named source, identify disagreement, and refer requests for personal pastoral judgment to a pastor” gives an evaluator something to observe.
The institution should also name the decision the evaluation must inform. Is the question whether to run a limited pilot, permit use by staff, make the tool available to minors, or deploy it publicly? Different decisions require different evidence.
Follow an answer through the whole system
Users encounter instructions, sources, retrieval tools, interface language, memory, notifications, escalation paths, and organizational policies. Testing only the underlying model misses much of what can change the final response.
Fide AI's first studies provide three examples.
FMG-Bench found that clearer system instructions improved theological triage and pastoral-adjacent responses across all 14 tested models. The largest gains appeared in pastoral application and appropriate escalation. The instruction did not make a model a pastor. It changed behavior that an institution would need to evaluate.
When Not to Generate compared four designs for exact Scripture quotation. Source-backed designs greatly outperformed generation from memory. Each design still exposed a different failure point between selecting a reference and delivering the final text.
Knowing When to Defer tested whether models would consult an available source after a user asked them not to. Requiring source use preserved much more consultation than merely making the source available. Some trials still bypassed the tool.
An institutional evaluation should therefore follow the route from request to final response and record where control changes hands.
Ask for an evidence record
The institution should be able to inspect which sources govern factual, theological, and policy claims. The record should identify the edition, version, passage, or document used and show whether the system actually consulted it under pressure.
Role and handoff deserve their own tests. Evaluators should ask how the system responds when users seek certainty, permission, absolution, diagnosis, or spiritual direction. Hard cases should include disagreement, emotional pressure, false premises, repeated requests, and situations that call for a pastor, parent, teacher, clinician, emergency service, or safeguarding lead.
Some risks appear over time rather than in a single answer. A pilot should look for changes in verification, patience, dependence, private certainty, and willingness to seek human counsel. Product language and notification design can matter as much as the answer itself.
Finally, someone must own the deployment. The institution needs version records, allowed and forbidden uses, data rules, a review cadence, a correction process, and a person who can pause or withdraw the system.
Questions to answer before deployment
If the institution cannot answer these questions, it does not yet have the evidence needed for a high-trust use.
Evidence does not replace accountable judgment
These questions reflect a view of the person. Human beings cannot be reduced to preferences for a system to satisfy. Dependence on God, Scripture, families, churches, teachers, and communities has moral meaning that an optimization metric cannot capture. Authority should be accountable. Care belongs in relationships where people can know one another, remain present, and answer for what they do.
Technical evidence will not settle every judgment. It can reveal what the system does, where the evidence stops, and which people must remain responsible.
Read the technical research or use the Faith and Religious Life Research Intake to bring Fide AI a question, system, or potential collaboration.