Featured study · Evaluation science
When AI Is Your Pastor
Do clearer instructions improve how AI handles theological, moral, and pastoral-adjacent questions?
What the study found
Guided responses scored higher in every question category.
- Frontier models evaluated
- 14
- Scored benchmark items
- 8,792
- Base scenarios
- 120
A benchmark of theological and pastoral-adjacent responses; higher rubric scores do not establish pastoral competence.