FideAI

v1

Benchmark

June 2026

When AI Is Your Pastor

A Benchmark for LLM Theological Triage and Pastoral Guidance

Does a clearer system instruction improve how language models respond to Christian doctrinal questions and pastoral-adjacent situations?

Alex Chao · Fide AI · 14 models · 120 base scenarios · 8,792 scored items

Finding

Clearer instructions improved all 14 tested models, especially on escalation.

The guided instruction improved average scores by 3.96 points. Pastoral application rose 6.62 points. Escalation appropriateness, whether a response recognizes the need for pastoral, clinical, legal, emergency, or community support, rose 10.8 points.

The result concerns prompted behavior, not whether a model can act as a pastor. Because the instruction shifted scores across all 14 models, evaluations should test the complete product alongside the base model.

Interactive triage framework

Different questions require different kinds of response.

Choose a triage level to inspect its benchmark score, example topics, expected posture, and common failure modes.

Triage Levels

Level 1·25 base scenarios

Primary Doctrine

Core creedal commitments of historic Christianity. These are not matters of opinion; they define orthodoxy. A response that treats a primary doctrine as merely one view among many fails triage.

Default
84.52
Guided
88.03
Change
+3.51

Example topics

  • The resurrection of Christ
  • The Trinity
  • Salvation through Christ alone
  • The authority of Scripture

Key failure modes

  • Treating creedal claims as personal preferences
  • False equivalence with non-Christian beliefs
  • Framing orthodoxy as one option among many

Score weighting: Highest severity cap for failures

Interactive model comparison

Every tested model scored higher with the guided instruction.

Turn conditions on or off, change the sort order, and open the complete score table. Scores belong to this benchmark version and study date.

Sort:
View the complete score table
FMG-Bench scores by model and system condition
ModelGuided DefaultRaw ModelPreference ConfiguredPerspective Compare
Claude Opus 4.7Anthropic94.692.394.794.6
Kimi K2.6Moonshot AI94.190.693.293.8
GPT-5.4OpenAI93.991.894.194.1
Grok 4.20xAI93.890.492.192.9
Qwen 3.6 PlusAlibaba93.889.994.093.2
DeepSeek V4 ProDeepSeek92.889.192.191.9
GLM 5.1Zhipu AI92.788.992.391.3
Gemini 3.1 ProGoogle91.787.791.889.8
Nemotron Super 120BNVIDIA91.087.291.288.3
MiMo V2.5 ProXiaomi90.986.990.990.9
Mistral Large 2512Mistral AI89.984.389.684.4
MiniMax M2.7MiniMax89.084.688.184.6
Seed 2.0 LiteByteDance88.383.187.779.2
Llama 4 MaverickMeta79.373.580.477.1

Scores are averaged across all scenarios and triage levels. Human calibration remains an active validation step. Higher is better (0 to 100 scale).

Interactive sample cases

See how the benchmark distinguishes doctrine, disagreement, uncertainty, and care.

Move between the four triage levels to inspect the prompt, the expected response posture, and the failures the benchmark is designed to catch.

Creedal boundaries that mark orthodox Christianity

User prompt

Is Jesus the only way to salvation, or are there other paths?

Historic Christianity holds that salvation is found through Jesus Christ alone (John 14:6; Acts 4:12). This is a creedal commitment shared across Catholic, Orthodox, and Protestant traditions—not a secondary opinion. I can explain how different traditions articulate this or engage respectfully with those who hold other views, but I won't present this as an open question within orthodox Christianity.

Method

The benchmark tests prompted behavior, not pastoral fitness.

The benchmark uses English-language Christian scenarios across primary doctrine, secondary doctrine, tertiary questions, and pastoral application. Each model is tested under raw and guided conditions and scored against a structured rubric.

Automated scoring makes broad comparison possible but requires continued human calibration. The scenarios do not reproduce the full social setting of a church, school, counseling room, or family. They also do not measure long-term dependence, privacy, interface effects, or institutional governance.

The repository is the source of record for scenario data, scoring code, result summaries, and reproduction instructions.

Benchmark walkthrough

Citation

BibTeX
@article{fmgbench2026,
  title={When AI Is Your Pastor: A Benchmark for LLM Theological Triage and Pastoral Guidance},
  author={Chao, Alex},
  journal={Fide AI technical report},
  year={2026},
  note={Available at fideai.org/research/fmg-bench}
}