FideAI

FID-002 · In progress

Validating Human and AI Judgments of Faith-Facing Systems

Can qualified human reviewers consistently evaluate how faith-facing AI systems use sources, handle authority, defer to people and institutions, preserve human agency, and respect pastoral boundaries? Where do automated model judges diverge from those human judgments?

Why the question remains open

Benchmarks are easy to publish and hard to validate. Fide AI should not rely on synthetic judges or aggregate scores unless it knows which dimensions are reliable enough to guide deployment, procurement, or public claims.

Working hypothesis

A proposition to test, not a finding.

Some dimensions will be reliable enough for decision support, while others will require narrower rubrics, more reviewer training, or removal from high-stakes claims.

Proposed method

How the question could be tested

  • 01Recruit a small calibrated reviewer panel.
  • 02Score a stratified subset of outputs from Fide AI's published faith-facing evaluation, FMG-Bench.
  • 03Measure inter-rater reliability, judge-human disagreement, strictness/leniency, failure-tag consistency, and disagreement concentration by scenario type.
  • 04Compare expert review, trained non-expert review, and model judge scores where feasible.

Needed controls

What must constrain the study

  • 01Reviewer conflict checks.
  • 02Blind model/system labels where practical.
  • 03Rubric versioning.
  • 04Adjudication protocol for high-disagreement items.
  • 05Public/private separation for sensitive notes.

Expected outputs

Artifacts the work should produce

  • 01Calibration report.
  • 02Reviewer protocol.
  • 03Updated guidance for faith-facing evaluation rubrics.
  • 04Evidence map classifying dimensions as decision-relevant, conditional, or not yet validated.

Open questions

Uncertainties the protocol must resolve

  • 01What minimum reliability threshold should Fide require for public claims?
  • 02How should theological tradition metadata be represented without tokenizing reviewers?
  • 03Which disagreement patterns are legitimate pluralism rather than rubric failure?

In progress

Open work

Primary need: expert reviewers, statistics

  • Serve as an expert reviewer.
  • Review scoring rubrics.
  • Help with reliability analysis.
  • Build annotation and adjudication tooling.