FideAI

FID-011 · Being scoped

Reviewer Reliability for Faith-Facing AI Evaluation

What reviewer configurations produce reliable, fair, and interpretable scores for faith-facing AI outputs?

Why the question remains open

Faith-facing evaluation depends on expert judgment, but expert disagreement is real. Fide AI needs to know when scores reflect stable constructs and when they reflect reviewer background, tradition, strictness, or rubric ambiguity.

Working hypothesis

A proposition to test, not a finding.

Reliability will vary by dimension. Factual source-grounding may show higher agreement than pastoral tone, theological judgment, or pluralism handling. Reviewer training and adjudication should improve reliability but may not solve tradition-bound disagreement.

Proposed method

How the question could be tested

  • 01Compare expert, trained non-expert, and model-judge scoring.
  • 02Track inter-rater reliability, disagreement tags, strictness, confidence, and reviewer background metadata.
  • 03Test rubric revisions on high-disagreement items.

Needed controls

What must constrain the study

  • 01Reviewer conflict-of-interest checks.
  • 02Blind model labels.
  • 03Balanced scenario sampling.
  • 04Public/private separation for reviewer metadata.

Expected outputs

Artifacts the work should produce

  • 01Reviewer reliability report.
  • 02Reviewer training protocol.
  • 03Adjudication policy.
  • 04Acceptable-evidence thresholds for public claims.

Open questions

Uncertainties the protocol must resolve

  • 01What reliability metric best fits mixed categorical and continuous rubrics?
  • 02How many reviewers are needed per high-stakes item?
  • 03When should disagreement be published rather than collapsed into a score?

Being scoped

Open work

Primary need: statistics, reviewer operations

  • Design reliability analysis.
  • Build reviewer assignment tooling.
  • Serve as reviewer or adjudicator.
  • Audit rubric wording.