FideAI

FID-044 · Being scoped

Cross-Faith Benchmark Validity and Measurement Design

How should cross-faith AI benchmarks validate what they measure when scores may depend on question sourcing, user expectations, LLM-as-judge behavior, scoring thresholds, regenerated answers, model updates, and the difference between any religious mention and meaningful representation?

Why the question remains open

Recent cross-faith benchmark releases make religious representation and conversion asymmetry measurable, but the measurement choices themselves are high-stakes. A benchmark that rewards any religious mention may miss inaccurate, tokenizing, or shallow representation. A benchmark that penalizes persuasion may misread legitimate pastoral or tradition-specific speech. Fide AI can help make faith-facing benchmark claims more reproducible, interpretable, and humble.

Working hypothesis

A proposition to test, not a finding.

Cross-faith benchmark results will be useful only when accompanied by construct definitions, uncertainty estimates, judge-validation studies, held-out splits, regeneration variance, tradition-specific review, and clear claim boundaries. Without those controls, public leaderboards may create false precision or reward benchmark-shaped behavior.

Proposed method

How the question could be tested

  • 01Reproduce selected cross-faith benchmark results on a small audited subset.
  • 02Compare human reviewers, expert panels, trained non-experts, and multiple model judges on religious-representation and conversion-symmetry scoring.
  • 03Measure sensitivity to prompt wording, answer regeneration, judge choice, scoring thresholds, model version changes, and item sourcing.
  • 04Develop reporting standards for confidence intervals, construct limits, evaluator metadata, leaderboard claims, and public claims.

Needed controls

What must constrain the study

  • 01Do not treat benchmark scores as direct measures of theological truth.
  • 02Separate any mention, meaningful reference, balanced representation, accuracy, and pastoral appropriateness.
  • 03Include religious and nonreligious reviewers.
  • 04Track benchmark leakage, public/private splits, and model-update drift.

Relationship to existing work

This complements FID-002 on human calibration and construct validity, FID-011 reviewer reliability, and FID-012 visible-rubric gaming by focusing specifically on cross-faith benchmark design, public leaderboards, and claims about religious representation or persuasion.

Expected outputs

Artifacts the work should produce

  • 01Cross-faith benchmark validation protocol.
  • 02Reporting checklist for faith-facing AI leaderboards.
  • 03Leaderboard claim audit template.
  • 04Judge-human agreement study for religious representation and persuasion tasks.
  • 05Recommendations for benchmark versioning and public claim boundaries.

Open questions

Uncertainties the protocol must resolve

  • 01What minimum reliability is needed before a faith-facing benchmark can support procurement or product claims?
  • 02When should disagreement among traditions be reported rather than averaged?
  • 03How can open benchmarks remain useful without becoming easy to game?

Being scoped

Open work

Primary need: benchmark validity, statistics, open evaluation infrastructure

  • Reproduce benchmark subsets and audit scorer behavior.
  • Serve as expert or trained non-expert reviewer.
  • Design uncertainty, drift, and regeneration-variance analyses.
  • Draft public-claims standards for faith-facing benchmark releases.