FID-044 · Being scoped
Cross-Faith Benchmark Validity and Measurement Design
How should cross-faith AI benchmarks validate what they measure when scores may depend on question sourcing, user expectations, LLM-as-judge behavior, scoring thresholds, regenerated answers, model updates, and the difference between any religious mention and meaningful representation?
Why the question remains open
Recent cross-faith benchmark releases make religious representation and conversion asymmetry measurable, but the measurement choices themselves are high-stakes. A benchmark that rewards any religious mention may miss inaccurate, tokenizing, or shallow representation. A benchmark that penalizes persuasion may misread legitimate pastoral or tradition-specific speech. Fide AI can help make faith-facing benchmark claims more reproducible, interpretable, and humble.
Working hypothesis
A proposition to test, not a finding.
Cross-faith benchmark results will be useful only when accompanied by construct definitions, uncertainty estimates, judge-validation studies, held-out splits, regeneration variance, tradition-specific review, and clear claim boundaries. Without those controls, public leaderboards may create false precision or reward benchmark-shaped behavior.
Proposed method
How the question could be tested
- 01Reproduce selected cross-faith benchmark results on a small audited subset.
- 02Compare human reviewers, expert panels, trained non-experts, and multiple model judges on religious-representation and conversion-symmetry scoring.
- 03Measure sensitivity to prompt wording, answer regeneration, judge choice, scoring thresholds, model version changes, and item sourcing.
- 04Develop reporting standards for confidence intervals, construct limits, evaluator metadata, leaderboard claims, and public claims.
Needed controls
What must constrain the study
- 01Do not treat benchmark scores as direct measures of theological truth.
- 02Separate any mention, meaningful reference, balanced representation, accuracy, and pastoral appropriateness.
- 03Include religious and nonreligious reviewers.
- 04Track benchmark leakage, public/private splits, and model-update drift.
Relationship to existing work
This complements FID-002 on human calibration and construct validity, FID-011 reviewer reliability, and FID-012 visible-rubric gaming by focusing specifically on cross-faith benchmark design, public leaderboards, and claims about religious representation or persuasion.
Expected outputs
Artifacts the work should produce
- 01Cross-faith benchmark validation protocol.
- 02Reporting checklist for faith-facing AI leaderboards.
- 03Leaderboard claim audit template.
- 04Judge-human agreement study for religious representation and persuasion tasks.
- 05Recommendations for benchmark versioning and public claim boundaries.
Open questions
Uncertainties the protocol must resolve
- 01What minimum reliability is needed before a faith-facing benchmark can support procurement or product claims?
- 02When should disagreement among traditions be reported rather than averaged?
- 03How can open benchmarks remain useful without becoming easy to game?
Related calls
Continue through this research area
FID-012
Optimization Pressure and Visible-Rubric Gaming
If builders can see Fide AI rubrics or optimize against public benchmark items, do systems become genuinely safer or merely better at passing the visible test?
FID-076
Authorization Boundaries and AI Control in Cybersecurity
Which controls keep capable agents within legitimate authorization when task pressure, untrusted inputs, or delegated work creates opportunities to exceed it?
FID-078
When Trustworthiness Evaluations Transfer Across Domains
Which measures of evidence use, authority boundaries, and human control transfer across high-trust domains, and which require domain-specific definitions and calibration?
Being scoped
Open work
Primary need: benchmark validity, statistics, open evaluation infrastructure
- Reproduce benchmark subsets and audit scorer behavior.
- Serve as expert or trained non-expert reviewer.
- Design uncertainty, drift, and regeneration-variance analyses.
- Draft public-claims standards for faith-facing benchmark releases.