FID-001 · Being scoped
Faith-Facing Model Comparison Platform
Can Fide AI build a public-interest evaluation platform that compares models, prompts, retrieval systems, agents, and full faith-facing product harnesses with the rigor expected from institutions like Arena, Artificial Analysis, and METR?
Why the question remains open
Faith institutions need evidence before adopting systems that appear theological, pastoral, or morally authoritative. Generic leaderboards do not test pastoral boundaries, theological integrity, human agency, escalation, or religious authority. Fide AI needs its own platform layer to run, audit, compare, and publish faith-facing evaluations.
Working hypothesis
A proposition to test, not a finding.
A platform built around run manifests, artifact custody, reviewer workflows, and public claims limits can produce more decision-relevant evidence than isolated benchmark scripts or one-off reports.
Needed controls
What must constrain the study
- 01Stable model endpoint/version metadata.
- 02Repeatable prompt and harness configuration.
- 03Artifact hashes for inputs, outputs, prompts, corpora, and scoring code.
- 04Publication review before claims are exposed publicly.
Expected outputs
Artifacts the work should produce
- 01Fide run manifest schema.
- 02Internal model/system comparison dashboard.
- 03Public comparison reports.
- 04Leaderboard snapshots with confidence intervals and claims limits.
Open questions
Uncertainties the protocol must resolve
- 01Whether Opik, Langfuse, Phoenix, or a custom trace layer should be used.
- 02How much raw output can be published safely.
- 03Whether public rankings should be aggregate scores, category scores, or evidence cards rather than a single leaderboard.
Related calls
Continue through this research area
FID-012
Optimization Pressure and Visible-Rubric Gaming
If builders can see Fide AI rubrics or optimize against public benchmark items, do systems become genuinely safer or merely better at passing the visible test?
FID-076
Authorization Boundaries and AI Control in Cybersecurity
Which controls keep capable agents within legitimate authorization when task pressure, untrusted inputs, or delegated work creates opportunities to exceed it?
FID-078
When Trustworthiness Evaluations Transfer Across Domains
Which measures of evidence use, authority boundaries, and human control transfer across high-trust domains, and which require domain-specific definitions and calibration?
Being scoped
Open work
Primary need: eval engineering
- Build runner adapters.
- Design the run manifest schema.
- Contribute reviewer workflow UI.
- Audit license and white-label constraints of candidate OSS platforms.