FideAI

FID-001 · Being scoped

Faith-Facing Model Comparison Platform

Can Fide AI build a public-interest evaluation platform that compares models, prompts, retrieval systems, agents, and full faith-facing product harnesses with the rigor expected from institutions like Arena, Artificial Analysis, and METR?

Why the question remains open

Faith institutions need evidence before adopting systems that appear theological, pastoral, or morally authoritative. Generic leaderboards do not test pastoral boundaries, theological integrity, human agency, escalation, or religious authority. Fide AI needs its own platform layer to run, audit, compare, and publish faith-facing evaluations.

Working hypothesis

A proposition to test, not a finding.

A platform built around run manifests, artifact custody, reviewer workflows, and public claims limits can produce more decision-relevant evidence than isolated benchmark scripts or one-off reports.

Needed controls

What must constrain the study

  • 01Stable model endpoint/version metadata.
  • 02Repeatable prompt and harness configuration.
  • 03Artifact hashes for inputs, outputs, prompts, corpora, and scoring code.
  • 04Publication review before claims are exposed publicly.

Expected outputs

Artifacts the work should produce

  • 01Fide run manifest schema.
  • 02Internal model/system comparison dashboard.
  • 03Public comparison reports.
  • 04Leaderboard snapshots with confidence intervals and claims limits.

Open questions

Uncertainties the protocol must resolve

  • 01Whether Opik, Langfuse, Phoenix, or a custom trace layer should be used.
  • 02How much raw output can be published safely.
  • 03Whether public rankings should be aggregate scores, category scores, or evidence cards rather than a single leaderboard.

Being scoped

Open work

Primary need: eval engineering

  • Build runner adapters.
  • Design the run manifest schema.
  • Contribute reviewer workflow UI.
  • Audit license and white-label constraints of candidate OSS platforms.