Evaluation archive
Published reports and reusable artifacts.
Each entry states its system scope, evidence status, and claims limit so readers can distinguish a bounded evaluation from an endorsement.
3 reports
Back to evaluation servicespublished
2026-08-26
6 systems
FID-056-P02 Source Delegation
A controlled study of 4,800 Scripture quotation requests. Delegation remained near 95% under neutral requests, but when users discouraged tool use it fell to 30.6% under discretionary availability and remained at 84.7% under a higher-priority source requirement.
Important limit: Behavior within the declared fixed panel and run window; not a model leaderboard, theological evaluation, product endorsement, or guarantee of source use.
Open report →
source delegation · tool use · instruction hierarchy
published
2026-08-05
6 systems
FID-056-P01 Scripture Quotation Fidelity
A public report for Paper 01 of the FID-056 research call, covering 8,640 matched Scripture requests routed through four delivery designs, where failure moves when a source is connected, and interpretation limits.
Important limit: Delivery-condition results within the declared study, not a model leaderboard; not theological correctness, pastoral safety, or legal compliance.
Open report →
quotation fidelity · retrieval · deterministic rendering
published
2026-06-01
6 systems
FMG-Bench v1 Model Comparison
A public report for the first FMG-Bench release, covering model comparison results, how guidance changed responses, caveats, and interpretation limits.
Important limit: Benchmark behavior only; not theological authority, pastoral authority, certification, or product endorsement.
Open report →
benchmark · model comparison · pastoral triage