Does it use sources faithfully?
Test whether a system preserves what its sources say, distinguishes evidence from interpretation, and handles missing or conflicting information.
Evaluation and assurance
We investigate how AI systems behave in consequential work. Independent evaluation helps you understand the failures, identify necessary changes, and decide whether the evidence supports the intended use.
Discuss an evaluation →What we investigate
Sources, instructions, tools, interfaces, and human decisions shape what an AI system actually does. We scope the evaluation around the complete workflow and the risks of its intended setting.
Test whether a system preserves what its sources say, distinguishes evidence from interpretation, and handles missing or conflicting information.
Investigate whether instructions, tool permissions, and delegation keep actions within the responsibilities the system has actually been given.
Examine whether people can inspect consequential behavior, correct mistakes, and intervene when the system should hand responsibility back.
Before release, before deployment, and as systems change. A new model, different tools, or greater autonomy can warrant reassessment. The scope determines which questions we can answer.
Published evidence
Our published evaluations begin in faith and religious life, where source fidelity and appropriate authority matter. Each report states its methods and limits; results in these settings do not establish performance in another domain.
published
2026-08-26
6 systems
A controlled study of 4,800 Scripture quotation requests. Delegation remained near 95% under neutral requests, but when users discouraged tool use it fell to 30.6% under discretionary availability and remained at 84.7% under a higher-priority source requirement.
Important limit: Behavior within the declared fixed panel and run window; not a model leaderboard, theological evaluation, product endorsement, or guarantee of source use.
Open report →
source delegation · tool use · instruction hierarchy
published
2026-08-05
6 systems
A public report for Paper 01 of the FID-056 research call, covering 8,640 matched Scripture requests routed through four delivery designs, where failure moves when a source is connected, and interpretation limits.
Important limit: Delivery-condition results within the declared study, not a model leaderboard; not theological correctness, pastoral safety, or legal compliance.
Open report →
quotation fidelity · retrieval · deterministic rendering
published
2026-06-01
14 systems
A public report for the first FMG-Bench release, covering model comparison results, how guidance changed responses, caveats, and interpretation limits.
Important limit: Benchmark behavior only; not theological authority, pastoral authority, certification, or product endorsement.
Open report →
benchmark · model comparison · pastoral triage
How an engagement works
A credible evaluation makes its assumptions visible. We state the relevant domain standards, values, and limits, then design the work around what the evidence needs to establish.
Start with one consequential workflow. A scoped research and engineering partnership can develop repeatable tests of source use, permissions, delegation, and human intervention, then use them to assess changes. Your team retains responsibility for the application and deployment.
Discuss a design partnership ↗Name the system, users, setting, and decision. Assess the fit between your question and the methods, expertise, and access available. Agree deliverables, schedule, and fee before work begins.
Identify the relevant instructions, sources, tools, controls, and operational records. Agree what can be inspected, minimize sensitive-data access, and document gaps that constrain the evaluation.
Run realistic and adversarial scenarios tied to the intended use. Examine available traces, actions, outcomes, and human intervention, including longer workflows where relevant.
Explain what happened, what remains uncertain, and which changes deserve priority. Agree any targeted retesting needed to assess a fix or a change in the system.
What you receive
The proposal sets the exact outputs. A typical evaluation brings the findings, their implications, and the practical next steps together.
Confidentiality and reporting terms are agreed in advance. Private access does not authorize publication: identifying an organization or sharing a method, dataset, or case study from an engagement requires separate permission.
Start a conversation
Tell us what the system does, who will use it, and what you need to establish before its release, deployment, or expansion. We start by assessing whether we can meaningfully answer your question.
Discuss your system ↗