Before release
Investigate a model or product's capabilities, failure modes, safeguards, and claims before wider access. Develop targeted methods when existing tests leave the question unresolved.
Evaluation and assurance
Fide AI develops evaluation methods and investigates complete AI systems to help organizations understand whether they can rely on them for consequential work. We test behavior, investigate failures, and identify the changes needed to support more confident decisions.
Discuss your system →A new model, different tools, or greater autonomy can change what a system can safely do. We scope each investigation around the system, access, expertise, and decision involved.
Investigate a model or product's capabilities, failure modes, safeguards, and claims before wider access. Develop targeted methods when existing tests leave the question unresolved.
Assess the assembled system against its intended workflow, users, sources, permissions, and human oversight. Build evidence for procurement, launch, or a controlled pilot.
Reassess changes to models, tools, sources, and autonomy. Investigate failures, verify fixes, and use periodic reviews or targeted retesting to inform continued use and expansion.
Evaluation method
Trustworthiness is not a general property or a single score. It is a bounded claim about a particular system, use, group of users, and set of conditions.
The assumptions, domain standards, and worldviews that shape the evaluation are documented so the result can be understood, challenged, and used responsibly.
In cybersecurity, an evaluation can examine whether a defensive agent stays within authorized scope, resists prompt injection, respects permissions, and leaves evidence that supports investigation and human intervention. Testing requires agreed access and a controlled environment; it is not a blanket security certification.
Name the system, intended use, users, setting, and the decision the evaluation must inform.
Document the model, instructions, sources, retrieval, tools, interface, logging, correction process, and human oversight.
Determine which operational records are needed to assess sources, permissions, tool calls, delegation, actions, outcomes, and human intervention while minimizing access to sensitive data.
Run realistic and adversarial scenarios, including long-horizon tasks when relevant, using methods tied to Fide AI research and the risks of the intended setting.
Document behavior, failure patterns, uncertainty, access limitations, response options, remediation priorities, and a bounded readiness judgment.
What we test
A capable base model can still fail when it is connected to weak sources, conflicting instructions, unsafe tools, misleading interface claims, or inadequate human oversight. The evaluation scope is determined by the decision, available access, and potential consequences of failure.
Agent systems
Pre-deployment tests cannot anticipate every path an agent may take. When access permits, we assess whether the deployed system leaves enough operational evidence to identify departures from policy and whether responsible people can intervene before a failure compounds.
Read our call for research →Observe
Are tasks, sources, tools, permissions, actions, and outcomes sufficiently visible?
Interpret
Can expected behavior be distinguished from error, drift, manipulation, or unauthorized action?
Respond
Can people correct, constrain, pause, or stop the system and verify recovery?
Operational evidence is scoped to the decision and the minimum necessary data. It does not assume access to private chain-of-thought or authorize general employee or user surveillance.
What you receive
Observed behavior, failure patterns, representative examples, uncertainty, and limitations under the agreed scope.
A bounded recommendation for the named use: proceed, proceed with restrictions, remediate and retest, keep internal, or do not deploy.
Prioritized changes to sources, prompts, permissions, controls, interface language, review, escalation, monitoring, or deployment policy.
A working session for the people responsible for the product, deployment, governance, or final decision.
With permission, a de-identified method, checklist, failure pattern, dataset, or case study that can strengthen the wider evaluation field.
Ways to work together
Some teams need a confidential decision. Others can contribute a method or finding to the wider field. Public learning is optional and requires separate agreement.
A confidential engagement for teams that need an independent findings report, remediation priorities, and a technical briefing. Fide AI does not identify the organization without written permission.
A focused assessment for a specific launch, procurement, expansion, or restriction decision. The report states what the evidence supports and what remains unknown.
An evaluation designed to produce both a decision for the participating organization and a reusable public contribution. Any public method, scenario, finding, or case study requires separate permission.
Design partnerships
We are seeking design partners developing AI agents for consequential work. Together, we will define the actions the system needs to perform reliably, test how it uses sources and stays within its authority, and develop repeatable evaluations for future changes.
Discuss a design partnership ↗Agree on the decision, operating conditions, and actions to evaluate.
Investigate failures and test whether scoped changes address them.
Build a repeatable basis for assessing changes to the model, tools, sources, or permissions.
Your team owns the application and deployment. Fide focuses on evaluation, evidence, and the changes that evidence shows are needed.
This is a scoped research and engineering collaboration, not a certification or a guarantee of safe deployment.
Evaluation reports
Each artifact states what was tested, the evaluation conditions, the result, and the limits on what can be concluded.
published
2026-08-26
6 systems
A controlled study of 4,800 Scripture quotation requests. Delegation remained near 95% under neutral requests, but when users discouraged tool use it fell to 30.6% under discretionary availability and remained at 84.7% under a higher-priority source requirement.
Important limit: Behavior within the declared fixed panel and run window; not a model leaderboard, theological evaluation, product endorsement, or guarantee of source use.
Open report →
source delegation · tool use · instruction hierarchy
published
2026-08-05
6 systems
A public report for Paper 01 of the FID-056 research call, covering 8,640 matched Scripture requests routed through four delivery designs, where failure moves when a source is connected, and interpretation limits.
Important limit: Delivery-condition results within the declared study, not a model leaderboard; not theological correctness, pastoral safety, or legal compliance.
Open report →
quotation fidelity · retrieval · deterministic rendering
published
2026-06-01
6 systems
A public report for the first FMG-Bench release, covering model comparison results, how guidance changed responses, caveats, and interpretation limits.
Important limit: Benchmark behavior only; not theological authority, pastoral authority, certification, or product endorsement.
Open report →
benchmark · model comparison · pastoral triage
From evaluation to research
Most evaluations begin with a concrete decision for one developer or institution. Repeated findings can also expose broader gaps in evaluation science.
When appropriate, Fide AI turns recurring failure patterns into public scenarios, benchmark updates, reviewer protocols, readiness criteria, or new research questions. Private access does not automatically become public evidence. Any public artifact requires separate permission, documentation, and claims review.
Work with Fide AI
Share what the system does, who will use it, where it will be used, what access can be provided, and what decision the evaluation should inform.