FideAI

Evaluation and assurance

Independent assurance for trustworthy AI.

We investigate how AI systems behave in consequential work. Independent evaluation helps you understand the failures, identify necessary changes, and decide whether the evidence supports the intended use.

Discuss an evaluation →

What we investigate

The model is only one part of the system.

Sources, instructions, tools, interfaces, and human decisions shape what an AI system actually does. We scope the evaluation around the complete workflow and the risks of its intended setting.

Does it use sources faithfully?

Test whether a system preserves what its sources say, distinguishes evidence from interpretation, and handles missing or conflicting information.

Does it act within its authority?

Investigate whether instructions, tool permissions, and delegation keep actions within the responsibilities the system has actually been given.

Can people meaningfully oversee it?

Examine whether people can inspect consequential behavior, correct mistakes, and intervene when the system should hand responsibility back.

Before release, before deployment, and as systems change. A new model, different tools, or greater autonomy can warrant reassessment. The scope determines which questions we can answer.

Published evidence

Inspect the work behind the approach.

Our published evaluations begin in faith and religious life, where source fidelity and appropriate authority matter. Each report states its methods and limits; results in these settings do not establish performance in another domain.

Browse all evaluation reports →

How an engagement works

One decision. An agreed scope.

A credible evaluation makes its assumptions visible. We state the relevant domain standards, values, and limits, then design the work around what the evidence needs to establish.

Building an agent? Explore a design partnership.

Start with one consequential workflow. A scoped research and engineering partnership can develop repeatable tests of source use, permissions, delegation, and human intervention, then use them to assess changes. Your team retains responsibility for the application and deployment.

Discuss a design partnership ↗
  1. 01

    Agree the question and scope

    Name the system, users, setting, and decision. Assess the fit between your question and the methods, expertise, and access available. Agree deliverables, schedule, and fee before work begins.

  2. 02

    Map the system and evidence

    Identify the relevant instructions, sources, tools, controls, and operational records. Agree what can be inspected, minimize sensitive-data access, and document gaps that constrain the evaluation.

  3. 03

    Test realistic failure modes

    Run realistic and adversarial scenarios tied to the intended use. Examine available traces, actions, outcomes, and human intervention, including longer workflows where relevant.

  4. 04

    Review findings and next steps

    Explain what happened, what remains uncertain, and which changes deserve priority. Agree any targeted retesting needed to assess a fix or a change in the system.

What you receive

Evidence you can act on.

The proposal sets the exact outputs. A typical evaluation brings the findings, their implications, and the practical next steps together.

Findings with supporting evidence
Observed behavior, representative examples, failure patterns, and explicit limitations under the agreed scope.
A bounded readiness judgment
A recommendation for the named use: proceed, restrict use, remediate and retest, keep internal, or do not deploy. Access and evidence limits remain part of the judgment.
Priorities for remediation
Practical changes to sources, permissions, controls, interfaces, escalation, or deployment policy, with a path to checking whether they help.
A briefing for the decision makers
A working session for the people responsible for the system, its governance, and the final deployment decision.

Confidentiality and reporting terms are agreed in advance. Private access does not authorize publication: identifying an organization or sharing a method, dataset, or case study from an engagement requires separate permission.

Start a conversation

What decision do you need to make?

Tell us what the system does, who will use it, and what you need to establish before its release, deployment, or expansion. We start by assessing whether we can meaningfully answer your question.

Discuss your system ↗