FideAI

Independent AI research lab

How do we know when AI deserves our trust?

Measurement science for trustworthy AI.

We help AI labs and organizations in high-trust domains understand and build trustworthy AI systems.

Public evidence · September 2026

The call for independent evaluation

01 / 13

Public statements, not endorsements of Fide AI.

Why now

Independent evaluation is becoming a shared priority.

Frontier Labs have announced commitments to independent evaluators with employee-like access. Turning that access into trustworthy judgments requires rigorous methods, reproducible evidence, and domain expertise.

“Scaling AI systems has to be constrained by our confidence in safety.”
Jakub Pachocki · Chief Scientist, OpenAIAn Alien Mind · September 6, 2026 ↗

Fide’s contribution

Trust depends on how AI behaves in context.

A trustworthy system must do more than produce a plausible answer. It must use the right evidence, respect the limits of its role, and keep people meaningfully in control. Fide investigates these questions through technical evaluation informed by domain expertise.

01Evidence
Do the sources, experiments, and records support the system’s claims?
02Authority
Does the system stay within the instructions, permissions, and responsibilities it has been given?
03Human control
Can people recognize failures, intervene, and retain meaningful authority over what happens next?

Research in development

Can we trust AI-generated safety evidence?

Fide is developing methods to help independent reviewers determine whether AI-generated safety reports are supported by the experiments and records behind them. We will compare report-only review with review grounded in those records, measuring missed failures, false alarms, and the cost of review.

Discuss a research collaboration →

Proposed research. This program has not yet produced published results.

  1. 01

    Evidence fidelity

    Does the report reflect the tests performed, including failed or incomplete runs?

  2. 02

    Authority boundaries

    Does the agent stop or seek approval when its permissions change?

  3. 03

    Effective review

    Which records help a reviewer detect failures at a practical cost?

Featured study brief

An open brief for the next investigation.

Explore developing research →

Cybersecurity

Protocol development

When Security Evaluations Go Stale

Fide is developing a study of whether targeted retesting catches meaningful performance declines, or gives false reassurance. We begin with AI malware-report analysis, comparing smaller retests against complete benchmark reruns.

Protocol and evaluation tooling in development. No model-performance findings yet.

Explore the study brief

Updated

Research independence

Methods you can inspect. Judgments you can question.

Every research program reflects judgments about what matters and how evidence should be interpreted. We make those choices visible so readers can examine our methods and challenge our conclusions.

  1. State the lens

    Disclose the assumptions, prior commitments, and worldviews that materially shape each project.

  2. Show the evidence

    Publish methods, data, limitations, and uncertainty so others can inspect the work.

  3. Separate evidence from judgment

    Distinguish measured findings from interpretation, recommendation, and moral judgment.

Fide retains control of its conclusions. Payment cannot determine favorable findings. We disclose material limits on access and what those limits mean for the conclusions we can draw.

Work with Fide

Independent evidence for consequential AI decisions.

An evaluation produces a scoped finding, its supporting evidence, and the limits of the conclusion. Where the evidence shows a failure, we identify changes to investigate and tests that can assess whether they help.

The starting point depends on the question you bring. We agree the decision, access, methods, and deliverables before work begins.

  1. 01

    AI labs and evaluators

    Develop methods to check whether safety reports match the experiments and records behind them.

    Explore a research collaboration

  2. 02

    Institutions and builders

    Bring a system and a release, deployment, or expansion decision. Start with one question and an agreed evaluation scope.

    Scope an evaluation

  3. 03

    Research supporters

    Help fund a defined study with a reproducible protocol, findings, and a public research contribution.

    Explore research support

Newsletter

Read new work by email.

Substack carries new essays and research notes. The permanent archive stays here on Fide AI.