FideAI
← All domains

Independent evaluation research

Cybersecurity

Trustworthy autonomous cyber defense.

Fide’s cybersecurity agenda applies measurement science to autonomous defense. We are developing independent evaluations to test whether AI defense teams withstand compromise, response actions preserve essential services, and repairs actually restore security.

Cybersecurity research priorities

Three priorities for dependable autonomous defense.

Our proposed research priorities address consequential challenges reflected in DARPA’s DICE, CASTLE, ANSR, and AI Cyber Challenge programs. Each brief connects a practical decision to a measurement gap and a concrete first experiment.

01 · Collective resilience

Resilient AI defense teams

Can a team of AI defenders keep protecting a network when one agent is compromised?

Proposed research · Seeking collaborators

Individual-agent scores do not establish how a team behaves when evidence, trust, and failures propagate between its members. Fide proposes to compare collective resilience at a fixed defensive task and resource budget.

First intended artifact: An executable evaluation suite for failure propagation and recovery in defensive agent teams.

Program contextDICECASTLE

02 · Effective defense under human control

Controlled autonomous response

Can an AI defender stop an attack without disrupting the systems it is meant to protect?

Proposed research · Seeking collaborators

A defender can contain a threat by taking unnecessarily disruptive action. Fide proposes to evaluate containment, service continuity, and authority compliance together, including cases where waiting for approval also carries a cost.

First intended artifact: A practitioner-reviewed evaluation protocol for autonomous response decisions.

Program contextCASTLEANSR

03 · Independent verification

Verified repair and recovery

How do we know an AI-generated fix actually restores security?

Proposed research · Seeking collaborators

A patch can pass the check that motivated it while failing a different security or functional check. Fide proposes to study how independent verification changes false acceptance and review cost, starting with released systems and reproducible cases.

First intended artifact: A reproducible independent verification protocol and a reviewed corpus of repair outcomes.

Program contextAI Cyber Challenge

These are proposed research directions, not reported findings. Program references explain the research context; they do not imply funding, partnership, endorsement, or an open solicitation.

Work in development · Protocol development

When Security Evaluations Go Stale

How much retesting is enough to detect a loss of reliability after an AI system changes?

Fide is developing a study of whether targeted retesting catches meaningful performance declines, or gives false reassurance. We begin with AI malware-report analysis, comparing smaller retests against complete benchmark reruns.

This initial study asks how much retesting is needed to detect a loss of reliability after an AI system changes. It develops our approach to checking whether earlier evaluation evidence still holds, starting with malware-report analysis. Our broader agenda applies independent evaluation to defense teams, response actions, and repairs.

Read the study brief →

No model-performance findings yet. Updated .

Selected research elsewhere

Research we build on.

Benchmarks, methods, and findings from researchers advancing defensive AI and its evaluation.

Cybersecurity

Anthropic

Research overview · 2025

Building AI for cyber defenders (external site)

Work on vulnerability discovery and patching, with a call for evaluations of defensive capabilities. A useful starting point for studying effective defense; the authors also discuss the gap between benchmarks and real security operations.

Source reviewed

AI monitoring

Anthropic Fellows, Redwood Research & Anthropic researchers

Research & benchmark · 2026

SLEIGHT-Bench: finding blind spots in AI monitors (external site)

A study of monitor failures using synthetic agent transcripts. A starting point for investigating oversight and shared blind spots, with limits on how synthetic cases represent deployed systems.

Source reviewed

These are the named organizations’ contributions. Inclusion does not imply partnership or endorsement.

Follow the research

Stay close to the question.

Interested in Fide’s cybersecurity research? Request updates on this topic by email, or subscribe to Fide’s broader essays and research notes.

Request topic updates →

Opens a draft in your email app. Send it to request updates.

Subscribe to all Fide updates on Substack ↗

Contribute to the research

Bring your perspective.

  • Security practitioners: help define the decisions and failure cases that evaluations should address.
  • Research teams: collaborate on defensive agent systems, simulation, runtime assurance, or independent verification.
  • Program teams and research supporters: help scope an independent evaluation contribution with concrete milestones and artifacts.
Discuss a contribution ↗

Opens our research inquiry form with this topic included.