FideAI
← Cybersecurity research agenda

Cybersecurity · Collective resilience

Resilient AI defense teams

Can a team of AI defenders keep protecting a network when one agent is compromised?

Help shape this research →

Why this question matters

What we want to understand.

Autonomous defense may depend on agents sharing observations, delegating work, and checking one another. A single unreliable participant could turn that cooperation into a channel for misleading evidence or harmful decisions. We want to measure whether a defense team can contain a local failure while continuing to protect the services people depend on.

The decision it could inform

Evidence someone can act on.

Give builders evidence for choosing coordination and verification mechanisms that preserve useful defense when individual agents fail, disagree, or become compromised.

Grounded in consequential research

The programs informing this priority.

DARPA · Program reference

DICE: resilient, controllable agent collectives

DICE targets decentralized agent collectives that remain controllable under agent loss or compromise. Fide’s proposed angle is independent measurement of failure propagation and containment in a cyber-defense setting.

DARPA · Program reference

CASTLE: AI agents for network defense

CASTLE makes realistic defensive environments and repeatable cyber evaluation central to its approach. It informs the environment and outcome requirements for this proposed comparison.

Fide’s proposed contribution is our interpretation of these research needs. References do not imply funding, partnership, endorsement, or an open solicitation.

Proposed first study

A concrete starting point.

Qualify an existing defensive simulation, then evaluate a small agent team under controlled communication and agent failures. The first comparison would ask whether independent evidence checks and bounded delegation reduce cascading errors without disabling useful defense.

  1. 01

    Select a contained environment with observable defensive goals and legitimate service activity. Establish centralized and peer-coordinated baselines using the same tasks, tools, and resource limits.

  2. 02

    Introduce prespecified faults: a missing agent, delayed messages, or an agent supplying incorrect observations. Compare direct sharing with source verification and independent review. Keep development scenarios separate from evaluation scenarios.

  3. 03

    Measure task outcomes and failure propagation across paired scenarios and repeated runs. Test whether containment and recovery preserve defense, and report uncertainty and communication costs alongside failures.

What we would measure

Defensive task success, legitimate service availability, propagation of incorrect evidence, unauthorized actions, recovery time, and communication and inference cost.

What a useful contribution could be

  • An executable evaluation suite for failure propagation and recovery in defensive agent teams
  • A comparative report identifying when verification contains failures and when it undermines useful coordination

Scope and readiness

Proposed research; no team-resilience experiments or findings yet. A small simulation would establish a bounded comparison, not demonstrate DICE-scale coordination or operational resilience. Environment suitability and the contribution beyond existing fault-tolerance research must be established first.

Follow the research

Stay close to the question.

Interested in Resilient AI defense teams? Request updates on this topic by email, or subscribe to Fide’s broader essays and research notes.

Request topic updates →

Opens a draft in your email app. Send it to request updates.

Subscribe to all Fide updates on Substack ↗

Contribute to the research

Bring your perspective.

  • Researchers in multi-agent systems, defensive cyber simulation, distributed systems, and AI control.
  • Help qualify the data, challenge the proposed comparison, or refine the practical decision this research should inform.
Discuss a contribution ↗

Opens our research inquiry form with this topic included.

Part of a wider agenda

More questions in trustworthy cyber defense.

Study in development

When Security Evaluations Go Stale →

Our current revalidation study examines targeted retesting for AI malware-report analysis. Explore its protocol, current stage, and collaboration needs.