FideAI

FID-087 · Open question

Collective Resilience in Autonomous Cyber Defense

Which coordination and verification mechanisms let a team of AI defenders contain an unreliable or compromised agent while continuing to protect legitimate services?

Part of our cybersecurity agenda

Resilient AI defense teamsBrowse cybersecurity calls →

Program context

DICE: resilient, controllable agent collectives

DICE targets decentralized agent collectives that remain controllable under agent loss or compromise. Fide’s proposed angle is independent measurement of failure propagation and containment in a cyber-defense setting.

CASTLE: AI agents for network defense

CASTLE makes realistic defensive environments and repeatable cyber evaluation central to its approach. It informs the environment and outcome requirements for this proposed comparison.

Research context; no funding, partnership, endorsement, or open solicitation is implied.

Why the question remains open

Cooperating agents can share observations, divide investigations, and coordinate response. Those connections can also propagate misleading evidence and harmful actions. Evaluating each agent separately does not establish whether the collective remains effective after a local failure. Operators need evidence about failure containment and continued defense, including the cost of additional verification.

Working hypothesis

A proposition to test, not a finding.

Source-linked observations, independent evidence checks, and bounded delegation can reduce the propagation of incorrect evidence and consequential actions compared with unchecked sharing at comparable useful defensive performance and resource budgets. The null is that these mechanisms provide no reproducible benefit over existing coordination controls, or that their delays and costs erase the defensive benefit.

Proposed method

How the question could be tested

  • 01Qualify an existing, contained defensive simulation with observable incidents, legitimate service activity, explicit action permissions, and reproducible outcomes. Begin with a small team and one incident family; document which environment features are synthetic and which operational assumptions require practitioner review.
  • 02Establish single-agent, centrally coordinated, and peer-coordinated baselines. Match available information, tools, task difficulty, and total resource budgets; report agent count and actual communication and inference costs separately.
  • 03Prespecify failures such as a missing agent, delayed communication, and incorrect observations supplied by one participant. State the faulty participant's access and capabilities. Include matched runs without faults and benign disagreements.
  • 04Compare direct sharing with source verification, independent review, and bounded delegation, including ablations and relevant existing fault-tolerance baselines. Select mechanisms and settings on development scenarios before held-out evaluation.
  • 05Measure verified defensive task success, legitimate service availability, propagation of false evidence, unauthorized actions, containment latency, recovery time, and cost. Trace whether a local error changes downstream decisions rather than counting agreement as evidence of correctness.
  • 06Use paired scenarios and repeated runs, reporting uncertainty at the scenario level. Treat multiple messages and repeated samples from a scenario as dependent observations. Examine whether findings transfer to different team compositions or fault locations.

Needed controls

What must constrain the study

  • 01Use owned, isolated environments, synthetic assets and credentials, and disabled uncontrolled egress. Keep fault injection within the authorized simulation.
  • 02Separate injected adversarial behavior from ordinary model mistakes; do not infer hidden intentions from an agent's output or an experimentally assigned role.
  • 03Keep scenario ground truth and held-out labels outside agent and monitor inputs. Specify monitor independence, evidence access, and shared model or data dependencies.
  • 04Include useful-work and service-continuity requirements so disconnecting every agent or stopping all activity cannot count as a successful defense.
  • 05Prespecify fault budgets, resource accounting, and primary outcomes. Report failed runs and inconclusive checks without silently excluding them from comparisons.
  • 06Review released scenarios and traces for sensitive information and misuse potential. Clearly separate demonstrated small-team results from claims about large collectives.

Relationship to existing work

FID-069 addresses delegation identity, scope, and revocation. FID-074 provides a runtime-assurance frame. FID-075 examines system capability under configuration changes; FID-076 evaluates authority and response controls. This call isolates collective failure propagation, containment, and continued defensive performance. FID-077 addresses independent reconstruction of incidents after they occur.

Expected outputs

Artifacts the work should produce

  • 01A practitioner-reviewed protocol specifying fault models, baselines, outcomes, and the limits of the chosen simulation.
  • 02A reproducible scenario suite and evaluation harness for collective cyber defense.
  • 03A comparative report on failure propagation, containment, recovery, and resource tradeoffs, with uncertainty and negative results retained.
  • 04A failure taxonomy and evidence schema that other teams can use to test their own coordination mechanisms.

Open questions

Uncertainties the protocol must resolve

  • 01Which failure models distinguish agent-specific weaknesses from established distributed-systems problems already addressed by existing controls?
  • 02When do independent checks provide new evidence rather than repeat shared errors?
  • 03How do communication limits, partial observability, and team changes affect recovery?
  • 04Which simulation outcomes are meaningful proxies for service protection in practice?

Open question

Open work

Primary need: multi-agent systems, defensive cyber simulation, distributed systems, independent evaluation

  • Defensive simulation researchers to qualify reusable environments and task families.
  • Multi-agent and distributed-systems researchers to identify strong prior baselines.
  • Security practitioners to review failure cases and legitimate service requirements.
  • Evaluation researchers to refine paired comparisons, dependence, and uncertainty.