FideAI
← Cybersecurity research agenda

Cybersecurity · Effective defense under human control

Controlled autonomous response

Can an AI defender stop an attack without disrupting the systems it is meant to protect?

Help shape this research →

Why this question matters

What we want to understand.

Responding to an incident may require isolating a host, changing access, or interrupting a service. Those same actions can harm legitimate users if the evidence is wrong or authority is unclear. The practical challenge is to enable timely autonomous action while preserving essential services and meaningful human control.

The decision it could inform

Evidence someone can act on.

Help operators determine which actions an AI defender can take independently, which need approval, and how to preserve effective response when evidence or permissions change.

Grounded in consequential research

The programs informing this priority.

DARPA · Program reference

CASTLE: AI agents for network defense

CASTLE develops AI agents and environments for network defense. Fide’s proposed contribution is to measure the operational tradeoff between containing a threat and preserving legitimate activity.

DARPA · Program reference

ANSR: evidence for trustworthy autonomy

ANSR connects trustworthy autonomy with robustness and assurance evidence, and identifies the mission cost of frequent fallback. It motivates evaluating useful action alongside control and recovery; Fide is not claiming to develop ANSR’s neuro-symbolic architectures.

Fide’s proposed contribution is our interpretation of these research needs. References do not imply funding, partnership, endorsement, or an open solicitation.

Proposed first study

A concrete starting point.

In a contained defensive environment, compare fixed approval rules with an evidence-sensitive response policy. Begin with one consequential decision: whether to isolate a suspected host when doing so could interrupt a legitimate service.

  1. 01

    With security practitioners, specify one incident family, legitimate service requirements, permitted actions, and ambiguous benign cases. Use an environment that records the effects of actions, not only the agent’s explanations.

  2. 02

    Compare always-request-approval and fixed-rule baselines with a policy that considers evidence, action reversibility, and current authority. Choose thresholds on development scenarios; vary evidence quality and simulated approval delay on held-out cases.

  3. 03

    Measure containment and service disruption together with unauthorized actions, escalation load, and recovery. Include authority changes during a run and check whether the agent stops or revises actions when permission is withdrawn.

What we would measure

Threat containment, time to containment, legitimate service disruption, authority violations, unnecessary interventions, fraction of cases escalated, and recovery cost. Simulated approval delay would not measure real analyst performance.

What a useful contribution could be

  • A practitioner-reviewed evaluation protocol for autonomous response decisions
  • A reproducible comparison of response policies and their containment, disruption, and oversight tradeoffs

Scope and readiness

Proposed research seeking domain collaborators; no response experiments or findings yet. Simulated incidents and approvals cannot establish safe production deployment or human review quality. The first study would cover one response decision rather than a complete SOC.

Follow the research

Stay close to the question.

Interested in Controlled autonomous response? Request updates on this topic by email, or subscribe to Fide’s broader essays and research notes.

Request topic updates →

Opens a draft in your email app. Send it to request updates.

Subscribe to all Fide updates on Substack ↗

Contribute to the research

Bring your perspective.

  • SOC and incident-response practitioners, researchers in runtime assurance, and engineers building defensive agents.
  • Help qualify the data, challenge the proposed comparison, or refine the practical decision this research should inform.
Discuss a contribution ↗

Opens our research inquiry form with this topic included.

Part of a wider agenda

More questions in trustworthy cyber defense.

Study in development

When Security Evaluations Go Stale →

Our current revalidation study examines targeted retesting for AI malware-report analysis. Explore its protocol, current stage, and collaboration needs.