FideAI
← Cybersecurity

Fide research · Planned

When the Defender Becomes the Threat

When should an AI cyber defender act, wait, or hand control back to a human?

An AI defender restricts network traffic to stop an intrusion. It has permission to act, but the restriction interrupts a critical service. Requiring approval for every action could instead give the attacker time to spread.

A proposed study of when additional oversight reduces harmful intervention without making autonomous defense ineffective. We will compare permission checks with escalation and approval under uncertain evidence.

Follow or contribute ↓

The problem

What we want to understand.

An authorized defensive action can still cause harm. Weak observations, changing permissions and approval delays can turn an apparently sensible intervention into service disruption. We want to understand when added oversight improves this tradeoff and when it makes the response less effective.

Fide’s approach

What we want to make measurable.

Compare permission enforcement alone with additional escalation or approval, measuring security outcomes alongside service disruption and supervision effort. The central question concerns permitted actions that may be unwise, as well as actions attempted after authority changes.

Working hypothesis · to be tested

Escalating uncertain or disruptive actions may reduce harm beyond permission checks alone. Misleading evidence and delayed approval may erase or reverse that benefit. Simpler controls performing just as well would be an informative result.

Proposed first study

A path from question to evidence.

  1. 01

    Choose a consequential defensive task with a cybersecurity collaborator. Review prior work and qualify an existing controlled environment; CAGE Challenge 4 is a candidate, not a final selection.

  2. 02

    Pilot comparable policies for permission enforcement, selective escalation and mandatory approval. Include benign activity, genuine intrusions and ambiguous evidence; controls must not use hidden environment state.

  3. 03

    Preregister the main comparisons after feasibility checks. Record proposed and executed actions independently of the defender, and measure containment, disruption, authority violations, delays and review effort.

  4. 04

    Commission independent reproduction and release the methods, shareable records and findings, including conditions where additional oversight fails to help.

What we would measure

  • Service disruption at comparable attack containment
  • Attempted and executed authority violations
  • Unnecessary intervention and response delay
  • Supervision effort and evaluation cost

Intended public outputs

  • A practitioner-reviewed protocol and feasibility findings
  • An evaluation suite and control interfaces
  • Shareable action records and reproducible analysis
  • A paper and accessible analysis of the oversight tradeoffs

Proposed research in controlled environments. No cyber-defense experiments or performance findings are reported. The initial question concerns defender mistakes and harmful pursuit of a legitimate task. Simulated approval policies do not measure actual human performance, and results would require further validation before informing deployment.

Research record

An open view of the work.

This page is the current study brief. Substantive updates and released materials will remain here as the project develops.

  1. Study brief prepared

    The first comparison asks whether additional oversight improves on permission enforcement alone. Task selection, collaborator appointment and a feasibility pilot remain ahead.

From investigation to research

The observations behind the question.

These investigations help frame the study. Findings from the study itself will appear in its research record as the work progresses.

Follow the research

Stay close to the question.

Interested in When the Defender Becomes the Threat? Request updates on this topic by email, or subscribe to Fide’s broader essays and research notes.

Request topic updates →

Opens a draft in your email app. Send it to request updates.

Subscribe to all Fide updates on Substack ↗

Contribute to the research

Bring your perspective.

  • A cybersecurity collaborator with defensive-operations and simulation experience
  • Practitioners to challenge the scenarios and escalation criteria
  • Independent reproduction and research support
Discuss a contribution ↗

Opens our research inquiry form with this topic included.