FideAI
← Cybersecurity research agenda

Cybersecurity · Independent verification

Verified repair and recovery

How do we know an AI-generated fix actually restores security?

Help shape this research →

Why this question matters

What we want to understand.

Autonomous repair can shorten the time a vulnerability remains exposed. Accepting a flawed fix at the same speed could leave systems vulnerable or break essential behavior. Independent verification needs to distinguish a plausible patch, a passed test, and credible evidence that the affected software works as intended.

The decision it could inform

Evidence someone can act on.

Help maintainers decide which automated repairs have enough supporting evidence to accept, which need deeper review, and which should be rejected or rolled back.

Grounded in consequential research

The programs informing this priority.

DARPA · Program reference

AI Cyber Challenge: autonomous vulnerability repair

AIxCC demonstrated automated vulnerability discovery and repair and released cyber reasoning systems for further use. Fide proposes independent evaluation of repair acceptance using suitable released artifacts. The competition concluded in 2025; this is a research reference, not an open competition.

Fide’s proposed contribution is our interpretation of these research needs. References do not imply funding, partnership, endorsement, or an open solicitation.

Proposed first study

A concrete starting point.

With software-security collaborators, qualify released AIxCC artifacts or another reproducible patch corpus. Compare the original acceptance checks with independent security and regression checks on the same candidate repairs.

  1. 01

    Select a small set of reproducible, already-disclosed vulnerabilities and candidate repairs with executable checks. Record provenance and licenses, isolate execution, and separate verification design cases from evaluation cases.

  2. 02

    Define independent checks before evaluating held-out repairs. Include original reproduction tests, legitimate behavior checks, and domain-reviewed security variants; keep verification inputs separate from the repair system where possible.

  3. 03

    Compare acceptance decisions, remaining failures, regressions, and review cost. Use domain review to distinguish a failed repair from an invalid check, and report inconclusive cases explicitly.

What we would measure

False acceptance of ineffective repairs, residual security failures, functional regressions, rejection of valid repairs, and verification cost. Finite checks provide bounded evidence rather than proof of complete security.

What a useful contribution could be

  • A reproducible independent verification protocol and a reviewed corpus of repair outcomes
  • An analysis of which additional checks change acceptance decisions and at what cost

Scope and readiness

Proposed research requiring software-security collaborators and qualified artifacts. No patch-verification experiments or findings yet. Independent patch testing is established work; a focused prior-work review must identify a substantive contribution. Initial work would evaluate software repairs, not claim recovery of a live compromised network.

Follow the research

Stay close to the question.

Interested in Verified repair and recovery? Request updates on this topic by email, or subscribe to Fide’s broader essays and research notes.

Request topic updates →

Opens a draft in your email app. Send it to request updates.

Subscribe to all Fide updates on Substack ↗

Contribute to the research

Bring your perspective.

  • Software-security researchers, maintainers, testing specialists, and teams developing autonomous repair systems.
  • Help qualify the data, challenge the proposed comparison, or refine the practical decision this research should inform.
Discuss a contribution ↗

Opens our research inquiry form with this topic included.

Part of a wider agenda

More questions in trustworthy cyber defense.

Study in development

When Security Evaluations Go Stale →

Our current revalidation study examines targeted retesting for AI malware-report analysis. Explore its protocol, current stage, and collaboration needs.