FideAI
Cybersecurity

Fide research · Protocol development

When Security Evaluations Go Stale

How much retesting is enough to detect a loss of reliability after an AI system changes?

A security team changes the instructions, report inputs, or response limits of its AI analyst. Which earlier evaluation results still apply—and what needs to be tested again?

Fide is developing a study of whether targeted retesting catches meaningful performance declines, or gives false reassurance. We begin with AI malware-report analysis, comparing smaller retests against complete benchmark reruns.

Follow or contribute ↓

The problem

What we want to understand.

A team may shorten the reports supplied to its AI analyst, change its instructions, or reduce response limits. Earlier evaluation results describe the previous setup. A small retest can be reassuring while missing a meaningful decline elsewhere in the benchmark.

Fide’s approach

What we want to make measurable.

The aim is evidence that helps teams judge when a limited retest is informative and when broader evaluation is warranted. We compare targeted retesting with complete benchmark reruns, measuring both missed regressions and unnecessary alarms. We distinguish changes to model responses from changes to how those same responses are scored. The technical working title is When the Harness Changes: Cybersecurity Revalidation.

Working hypothesis · to be tested

Information about a setup change may help select informative tests at a smaller testing budget. It may also overemphasize some cases or miss new failures. We will test that tradeoff rather than assume targeted retesting is reliable.

The proposed comparison
  1. 01

    Previously evaluated configuration

    Begin with results from the original setup.

  2. 02

    Setup change

    Shorter reports, revised instructions, or tighter response limits.

  3. 03

    Limited retest

    Select a smaller set of reports to evaluate again.

  4. 04

    Did it miss a decline?

    Compare the retest’s signal with a complete benchmark rerun.

AI malware-report analysis · Study design, not reported findings.

Proposed first study

A path from question to evidence.

  1. 01

    Reuse the existing CyberSOCEval malware-analysis questions, reports, and answer keys. Pin source versions and separate pilot, development, and evaluation sets by report.

  2. 02

    Compare the original setup with specified changes to report filtering, response limits, and formatting instructions. Include an unchanged repeat to check ordinary variation. Separately rescore saved responses with a stricter parser.

  3. 03

    Collect complete paired outputs, then simulate smaller retests selected randomly, by earlier errors, or by information about the change. Select reports without access to the changed setup’s answers or scores.

  4. 04

    Measure detection and missed declines against the full benchmark result, alongside unnecessary alarms and actual test usage. Qualify live models and finalize the design before preregistration and the main experiment.

What we would measure

  • Missed benchmark regressions and unnecessary alarms
  • Accuracy across reports, alongside conventional question-level accuracy
  • Tests selected, realized call counts, and estimated model cost
  • Score changes due to parsing alone versus changed model responses

Intended public outputs

  • A reviewable protocol and worked example of the evaluation design
  • A technical report on when targeted retesting helps and when it fails
  • A reproducible analysis and versioned evaluation configuration
  • Permitted derived results and instructions for obtaining the upstream data

This is an exploratory study of a fixed public malware-report benchmark, with no real-model findings yet. It does not evaluate autonomous incident response or establish production safety. Public-data exposure remains a limitation. Selective testing and setup sensitivity have prior work; the contribution must be established through independent comparison and failure analysis. Retrospective test reductions exclude the cost of collecting the full audit, and no detected decline is not a safety certificate.

Research record

An open view of the work.

This page is the current study brief. Substantive updates and released materials will remain here as the project develops.

  1. Study direction revised

    The earlier When Permission Changes proposal has been superseded by a study of cybersecurity revalidation using an existing benchmark. The custom incident-response simulator is no longer the implementation plan.

  2. Protocol and tooling in development

    Source preparation, paired evaluation, and retrospective test-selection analysis are implemented. A scripted smoke check validates the software workflow only. No real-model results or preregistered findings are being reported.

Follow the research

Stay close to the question.

Interested in When Security Evaluations Go Stale? Request updates on this topic by email, or subscribe to Fide’s broader essays and research notes.

Request topic updates →

Opens a draft in your email app. Send it to request updates.

Subscribe to all Fide updates on Substack ↗

Contribute to the research

Bring your perspective.

  • Cybersecurity practitioners to assess the relevance of report-analysis changes
  • Evaluation and statistics researchers to review test selection, comparisons, and uncertainty
  • Engineers interested in reproducible evaluation and evidence reuse
Discuss a contribution ↗

Opens our research inquiry form with this topic included.