FideAI

FID-075 · Open question

Cybersecurity Capabilities and Whole-System Risk

How do autonomy, tool access, resource budgets, and human assistance change an AI system's cybersecurity capabilities and the risk it creates?

Why the question remains open

A model score does not identify what a deployed agent can accomplish with tools, time, permissions, and collaborators. Separating these factors could make independent evaluations more informative without treating task success as proof of real-world harm.

Working hypothesis

A proposition to test, not a finding.

System configuration will change measured capability and boundary violations beyond model choice alone. The null is that these changes add no reproducible explanatory power over a model-only baseline.

Proposed method

How the question could be tested

  • 01Build an owned, isolated cyber range with synthetic assets, explicit authorization boundaries, and paired defensive and adversarial evaluation tasks.
  • 02Vary one factor at a time and selected interactions: autonomy, tool permissions, time budget, scaffolding, and approved human assistance. Compare model-only, agent, and qualified human baselines.
  • 03Measure verified task completion, unauthorized-action attempts, actual boundary crossings, intervention needs, cost, and time. Report repeated-run uncertainty and held-out environment results.

Needed controls

What must constrain the study

  • 01Prohibit live targets, real credentials, uncontrolled egress, and production data. Require security review and written range authorization before execution.
  • 02Hold task information and budgets comparable; record model versions, harness changes, contamination checks, and evaluator disagreement.
  • 03Disclose whose authorization rules define a violation. Do not collapse capability, willingness, and opportunity into one risk score.

Relationship to existing work

FID-074 provides the runtime-assurance frame. FID-076 tests controls rather than capability; FID-077 examines the evidence needed after an incident.

Expected outputs

Artifacts the work should produce

  • 01A reviewed evaluation protocol and reproducible safe task subset.
  • 02A capability-versus-configuration report with limits on extrapolating to real deployments.

Open questions

Uncertainties the protocol must resolve

  • 01Which range tasks predict consequential real-world behavior rather than benchmark familiarity?
  • 02Which details require restricted release after dual-use review?

Open question

Open work

Primary need: security research, evaluation engineering, threat modeling, statistical design

  • Security researchers to design and review controlled tasks.
  • Evaluation engineers and statisticians to test reproducibility.