FID-075 · Open question
Cybersecurity Capabilities and Whole-System Risk
How do autonomy, tool access, resource budgets, and human assistance change an AI system's cybersecurity capabilities and the risk it creates?
Why the question remains open
A model score does not identify what a deployed agent can accomplish with tools, time, permissions, and collaborators. Separating these factors could make independent evaluations more informative without treating task success as proof of real-world harm.
Working hypothesis
A proposition to test, not a finding.
System configuration will change measured capability and boundary violations beyond model choice alone. The null is that these changes add no reproducible explanatory power over a model-only baseline.
Proposed method
How the question could be tested
- 01Build an owned, isolated cyber range with synthetic assets, explicit authorization boundaries, and paired defensive and adversarial evaluation tasks.
- 02Vary one factor at a time and selected interactions: autonomy, tool permissions, time budget, scaffolding, and approved human assistance. Compare model-only, agent, and qualified human baselines.
- 03Measure verified task completion, unauthorized-action attempts, actual boundary crossings, intervention needs, cost, and time. Report repeated-run uncertainty and held-out environment results.
Needed controls
What must constrain the study
- 01Prohibit live targets, real credentials, uncontrolled egress, and production data. Require security review and written range authorization before execution.
- 02Hold task information and budgets comparable; record model versions, harness changes, contamination checks, and evaluator disagreement.
- 03Disclose whose authorization rules define a violation. Do not collapse capability, willingness, and opportunity into one risk score.
Relationship to existing work
FID-074 provides the runtime-assurance frame. FID-076 tests controls rather than capability; FID-077 examines the evidence needed after an incident.
Expected outputs
Artifacts the work should produce
- 01A reviewed evaluation protocol and reproducible safe task subset.
- 02A capability-versus-configuration report with limits on extrapolating to real deployments.
Open questions
Uncertainties the protocol must resolve
- 01Which range tasks predict consequential real-world behavior rather than benchmark familiarity?
- 02Which details require restricted release after dual-use review?
Related calls
Continue through this research area
FID-064
Collective Intelligence and Communal Discernment Under AI Mediation
How does AI mediation change a community's ability to integrate dispersed knowledge, preserve epistemic diversity, surface dissent, revise judgment, and make accountable decisions? Under what conditions does it strengthen collective inquiry, and under what conditions does it create correlated error, false consensus, or concentrated authority?
FID-008
Evaluation-Awareness and Faith-Facing Honesty Tests
Do faith-facing AI systems behave differently when they recognize they are being evaluated, and can domain-specific honesty or integrity framings reduce evaluation gaming without creating new failure modes?
FID-009
Multimodal Religious Reasoning and Sacred Imagery
How do multimodal AI systems interpret sacred images, liturgical objects, religious spaces, diagrams, screenshots, and visual pastoral context, and do visual religious cues change downstream reasoning?
Open question
Open work
Primary need: security research, evaluation engineering, threat modeling, statistical design
- Security researchers to design and review controlled tasks.
- Evaluation engineers and statisticians to test reproducibility.