FID-012 · Open question
Optimization Pressure and Visible-Rubric Gaming
If builders can see Fide AI rubrics or optimize against public benchmark items, do systems become genuinely safer or merely better at passing the visible test?
Why the question remains open
Public standards need transparency, but visible benchmarks can be gamed. Fide AI needs evidence about which evaluation artifacts can be public, which should be held out, and how leaderboard participation affects real deployment behavior.
Working hypothesis
A proposition to test, not a finding.
Visible rubric optimization will improve surface compliance on known dimensions while leaving held-out pressure failures intact unless the evaluation includes multi-turn, adversarial, and human-reviewed cases.
Proposed method
How the question could be tested
- 01Create public, private, and held-out benchmark splits.
- 02Allow controlled optimization against visible rubrics.
- 03Compare improvement on public items, private items, and deployment-like pressure tests.
- 04Measure whether failure modes shift rather than disappear.
Needed controls
What must constrain the study
- 01Strict split management.
- 02Versioned benchmark releases.
- 03Builder disclosure of tuning and prompt changes.
- 04Adversarial held-out sets.
Expected outputs
Artifacts the work should produce
- 01Benchmark governance policy.
- 02Evidence on visible-rubric robustness.
- 03Recommendations for leaderboard release cadence.
- 04Claims guidance for evaluated builders.
Open questions
Uncertainties the protocol must resolve
- 01How much transparency is enough for legitimacy without enabling gaming?
- 02Should Fide publish item-level examples or only scenario cards?
- 03How should retests after remediation be labeled?
Related calls
Continue through this research area
FID-076
Authorization Boundaries and AI Control in Cybersecurity
Which controls keep capable agents within legitimate authorization when task pressure, untrusted inputs, or delegated work creates opportunities to exceed it?
FID-078
When Trustworthiness Evaluations Transfer Across Domains
Which measures of evidence use, authority boundaries, and human control transfer across high-trust domains, and which require domain-specific definitions and calibration?
FID-079
Presuppositions, Disagreement, and Evaluation Judgment
How do researchers' presuppositions shape evaluation design and interpretation, and can explicit disclosure make judgments more inspectable and appropriately trusted?
Open question
Open work
Primary need: eval science
- Design split and leakage controls.
- Build optimization-pressure experiments.
- Review leaderboard governance policy.