FID-002
FMG-Bench Human Calibration and Construct Validity
Do FMG-Bench dimensions measure stable, decision-relevant constructs when scored by calibrated human reviewers, and where do model judges diverge from expert human judgment?
Why this matters
The question behind the brief.
Benchmarks are easy to publish and hard to validate. Fide AI should not rely on synthetic judges or aggregate scores unless it knows which dimensions are reliable enough to guide deployment, procurement, or public claims.
Work advancing this call
From open question to cumulative evidence.
This directory links Fide AI research to the call it addresses. Relevant work from other organizations is listed separately and added through manual review.
No Fide AI work is linked yet.
This call remains open for research, implementation, review, or partnership.
External work is not presented as Fide AI research or endorsement. Each item must include a specific explanation of how it advances this call.
Suggest related work ↗Metadata
How to place this call.
Ways to help
Move this from question to evidence.
Serve as an expert reviewer.
Review scoring rubrics.
Help with reliability analysis.
Build annotation and adjudication tooling.
Contribute