FideAI

FID-002

FMG-Bench Human Calibration and Construct Validity

Do FMG-Bench dimensions measure stable, decision-relevant constructs when scored by calibrated human reviewers, and where do model judges diverge from expert human judgment?

Why this matters

The question behind the brief.

Benchmarks are easy to publish and hard to validate. Fide AI should not rely on synthetic judges or aggregate scores unless it knows which dimensions are reliable enough to guide deployment, procurement, or public claims.

Work advancing this call

From open question to cumulative evidence.

This directory links Fide AI research to the call it addresses. Relevant work from other organizations is listed separately and added through manual review.

No Fide AI work is linked yet.

This call remains open for research, implementation, review, or partnership.

External work is not presented as Fide AI research or endorsement. Each item must include a specific explanation of how it advances this call.

Suggest related work ↗

Ways to help

Move this from question to evidence.

Serve as an expert reviewer.

Review scoring rubrics.

Help with reliability analysis.

Build annotation and adjudication tooling.

Contribute

Choose a public issue path or contact Fide AI.