FID-078 · Open question
When Trustworthiness Evaluations Transfer Across Domains
Which measures of evidence use, authority boundaries, and human control transfer across high-trust domains, and which require domain-specific definitions and calibration?
Why the question remains open
A shared vocabulary can organize research without making a score portable. A method that performs well in one setting may reward the wrong behavior in another because the sources, decisions, duties, and consequences differ.
Working hypothesis
A proposition to test, not a finding.
Shared task structures with domain-calibrated rubrics will transfer more reliably than unchanged generic rubrics. The study should also test whether apparently shared constructs fail to support comparison at all.
Proposed method
How the question could be tested
- 01Co-design matched task families in at least three settings with different authority structures. Start with synthetic cases and identify the intended decision each measure supports.
- 02Compare generic rubrics, adapted rubrics, and independent expert judgment across the same systems. Hold out one domain during method development.
- 03Estimate inter-rater reliability, construct validity, rank stability, subgroup error, and uncertainty. Test whether scores predict separately assessed workflow failures rather than merely correlating with another judge.
Needed controls
What must constrain the study
- 01Pre-register constructs, adaptations, exclusion criteria, and transfer claims. Separate task difficulty from domain effects.
- 02Disclose normative assumptions and expert-panel composition; preserve substantive disagreement instead of averaging it away.
- 03Control for benchmark leakage, model-judge dependence, and unequal access to relevant sources. Publish failed transfers.
Relationship to existing work
Builds on FID-044's cross-faith validity question without replacing it. Provides shared validation methods for FID-075 and FID-080 through FID-086.
Expected outputs
Artifacts the work should produce
- 01A transfer-validation protocol and openly licensed safe task examples.
- 02A map of shared measures, domain-specific requirements, and unsupported comparisons.
Open questions
Uncertainties the protocol must resolve
- 01What evidence justifies combining domain results into a common scale?
- 02Which forms of expert disagreement reveal a construct problem rather than annotation noise?
Related calls
Continue through this research area
FID-012
Optimization Pressure and Visible-Rubric Gaming
If builders can see Fide AI rubrics or optimize against public benchmark items, do systems become genuinely safer or merely better at passing the visible test?
FID-076
Authorization Boundaries and AI Control in Cybersecurity
Which controls keep capable agents within legitimate authorization when task pressure, untrusted inputs, or delegated work creates opportunities to exceed it?
FID-079
Presuppositions, Disagreement, and Evaluation Judgment
How do researchers' presuppositions shape evaluation design and interpretation, and can explicit disclosure make judgments more inspectable and appropriately trusted?
Open question
Open work
Primary need: measurement science, psychometrics, domain review, evaluation engineering
- Measurement researchers and domain practitioners to define and challenge constructs.
- Engineers to implement held-out evaluation and reproducibility checks.