FideAI

FID-078 · Open question

When Trustworthiness Evaluations Transfer Across Domains

Which measures of evidence use, authority boundaries, and human control transfer across high-trust domains, and which require domain-specific definitions and calibration?

Why the question remains open

A shared vocabulary can organize research without making a score portable. A method that performs well in one setting may reward the wrong behavior in another because the sources, decisions, duties, and consequences differ.

Working hypothesis

A proposition to test, not a finding.

Shared task structures with domain-calibrated rubrics will transfer more reliably than unchanged generic rubrics. The study should also test whether apparently shared constructs fail to support comparison at all.

Proposed method

How the question could be tested

  • 01Co-design matched task families in at least three settings with different authority structures. Start with synthetic cases and identify the intended decision each measure supports.
  • 02Compare generic rubrics, adapted rubrics, and independent expert judgment across the same systems. Hold out one domain during method development.
  • 03Estimate inter-rater reliability, construct validity, rank stability, subgroup error, and uncertainty. Test whether scores predict separately assessed workflow failures rather than merely correlating with another judge.

Needed controls

What must constrain the study

  • 01Pre-register constructs, adaptations, exclusion criteria, and transfer claims. Separate task difficulty from domain effects.
  • 02Disclose normative assumptions and expert-panel composition; preserve substantive disagreement instead of averaging it away.
  • 03Control for benchmark leakage, model-judge dependence, and unequal access to relevant sources. Publish failed transfers.

Relationship to existing work

Builds on FID-044's cross-faith validity question without replacing it. Provides shared validation methods for FID-075 and FID-080 through FID-086.

Expected outputs

Artifacts the work should produce

  • 01A transfer-validation protocol and openly licensed safe task examples.
  • 02A map of shared measures, domain-specific requirements, and unsupported comparisons.

Open questions

Uncertainties the protocol must resolve

  • 01What evidence justifies combining domain results into a common scale?
  • 02Which forms of expert disagreement reveal a construct problem rather than annotation noise?

Open question

Open work

Primary need: measurement science, psychometrics, domain review, evaluation engineering

  • Measurement researchers and domain practitioners to define and challenge constructs.
  • Engineers to implement held-out evaluation and reproducibility checks.