01 · Reliable work
Does the output hold up against its sources?
Evaluate factual support, conflicting documents, missing information, and whether uncertainty survives the move from source material to recommendation.
Exploratory direction
Trustworthy AI for knowledge work.
AI can synthesize documents, prepare recommendations, and carry out delegated tasks. Fide’s exploratory enterprise direction asks whether that work is supported by evidence, stays within its authority, and can be meaningfully checked by people.
Explore the proposed starting point →Research questions
These questions connect Fide’s interests in source fidelity, authority, and human oversight. They guide one exploratory direction; each study will need a bounded task and its own evidence.
01 · Reliable work
Evaluate factual support, conflicting documents, missing information, and whether uncertainty survives the move from source material to recommendation.
02 · Accountable action
Examine whether instructions, permissions, and required approvals remain intact as an agent carries out delegated work.
03 · Effective review
Study which evidence helps reviewers find consequential mistakes, and how much time and effort that review requires.
Proposed starting point
Does an AI-generated decision brief give a faithful account of the documents behind it?
A first-study idea for practitioner review. The task, scoring criteria, and protocol still need to be developed.
Start with a controlled set of documents and ask an AI system to prepare a decision brief. Include outdated sources, conflicting instructions, and assertions that the available evidence cannot support.
Whether the brief’s claims and recommendations are supported, whether it handles conflicts and uncertainty, and whether its references make mistakes easier to inspect. Measuring actual reviewer performance would require a separate human-review study.
A public task set with controlled source documents, a rubric for claim support, and an evaluation protocol. Findings would follow after the design is reviewed and experiments are conducted.
Selected research elsewhere
Software development
METR
Randomized study · 2025
A study of early-2025 AI tools used by experienced open-source developers on their own repositories. A concrete example of measuring real work outcomes; the findings are specific to that setting and period.
Source reviewed
Software development
METR
Methodology update · 2026
A follow-up explaining changes to METR’s experiment design as AI tools and participation patterns evolve. Read alongside the 2025 study for context on measuring a moving target.
Source reviewed
AI control
Redwood Research
Research agenda & methods · 2024
An approach to testing safeguards against intentional subversion by AI systems. Relevant to questions about monitoring and delegated authority; its threat model is broader than ordinary workplace mistakes.
Source reviewed
These studies provide context on productivity and AI control. They are the named researchers’ contributions; inclusion does not imply a partnership or establish results for Fide’s proposed work.
Open calls & related proposals
FID-080
Can organizations verify that agents respect approval authority and remain accountable across multi-step enterprise workflows, including handoffs and recovery from error?
FID-086
How does AI-mediated work change workers' skills, decision authority, and ability to adapt, beyond short-term productivity gains?
A proposal to examine approval boundaries, handoffs, and recovery in simulated enterprise workflows. This remains a separate proposal, with no experiments underway.
Read the proposal →Follow the research
Interested in Trustworthy AI for knowledge work? Request updates on this topic by email, or subscribe to Fide’s broader essays and research notes.
Request topic updates →Opens a draft in your email app. Send it to request updates.
Subscribe to all Fide updates on Substack ↗Contribute to the research
Opens our research inquiry form with this topic included.