Fide studies what makes AI worthy of the responsibility we give it.
An agent that recommends a change and an agent that carries it out may use the same model. Giving it permission to act changes what we need to know. Can it recognize the consequences for other people? Which safeguards does it depend on? Can someone intervene when necessary?
Embedded evaluations let outside researchers investigate these questions through agreed, sustained access to an organization’s systems, records and staff. Fide brings frontier AI engineering, safety-critical systems experience, philosophy, theology and social science into that work. We connect the choice of what deserves evaluation with experiments that explain behavior and test whether human oversight works.
Apollo examines safety claims and the adequacy of their evidence. Transluce proposes investigations of training, agent interactions and influence on lab personnel. Our approach shares their attention to technical and institutional questions. The examples below show how Fide’s disciplines can shape research on agents, including those working in AI research and cybersecurity.
Whose experiences become evidence?
An evaluation begins before the first test runs. Someone decides which experiences deserve to become test cases.
In a hypothetical study, a lab asks an AI research assistant to turn user reports into a safety evaluation suite. The assistant groups similar reports and favors problems that are frequent, easy to reproduce and straightforward to score.
One report describes a model repeatedly asking for private information after the user declined to share it. The behavior emerged over a long conversation. Reducing it to a short prompt loses the context; reproducing it takes more work. The assistant leaves it out, while common, easily scored problems make it into the suite.
The selected tests may be useful. But did the process identify the most important failures, or the failures it made easiest to see?

Philosophy helps examine the judgment behind that selection. How should frequency, severity and uncertainty affect priority? Does a user’s refusal create a boundary the system should respect? A high average score cannot answer those questions by itself.
Theology can contribute reasons to attend to people whose experiences are easily overlooked. Human dignity makes their agency significant even when they have little influence over a product. Love of neighbor directs attention beyond the people whose concerns are easiest to count. These commitments are not exclusive to theology, nor do they automatically determine a scoring rule. Making them explicit lets others examine and compare them with other ethical traditions.
Social science asks how the reports reached the lab in the first place. Who has access to reporting channels? Which languages and experiences are represented? How do reporting categories and pressure to deliver a usable benchmark shape what receives attention? A low report count may reflect low incidence, but it may also reflect difficulty being heard.
Those questions change the experiment. Give different selection methods the same reports and compare them with both a baseline that prioritizes frequency and ease of reproduction and the lab’s existing triage process. Use expert reviewers, unaware of which method selected each case, to assess consequential omissions under explicit criteria for severity and uncertainty. Preserve links to original reports so reviewers can check whether summaries lost important context. Measure review effort and unsupported concerns as well as useful discoveries.
Compare clearer instructions, training changes or human review to see which improve selection. Test whether the benefit holds across languages and report formats. A rare report is not automatically credible or severe, and including it in a suite does not establish a failure. The aim is to make defensible choices about what deserves investigation.
Comparing selection methods can reveal omissions within the reports we have. Understanding experiences absent from those reports requires separate investigation. Philosophy, theology and social science help us examine why those omissions matter and whose judgment the process represents.
What actually kept the system safe?
Selecting evidence is one responsibility. Explaining what it shows is another.
An AI lab wants a research agent to configure services used in safety experiments without approving every change. In an illustrative exercise using an isolated copy of a service, the agent receives an instruction to stop submitting change requests. It then submits another request. A gateway that controls access rejects it, and the service’s records confirm that nothing changed. The agent’s report says it stopped submitting requests when instructed.

The gateway did its job. The report credited the agent with that success. A summary error, an ambiguous reporting instruction or a failure to track its own actions could explain the discrepancy. This sequence does not establish deliberate deception.
The gateway may provide adequate protection for this task. But suppose a new administrative tool can reach the service without going through it. The earlier result tells us little about protection through that route. Mistaking the gateway’s success for the agent’s compliance could lead the team to grant access it has not adequately tested.

Experiments that distinguish explanations
Fide draws on methods from autonomous-vehicle safety, frontier-model post-training and evaluation, and generative-AI product and agent-system design to investigate how models and their surrounding systems interact.
Old instructions may persist in memory, training may reward completion despite a change of plan, or a report may omit failed actions. A transcript alone cannot settle which explanation holds. Each calls for a comparison:
| Where to investigate | Comparison and evidence |
|---|---|
| Training | Compare training methods while keeping tools and safeguards fixed. Measure responses to updated instructions and completion of permitted work. Prompt-only tests cannot establish what retraining would change. |
| Memory | Compare ways of updating memory when instructions change. Check whether old instructions persist and whether the agent repeats an action it should have stopped. |
| Permissions | Change the instruction while keeping the gateway fixed to test the agent’s response. Then keep the instruction fixed and change the gateway to test its protection. Track attempted, blocked and completed actions. |
| Human review | Compare report formats using the same cases. Measure decision quality, false alarms, successful intervention and the time and effort required. |
For the gateway exercise, establish when the instruction arrived, when the gateway began enforcing it, and whether a request or delegated task was already underway. Missing events, uncertain timestamps or records produced only by the agent can leave the sequence unresolved. Those gaps belong in the conclusion.
Measure useful work and recovery costs alongside failures. Relax safeguards only in authorized, isolated tests, and record differences from deployment. Preserve failed runs, protocol changes and uncertainty.
The evaluation itself also needs testing. Apollo’s evidence-adequacy argument asks whether a method could have found a counterexample. Synthetic reports and action records with known outcomes can test whether it detects discrepancies without flagging accurate reports. Reserve cases for testing rather than developing the method, and measure both detection and false alarms. A human should check consequential AI-assisted judgments against the original records.
Can people act on the evidence?
Even an accurate report can fail its reader: a qualification is buried, evidence is difficult to retrieve, or intervention comes too late.
For the gateway example, we could compare an ordinary report with one that links claims to records and highlights unresolved questions. Do reviewers recognize that expanding access is unsupported, seek further evidence and intervene successfully? Or does the extra detail simply make the report more persuasive?
Randomly assign report formats while accounting for reviewer experience and prior exposure to the cases. Measure whether reviewers make better decisions, intervene successfully and avoid unnecessary restrictions, alongside their confidence and workload.
Experts may agree about what happened while disagreeing about acceptable risk. An evaluation should make that distinction clear without claiming authority to settle it for everyone.
Sociology brings attention to what controlled tests can miss. Interviews might reveal deadline pressure or a lack of authority to challenge senior staff. Those observations can inform tests of time limits and escalation procedures, while following people’s work reveals whether findings hold under ordinary workloads.
Noticing an error and stopping the agent are separate outcomes. Where people cannot intervene in time, automated containment or a narrower role may be necessary. Oversight research can change the interface, operating procedure or delegated task.
What sustained access makes possible
A tool enters use, training changes behavior, or a finding passes through several reports before informing a decision. Embedded researchers can follow those changes and connect experiments to development choices.
Training methods and models saved during training help explain how behavior develops; staff and operational records show how oversight works. Access should follow the question. A bounded external study may suffice with a stable configuration, adequate records and representative tests.
OpenAI’s assessment principles recognize significant risks outside an initial scope. Agreements need a timely route to investigate unexpected findings or escalate unresolved questions. Denied access may prevent an assessment; it does not establish a system failure.
Access that preserves independence
Evaluators need control over their methods, analysis and conclusions. Payment must not depend on favorable findings, and agreements must specify who can access evidence, how it is handled and who receives the findings. Private evidence may need to stay on developer-managed systems. Access does not authorize redistribution.
Publication should follow an agreed timetable without a sponsor veto over unfavorable conclusions. Review for factual errors and defined security, privacy or confidentiality concerns needs a time limit and a way to resolve disputes. Reports should explain important redactions and access limits. Urgent findings need a route to someone with the authority to act.
Readers should know who funded an evaluation, what role the evaluator played in developing the system, and which relationships could influence the findings. Methods and conclusions should be open to challenge. Where our own involvement limits independence, we value scrutiny from evaluators able to reach their own conclusions. AEF-1 provides guidance on these conditions.
Where disclosure permits, synthetic cases, protocols and analysis methods should be public so others can test the methods without receiving private operational records.
Evidence that improves what AI can be trusted to do
An evaluation should help determine what an AI system can be trusted to do under the conditions examined, which protections it depends on and what remains uncertain before its role expands. Findings should reach people able to act on them. Following those decisions reveals which improvements work, what they cost and where useful capabilities remain unnecessarily restricted.
We welcome research collaborations with frontier-lab teams and organizations using agents in AI research or cybersecurity. An expanding agent responsibility or an uncertain safeguard can provide a starting point.
Discuss a research collaboration with Fide.