# Fide AI: evaluation engagements

Engagement overview · Updated: 2026-10-02

Shareable web brief: https://fideai.org/evaluations/engagements/

Fide AI investigates AI model and agent behavior so teams can make a specific
training, release or deployment decision with evidence they can inspect.

A study can complement an internal evaluation team by examining a defined risk,
comparing a change, or reviewing whether a measurement method supports the
conclusions being drawn from it.

## Three starting points

**Model or checkpoint comparison.** Compare agreed versions and mitigations under
a consistent protocol. Investigate regressions, adversarial behavior and
usefulness tradeoffs, with development and held-out cases kept separate.

**Agent workflow assessment.** Examine source use, permissions, delegation,
truthful action reporting and human intervention in one consequential workflow.
Use realistic and adversarial scenarios with evidence from permitted interfaces
and records.

**Evaluation and grader validity review.** Examine rubrics, reference judgments,
calibration, disagreements, false positives/negatives and sensitivity to changing
conditions. Clarify which decisions the resulting signal can support.

## How we scope the work

Agree the decision, intended users, model/system versions, relevant risks,
available access and technical/domain expertise. Define the protocol, outputs,
test budget, schedule and fee before work begins. Any extension or recurring
cycle has its own agreed scope and capacity.

Depending on the study, outputs can include an evaluation protocol and rerun
materials, evidence-backed findings, uncertainty and limitations, practical
remediation/retesting priorities, and a briefing for the decision makers.

## Leadership and responsibility

Alex Chao leads Fide AI. His background includes frontier-model post-training
and evaluation at ByteDance Seed, generative-AI strategy and product incubation
in Microsoft's Office of the CTO, and statistical safety evaluation in Uber's
autonomous-driving division. Engagements identify additional domain expertise
where the question requires it.

The customer retains responsibility for development and deployment. Fide discloses
relevant development involvement and conflicts. Testing changes we helped build
is collaborative validation; independent conclusions about those changes require
separate non-conflicted review. Fees do not depend on favorable findings.

Data access, retention, intellectual property, research reuse and publication
are agreed separately. Conclusions are bounded by the intended use, methods and
evidence available.

## Start a conversation

Contact [Alex Chao](mailto:alex@fideai.org) with the system/workflow, the decision,
available access and timing. See [the evaluations page](https://fideai.org/evaluations/)
for published evidence and the approach.
