01
Evaluation validity
Do our benchmarks measure properties that matter outside the test? We study construct validity, human and model evaluator reliability, benchmark gaming, distribution shift, and the limits of evaluation evidence.
About Fide AI
Fide AI is an independent research lab building the measurement science needed to understand and create trustworthy AI systems. We study evidence, authority, oversight, and human impact in high-trust domains, including cybersecurity, enterprise, finance, law, healthcare, and faith. Faith is our first domain of published empirical work; broader applications require their own evidence and domain expertise.
Why now
Advanced AI systems are becoming more capable and more deeply embedded in consequential decisions. Yet capability alone does not tell us whether a system uses evidence faithfully, follows the right authority, preserves human agency, resists manipulation, or recognizes when it should defer.
These are technical questions with implications for AI safety, alignment, ethics, and governance. We develop empirical methods for answering them before people and institutions are asked to depend on a system.
What trust requires
01
Do our benchmarks measure properties that matter outside the test? We study construct validity, human and model evaluator reliability, benchmark gaming, distribution shift, and the limits of evaluation evidence.
02
Does a system use the sources it claims to use? Can it distinguish authoritative instructions from untrusted context? We study retrieval, citation fidelity, source delegation, prompt injection, and verifiable authority boundaries.
03
Does an AI system preserve meaningful human judgment and control? We study delegation, escalation, revocation, relational substitution, and the conditions under which a system should defer to responsible people or institutions.
04
Do agents remain aligned with human intent while they plan, use tools, delegate work, and act? We study operational trace sufficiency, policy adherence, drift, escalation, intervention, and recovery without assuming unrestricted access to private reasoning.
05
How does repeated AI use affect attention, formation, work, community, trust, and moral responsibility? We study consequences that are difficult to capture in short benchmark interactions but central to human flourishing.
Research transparency
Every research program begins with judgments about which risks matter, what counts as harm, and what technology is ultimately for. Those judgments can come from disciplinary traditions, domain expertise, moral commitments, or religious worldviews.
Fide AI documents the assumptions and worldviews that materially shape a project, including how they influence its framing, interpretation, and handling of uncertainty. Making the lens visible gives readers a fair way to inspect the reasoning and locate disagreement.
Public methods, evidence, and stated limits keep conclusions accountable. Researchers and collaborators do not need to share the same worldview to contribute to or evaluate the work.
A demanding research setting
Questions involving theology, moral guidance, formation, and pastoral-adjacent care create unusually demanding tests of source fidelity, disagreement, authority, human dependence, and appropriate deference. They give Fide AI a concrete setting in which to study broader problems in trustworthy AI.
Some of our research therefore focuses directly on religious and Christian contexts. Other projects address problems shared across AI safety, alignment, ethics, security, and governance.
Research practice
Benchmarks are one part of the work. Fide AI also develops evaluation protocols, interactive research artifacts, source-verification methods, research agendas, governance analysis, and practical guidance.
Our research develops the methods. Our assurance engagements apply them to consequential decisions. We help organizations understand what the evidence supports, what needs to improve, and when a system should be reassessed. Commercial work supports further public research while findings remain independent.
AI should raise human dignity, not erode it. In faith-facing contexts, that requires more than good intentions; it requires evidence.
I started Fide AI because AI systems are already entering spaces where people ask sacred, moral, and pastoral-adjacent questions, but the public evidence base is still thin. Too much of the conversation relies on intuition, anecdotes, or generic AI safety language that was not designed for theology, formation, religious education, or pastoral care.
My Christian faith shapes why I care about truth, human dignity, moral agency, love, and responsible authority. It is part of why I founded Fide AI. The lab is built for researchers and collaborators across beliefs and disciplines. What we ask is that relevant assumptions are disclosed and claims remain answerable to evidence.
How Fide AI works
Fide AI's work begins with a defined question and a method that can be criticized. We separate research findings from recommendations, state the limits of each result, and publish the strongest technical package the work safely permits.
Study consequential AI behavior through experiments, benchmarks, evaluations, and technical analysis. Publish methods, evidence, and limitations that others can inspect.
Test models and deployed systems under conditions that reflect real sources, tools, instructions, adversarial pressure, and human workflows.
Make technical work accessible without separating conclusions from their methods, uncertainty, or limits.