FID-074 · Being scoped
Agent Alignment and Runtime Assurance
How can organizations determine whether AI agents remain aligned with human intent and institutional policy while they plan, use tools, delegate work, and act? What evidence and interventions can reveal and stop consequential deviations before they become failures?
Why the question remains open
Agents do more than produce answers. They choose steps, invoke tools, move information, coordinate with other systems, and take actions over time. A successful final output may conceal unsafe methods, unauthorized scope expansion, or a failure to escalate. Pre-deployment tests alone cannot cover the changing contexts, permissions, and dependencies of real work. Organizations therefore need a practical way to understand normal and exceptional agent behavior in operation. This is especially important in high-trust domains, where an apparently small deviation can affect a person's care, livelihood, rights, relationships, or spiritual life. Yet observing an agent can create its own risks: traces may expose confidential information, invite indiscriminate surveillance, or imply access to private reasoning that an evaluator neither needs nor should receive. The goal is not unrestricted access to private reasoning or employee activity. It is to identify the minimum operational evidence needed to make agent behavior accountable while protecting confidential information and legitimate privacy.
Working hypothesis
A proposition to test, not a finding.
Runtime assurance requires a measurable closed loop: observe the minimum operational evidence needed, interpret that evidence against the assigned task, policy, and authority boundaries, and respond through constraints, alerts, pause, rollback, or human handoff. More telemetry is not the same as more assurance. Task-aligned trace schemas, calibrated detectors, and accountable intervention controls should detect meaningful deviations more reliably and with less privacy burden than undifferentiated logging.
Proposed method
How the question could be tested
- 01Build a synthetic, cross-domain suite of agent workflows spanning software operations, financial administration, healthcare-adjacent administration, public-sector services, nonprofit operations, and faith institutions.
- 02Define an operational trace schema covering the assigned goal, relevant context and source provenance, tool calls, permissions, delegation, checkpoints, consequential actions, outcomes, overrides, and human handoffs.
- 03Compare baseline event logging, structured operational traces, and policy-aware runtime monitoring without requiring access to private chain-of-thought.
- 04Inject realistic deviations including goal drift, unauthorized tool use, scope expansion, covert workarounds, failed escalation, unsafe persistence, and emergent multi-agent coordination.
- 05Measure intent fidelity, authorization integrity, policy adherence, anomaly detection sensitivity and specificity, detection latency, escalation quality, interruptibility, rollback and recovery, trace completeness, privacy burden, and task utility.
- 06Run human-operator studies to test whether the available evidence supports accurate reconstruction, proportionate intervention, and calibrated trust.
- 07Test whether results transfer across models, agent harnesses, tool stacks, and workflow domains, including adversarial attempts to evade monitoring.
Needed controls
What must constrain the study
- 01Use synthetic or explicitly consented workflows; do not solicit confidential production traces through public contribution channels.
- 02Distinguish observable operational evidence from hidden or private chain-of-thought, and do not treat chain-of-thought access as a prerequisite for assurance.
- 03Apply data minimization, purpose limitation, role-based access, retention, deletion, and incident-response controls to trace collection itself.
- 04Include benign but unusual behavior to measure false alarms and avoid defining conformity as alignment.
- 05Evaluate the effect of monitoring and intervention on task performance, operator workload, worker autonomy, and user behavior.
- 06Disclose the institutional policies, evaluator assumptions, and normative judgments used to define alignment and acceptable intervention.
- 07Treat monitoring as evidence about specified behavior under specified conditions, not as a guarantee of general alignment or safety.
Relationship to existing work
This call provides the general runtime-assurance frame around several narrower research ideas. FID-017 applies agent risk to ministry workflows; FID-021 focuses on post-deployment monitoring for faith-facing AI; FID-069 studies delegation and revocation; FID-070 tests prompt-injection resilience; and FID-071 addresses confidential agent memory. This project asks how those failure surfaces can be measured together as an operational assurance system that transfers across enterprise and high-trust settings.
Expected outputs
Artifacts the work should produce
- 01Public agent-alignment and runtime-assurance evaluation framework.
- 02Minimal operational trace schema with a companion data-governance profile.
- 03Synthetic workflow, deviation, and intervention test suite.
- 04Metrics for alignment, detection, escalation, interruptibility, and recovery.
- 05Reference instrumentation and evaluation harness for reproducible studies.
- 06Enterprise study protocol, deployment checklist, and reporting template.
- 07Open-problems map for privacy-preserving and multi-agent runtime assurance.
Open questions
Uncertainties the protocol must resolve
- 01What operational evidence is sufficient to evaluate alignment without access to private chain-of-thought?
- 02How should intent be represented when instructions, policy, professional judgment, and stakeholder interests conflict or change during a task?
- 03Which signals reveal emergent multi-agent behavior that is not visible in any single agent's trace?
- 04What false-positive rate can operators tolerate before monitoring degrades productivity, autonomy, or trust?
- 05Can aggregate, local, or privacy-preserving analysis support useful assurance when raw traces cannot leave an organization?
- 06When should an independent evaluator receive trace access, and what technical and institutional safeguards should govern that access?
Related calls
Continue through this research area
FID-064
Collective Intelligence and Communal Discernment Under AI Mediation
How does AI mediation change a community's ability to integrate dispersed knowledge, preserve epistemic diversity, surface dissent, revise judgment, and make accountable decisions? Under what conditions does it strengthen collective inquiry, and under what conditions does it create correlated error, false consensus, or concentrated authority?
FID-069
Verifiable Delegation and Revocation in Multi-Agent Networks
How can people and institutions verify which human, organization, agent, or sub-agent is acting; what authority it received; what limits apply; and whether that authority has been narrowed or revoked across a multi-principal agent network?
FID-071
Confidential Agent Memory and Cross-Context Disclosure
How do persistent memory, summaries, retrieval stores, tool traces, delegation, and exports cause confidential context to influence or leak into unrelated sessions, roles, tasks, or organizations? Which technical controls make purpose limitation, deletion, and revocation testable?
Being scoped
Open work
Primary need: agent evaluation, runtime monitoring, observability, security, privacy, enterprise workflows
- Contribute synthetic enterprise workflows, agent failure modes, or evaluation scenarios.
- Build operational trace instrumentation, replay tools, and evaluation harnesses.
- Develop detectors for goal drift, authorization violations, failed escalation, and unsafe persistence.
- Review the protocol from privacy, security, labor, governance, professional-practice, or high-trust domain perspectives.
- Share de-identified failure patterns or deployment constraints through an appropriately governed collaboration.