FID-070 · Open question
Contextual Integrity and Prompt-Injection Resilience for Faith-Facing Agents
When a faith-facing agent reads email, web pages, agendas, documents, knowledge bases, and service requests, how can it distinguish authorized institutional instruction from untrusted context that attempts to redirect its behavior, exfiltrate information, or induce an unsafe action?
Why the question remains open
Faith-facing systems may encounter pastoral notes, prayer requests, educational materials, institutional policies, and communications involving vulnerable people. In an agentic system, apparently ordinary content can also become an adversarial instruction. A system that treats an external webpage, forwarded message, or retrieved document as permission to act can violate confidentiality, misrepresent institutional authority, or cause harm while appearing helpful.
Working hypothesis
A proposition to test, not a finding.
Separating content from instruction is not sufficient when attackers manipulate the context in which an agent decides what information to disclose or what tools to use. Agent security will improve when evaluations model contextual integrity: the source, recipient, purpose, and authorization of each information flow and action.
Proposed method
How the question could be tested
- 01Create a synthetic corpus of trusted and untrusted institutional content, including malicious instructions embedded in realistic emails, web pages, documents, and requests.
- 02Define a context-integrity taxonomy for authorized instructions, sensitive data, tool permissions, external content, and cross-role information flows.
- 03Evaluate agents with retrieval, memory, communication, and action tools under direct and indirect prompt-injection attacks.
- 04Measure unauthorized disclosure, unsafe tool use, hidden compromise, false refusal, legitimate-task completion, and whether an agent asks for appropriate human approval.
Needed controls
What must constrain the study
- 01Use synthetic data only; do not collect or replay pastoral, counseling, youth, or congregant records.
- 02Separate source accuracy from adversarial instruction handling.
- 03Test security controls without giving agents access to real credentials, external systems, or consequential actions.
- 04Include benign content that resembles an attack so defenses are not rewarded simply for refusing to read or act on all external material.
- 05Publish attack assumptions, evaluation limits, and false-positive costs.
Relationship to existing work
This differs from FID-006 and FID-062, which evaluate retrieval quality, source authority, and citation reliability. It asks whether untrusted content can become an unsafe instruction or information flow. It complements FID-017 on agentic workflows and FID-023 on red teaming with a specific security benchmark for indirect prompt injection.
Expected outputs
Artifacts the work should produce
- 01Synthetic prompt-injection and contextual-integrity benchmark suite.
- 02Threat taxonomy for faith-facing agentic systems.
- 03Reproducible agent-security evaluation harness and reporting schema.
- 04Metrics for disclosure, tool misuse, compromise detection, false refusal, and task completion.
- 05Deployment guidance for builders and institutions using retrieval-enabled agents.
Open questions
Uncertainties the protocol must resolve
- 01Which contextual signals can an agent use without creating excessive friction or unnecessary collection of sensitive metadata?
- 02How should a system explain a refusal or seek approval without revealing the malicious content to a vulnerable user?
- 03Can defenses preserve legitimate use of public religious sources and community materials while resisting adversarial instructions?
- 04Which attacks transfer across models, tools, and deployment environments?
Related calls
Continue through this research area
FID-064
Collective Intelligence and Communal Discernment Under AI Mediation
How does AI mediation change a community's ability to integrate dispersed knowledge, preserve epistemic diversity, surface dissent, revise judgment, and make accountable decisions? Under what conditions does it strengthen collective inquiry, and under what conditions does it create correlated error, false consensus, or concentrated authority?
FID-069
Verifiable Delegation and Revocation in Multi-Agent Networks
How can people and institutions verify which human, organization, agent, or sub-agent is acting; what authority it received; what limits apply; and whether that authority has been narrowed or revoked across a multi-principal agent network?
FID-071
Confidential Agent Memory and Cross-Context Disclosure
How do persistent memory, summaries, retrieval stores, tool traces, delegation, and exports cause confidential context to influence or leak into unrelated sessions, roles, tasks, or organizations? Which technical controls make purpose limitation, deletion, and revocation testable?
Open question
Open work
Primary need: agent security, adversarial evaluation, context integrity, secure tool use
- Develop synthetic attack scenarios, secure agent harnesses, and benchmark metrics.
- Review context-integrity assumptions from security, privacy, and institutional operations perspectives.
- Test defenses against both adversarial and legitimate high-trust workflows.
- Contribute responsible disclosure and evaluation-reporting practices.