FideAI

FID-070

Contextual Integrity and Prompt-Injection Resilience for Faith-Facing Agents

When a faith-facing agent reads email, web pages, agendas, documents, knowledge bases, and service requests, how can it distinguish authorized institutional instruction from untrusted context that attempts to redirect its behavior, exfiltrate information, or induce an unsafe action?

Why this matters

The question behind the brief.

Faith-facing systems may encounter pastoral notes, prayer requests, educational materials, institutional policies, and communications involving vulnerable people. In an agentic system, apparently ordinary content can also become an adversarial instruction. A system that treats an external webpage, forwarded message, or retrieved document as permission to act can violate confidentiality, misrepresent institutional authority, or cause harm while appearing helpful.

Work advancing this call

From open question to cumulative evidence.

This directory links Fide AI research to the call it addresses. Relevant work from other organizations is listed separately and added through manual review.

No Fide AI work is linked yet.

This call remains open for research, implementation, review, or partnership.

External work is not presented as Fide AI research or endorsement. Each item must include a specific explanation of how it advances this call.

Suggest related work ↗

Metadata

How to place this call.

agent securityadversarial evaluationcontext integritysecure tool useformationgovernanceresearcher

Ways to help

Move this from question to evidence.

Develop synthetic attack scenarios, secure agent harnesses, and benchmark metrics.

Review context-integrity assumptions from security, privacy, and institutional operations perspectives.

Test defenses against both adversarial and legitimate high-trust workflows.

Contribute responsible disclosure and evaluation-reporting practices.

Contribute

Choose a public issue path or contact Fide AI.