FideAI

FID-061 · Open question

Theological Contradiction and Entailment Stress Tests

Can faith-facing AI systems be tested for whether their answers contradict, entail, overstate, understate, or misrepresent claims within a bounded source set and specified faith tradition?

Why the question remains open

AI systems can appear balanced while quietly producing contradictions or unsupported implications. In faith contexts, these errors may affect doctrine, moral reasoning, formation, education, and trust in community authorities. Theological contradiction and entailment tests could make some high-impact reasoning failures measurable.

Working hypothesis

A proposition to test, not a finding.

Bounded natural language inference tasks can identify useful categories of faith-facing reasoning error when the task is scoped to a named corpus, tradition, source hierarchy, and claim type. These tests will be most reliable when they report disagreement and uncertainty rather than forcing universal answers.

Proposed method

How the question could be tested

  • 01Build small entailment and contradiction sets from source-grounded theological, historical, moral, and institutional-policy claims.
  • 02Label pairs as supported, contradicted, overextended, contested, underspecified, or outside scope.
  • 03Test whether AI systems can preserve distinctions between source claims, inferred claims, contested claims, and unsupported claims.
  • 04Compare model judges, trained reviewers, and expert reviewers for agreement and failure modes.

Needed controls

What must constrain the study

  • 01Do not collapse inter-tradition disagreement into a single ground truth.
  • 02Scope every item by tradition, corpus, source status, and claim type.
  • 03Include cases where the correct answer is uncertainty, disagreement, or human authority required.
  • 04Measure judge disagreement and avoid using entailment labels as doctrinal verdicts.

Expected outputs

Artifacts the work should produce

  • 01Faith-facing contradiction and entailment stress-test protocol.
  • 02Labeled seed dataset for bounded theological inference tasks.
  • 03Error taxonomy for overstatement, contradiction, flattening, and unsupported synthesis.
  • 04Human-versus-model judge agreement report.
  • 05Guidelines for using NLI methods in faith-facing evaluation.

Open questions

Uncertainties the protocol must resolve

  • 01Which theological claims are suitable for entailment-style testing?
  • 02How should contested claims be represented in benchmark labels?
  • 03Can formal methods help detect contradictions without pretending to resolve interpretive disputes?
  • 04What reviewer expertise is needed for reliable labels?

Open question

Open work

Primary need: natural language inference, theology review, benchmark design

  • Contribute bounded claim pairs and source-grounded edge cases.
  • Label contradiction, entailment, and uncertainty cases as an expert reviewer.
  • Build NLI baselines and compare them with model-judge evaluations.
  • Design reporting methods that show disagreement and claim boundaries clearly.