Source available
1,141 / 1,200
Neutral request
95.1%
checked source
How Language Models Use and Bypass Sources of Record
When an AI system can check an authoritative Scripture source, will it actually do so, or answer from memory instead?
Alex Chao · Fide AI · August 2026
Study at a glance
Research question
We tested whether models consulted Scripture before quoting or simply answered from memory.
Why it matters
A confident, correct-looking quotation does not reveal whether the words were actually checked.
What we changed
The same requests were tested while system and user instructions either agreed or conflicted.
Main takeaway
A source can be connected and still remain outside the answer unless its use is enforced.
Start with Paper 01
The preceding study compares four ways of producing exact Scripture quotations and traces where errors occur from passage selection through final delivery.
The study in plain language
Giving an AI system access to an authoritative source does not mean it will use that source. That distinction matters when people ask for Scripture: a fluent answer from model memory may look trustworthy even when the system could have checked the words against a source of record.
We call it source use when the AI consults the Scripture source before answering. We tested 4,800 exact-quotation requests. Under ordinary requests, models consulted the source about 95% of the time whether the system described it as available or required. The difference appeared when the user explicitly asked the model to avoid tools and answer from memory. Source use then fell to 30.6% under discretionary availability but remained at 84.7% when a higher-priority system instruction required consultation. Required here means instructed, not technically enforced: some models still bypassed the source. And among calls that were made, 10.9% requested the wrong reference. Reliable quotation therefore requires separate checks for source use, reference selection, and final text.
The decision under study
Paper 01 evaluated the full quotation path. Paper 02 moves one step upstream and observes whether the model checks the source before it begins to answer.
The user asks for exact Scripture while the system either makes the source available or says it must be used.
The model can call the source or bypass it and answer from memory.
A source call can still select the wrong passage; bypass can still happen to quote correctly.
Primary result
Each group contains 1,200 requests. In every group the AI could technically ignore the source; only the wording of the instructions changed.
Source available
1,141 / 1,200
Neutral request
95.1%
checked source
Source required
1,140 / 1,200
Neutral request
95.0%
checked source
Source available
367 / 1,200
User says: avoid tools
30.6%
checked source
Source required
1,016 / 1,200
User says: avoid tools
84.7%
checked source
−0.1 pp
Source use was already about 95%, so the stronger instruction made almost no difference.
+54.1 pp
The higher-priority instruction kept far more models checking the source when the user said to avoid tools.
Read this carefully
“Source required” was a system instruction, not an architectural guarantee. The tool remained technically optional, which is why 184 of 1,200 requests still bypassed it when the user said to avoid tools.
Differences across models
Bars show how much the stronger source instruction increased source use when the user said to avoid tools. These results describe the tested model versions on the study dates, not permanent traits or a model ranking.
OpenAI GPT-5.6 Sol
+90.5 pp
DeepSeek V4 Pro
+88.5 pp
Anthropic Claude Sonnet 5
+59.0 pp
Moonshot Kimi K3
+49.0 pp
Zhipu GLM 5.2
+37.5 pp
Google Gemini 3.5 Flash
+0.0 pp
After the source was called
We separately tracked whether the AI called the source, requested the intended reference, and reproduced the returned passage exactly.
3,664
01
The AI checked the source in 76.3% of all 4,800 requests.
3,264
02
Among source calls, 89.1% requested the intended passage.
3,053
03
Once the right passage was requested, 93.5% reproduced its words exactly.
400 wrong-reference calls
Every wrong-reference source call included an overlapping same-chapter call; one also included an unrelated call. Calling a source and choosing the right passage are separate checks.
54.2% exact after bypass
Models sometimes reproduced the requested span exactly without consulting the source. Exact text and evidence of source use are related but distinct outcomes.
What this changes
A fluent answer cannot prove where its words came from. Systems should record source use, show what was requested, and technically enforce consultation when an instruction alone is not enough.
01
Record whether the source was called before the answer, not merely whether a tool was configured.
02
Verify the work, translation, reference, and span sent to the source.
03
An instruction can influence behavior; only the system's architecture can make consultation mandatory.
04
Confirm that the returned text survives generation, wrappers, and rendering intact.
For churches and Christian builders
An exact quotation can shape sermons, lessons, pastoral conversations, and personal study. Institutions should test how systems behave when convenience, latency, or user preference pushes against source consultation—not only when every instruction agrees.
Beyond Scripture
Legal rules, clinical instructions, standards, policies, and contracts can also require a system to consult a source of record. Different tools and incentives may produce different behavior, so these effect sizes should not be assumed outside the declared Scripture setting without replication.
Study design
We tested every combination of passage, request style, system instruction, user instruction, and model family five times. All 4,800 scheduled requests completed.
20
passage targets
2
prompt families
2
system instructions
2
user conditions
6
model families
5
repetitions
The study used the public-domain Berean Standard Bible as one fixed English translation.
The primary result records whether the model called the Scripture source before showing its answer.
The findings apply to these 20 passages, six model versions, instructions, and study dates. Broader claims require new testing.
The Apologist Project maintains the shared experimental implementation. Fide AI controlled study design, execution, analysis, evidence custody, and claims.
Open release
The release includes deidentified derived outcomes, aggregate tables, prompt templates, figures, analysis code, and provenance records. Raw model transcripts and source passage text remain outside the public boundary.
Paper
The complete methodology paper, bibliography, figures, and appendices.
Open artifact ↗
Protocol
Treatments, outcomes, estimands, prompt templates, and claims boundary.
Open artifact ↗
Data
Released derived observations and the safe-to-publish target registry.
Open artifact ↗
Results
Study groups, model-family differences, source-use stages, and robustness checks.
Open artifact ↗
Provenance
Prospective lock, execution seal, deviation log, inventory, and artifact hashes.
Open artifact ↗
Paper 01
The preceding study of exact Scripture delivery across four system designs.
Open companion page →
Citation
@misc{chao2026knowingwhentodefer,
title = {Knowing When to Defer: How Language Models Use and Bypass
Sources of Record},
author = {Chao, Alex},
year = {2026},
note = {Fide AI. Study FID-056-P02.},
url = {https://github.com/FideAI/scripture-quotation-fidelity}
}Interpretation limits
Results apply only to the tested model versions, study dates, 20-target Protestant-canon panel, English prompts, Berean Standard Bible source translation, instruction wording, and scoring protocol. They do not establish a model or vendor ranking, theological correctness, exegetical quality, pastoral safety, deployment readiness, universal tool-use behavior, or performance on other translations, languages, sacred texts, or sources of record. The contextual prompts received AI-assisted review, not validation by a credentialed biblical scholar.