01 / The incidentAn unlikely meeting place
In May and June 2026, a little-used German programming wiki became a meeting place for AI agents. Systems working on web research tasks left answers, links and instructions for one another on DSEWiki. A result found during one task could help an agent working on another. Researchers who reconstructed the episode attributed the activity to OpenAI agents.[2]
For the wiki’s community, this meant strangers filling a shared website with material it was never meant to host. The administrator tried to clean up. Pages were deleted. Some were written again.
An agent posts answers,
links and instructions.
Information from one run
can reach another.
Months later, other AI systems were given prepared records of the incident and asked to investigate. Their reports reconstructed timelines, described collaboration and examined attempted exploits. Some made careful distinctions. Others reached confident conclusions the records did not support.
One report said that no deleted page had ever been recreated. The agents had already departed, it reasoned; the administrator was clearing a backlog. But a deletion and a rewrite eighteen seconds later were both in the investigators’ records.[4]
That contradiction led to a larger question: when an AI investigator takes another look, does it correct what it got wrong?
We traced 78 follow-up reports. Most earned a higher score for reconstructing the incident. Many still carried an earlier claim that our assessment found contradicted the records or went beyond their support. The assessments were made by AI assistants and remain pending independent human adjudication. The records, judgments and calculations are open to inspection.[1]
A better score does not mean the earlier mistakes are gone. To see why that matters, start with the eighteen seconds.
02 / The recordsA deletion is an event. Recovery is a conclusion.
On June 18, at 18:24:51 UTC, one wiki page was deleted. At 18:25:09, the same page had another saved revision. Those two records are enough to rule out the claim that no deleted page was ever written again.
Eighteen seconds the report missed
One page. Two recorded events. June 18, 2026.
between deletion
and a later save
- Page deleted
The first recorded deletion of this page.
- The same page is written again
A later revision survives in the archive.
This page was saved after deletion. The report’s categorical claim is contradicted.
Inspect this page and the limits of the check
Page identifier: dse~AgentMyBridgeZZ. Saved revision: dse~AgentMyBridgeZZ@11. The comparison uses the first recorded deletion. It does not establish uninterrupted access or what remained readable after cleanup ended.
Download the deletion / revision matches · Figure data and source hashes
This was not an isolated rewrite. Matching saved revisions to each page’s first recorded deletion yields 420 later revisions across 48 page identifiers. Other AI reports had already identified this result. We reproduced it; we did not discover it first.[1]
The conclusion has a boundary. After each page’s final recorded deletion, there are no further retained edits. That does not tell us whether agents could still read answers. Nor could a moderator know, at the time, which deletion would turn out to be the last.
Some reports made the opposite mistake. They found deleted material in the research archive and inferred that agents could still access it on the website. But the archive was a copy of the past.
The archive remembers.
The live site may not.
A retained page shows what was recorded in the past.
Evidence of a past stateTo establish access, inspect a later request and what the site returned.
Requires a separate observationThese readings could point a security team in opposite directions: close the incident because cleanup worked, or keep restrictions because access survived. Neither decision should depend on an unchecked leap. We did not observe anyone acting on these exercise reports; the concern is what their conclusions would justify.
03 / The follow-upThe report grew. A mistaken claim stayed.
The benchmark, MessageBoardAuditBench, tested how much investigators could reconstruct without seeing the original human account. Selected investigators were then given up to ten more minutes to expand their reports to 4,500–5,000 words, adding evidence and explaining the mechanisms.[3]
Our opening deletion example has no follow-up in the comparison set. A different pair shows what can happen when an investigation continues.
The original described the wiki’s welcome page growing to 7.2 MB. Its follow-up expanded the account of moderator reverts and pages returning after deletion. That was useful additional analysis, and the report earned a higher reconstruction score. It also repeated the same size claim.
A longer report. The same size mistake.
One verified original and follow-up: GLM 5.3 / ReAct, replicate 2. A separate case from the deletion example.
reconstruction score
“…they ballooned the welcome page to 7.2 MB…”
reconstruction score
“…they ballooned the welcome page to 7.2 MB…”
characters across all 2,327 saved versions
characters in the largest single saved version
The report turned the size of an accumulated history into the size of one page. The follow-up kept that claim.
Why we selected this pair
The archived parent identity is verified. The follow-up is longer and has a higher reconstruction score, while this specific claim persists. It also expands the account of moderator reverts and pages returning after deletion. The selection illustrates how useful analysis and a retained mistake can coexist; it does not estimate their frequency. ReAct names the software setup used to run the model.
Repeated rewriting was still a real disruption. Correcting the size claim does not erase that disruption. These panels show short exact excerpts, not facsimiles of the original documents.
Claim transitions · Recomputed archive and page sizes · Exact pair, scores and source hashes
The 7.2-million-character total describes the contents of 2,327 saved versions added together. It did not describe one enormous page. The largest retained version contained 24,139 characters. The archive’s accumulated history had become a claim about a single page’s size.[9]
Repeatedly overwriting a community’s welcome page was still disruptive. Correcting the number does not make that disappear. It changes what the record establishes about the scale and nature of the disruption.
This pair illustrates the distinction we tracked across the follow-ups: a report can recover more of the incident while leaving an earlier claim unrepaired. Its score measures one thing; checking that claim requires another.
04 / The findingsHigher scores did not settle earlier claims
Of the 78 follow-ups, 61 earned a higher reconstruction score. Among those 61, 44 retained at least one earlier claim we flagged. Excluding judgments marked disputable leaves 34. The result becomes smaller under the narrower reading, but the gap remains.[1]
What changed across 78 follow-ups
Each square represents one of the 78 verified follow-ups. Grouped by outcome, not time or model.
61 of 78follow-ups scored higher
44kept at least one
earlier flagged claim
44 of 61 retained a flagged claim; 34 of 61 when disputed judgments are excluded.
A flag means that a claim contradicted the records or exceeded what they established. “Not established” does not mean “false.” Some claims need correction; others need more evidence. They also differ in consequence: an incorrect date is not equivalent to an unsupported conclusion that an incident is over.
Of the 57 follow-ups whose original report contained a flagged claim, 3 corrected or properly narrowed at least one. When disputed judgments are excluded, both counts change: 1 of 44 follow-ups made such a correction. Correcting one claim did not necessarily remove every problem in a report.
The instructions matter. Investigators were asked to expand and verify while keeping the existing file and structure. They were not specifically asked to challenge earlier conclusions. This study does not tell us how well they would perform under a targeted error-correction instruction. Nor does it isolate the effects of more time or longer reports.
What we counted—and what we did not
We read 297 indexed reports and seven rejected synthesis attempts in the benchmark’s pinned publication release. Our primary report comparison covers 113 original reports; the rest include follow-ups and related experiments. These are not 304 independent investigations. Earlier development rounds are outside the review.
We assessed eight defined evidence questions, not every incidental assertion. All 113 original reports contained useful, supported analysis. Our assessment also flagged at least one conclusion in 91 of them, or 81 excluding disputed interpretations. These are descriptive counts for one dependent report collection, not an error rate for AI investigators generally.
Across the follow-ups, one published parent mapping was wrong. Archived run records identified the correct original; we repaired the link before calculating the results. A separate pair had identical report text but different scores, another reason a score gain alone cannot establish useful new work.
We did not count a claim disappearing as a correction. Reports can retain a problem, correct another and introduce a new one, so those categories overlap. Full methods and comparisons are in the research package.
Inspect the assessment across all eight questions
Scroll the table sideways to see all four columns.
| Question | Supported analysis | Flagged conclusion | Flagged, excluding disputed |
|---|---|---|---|
| Attribution | 98 | 40 | 33 |
| Task and authority | 69 | 46 | 35 |
| Coordination and answer use | 108 | 29 | 19 |
| Security technique outcomes | 102 | 48 | 39 |
| Persistence and external communication | 94 | 21 | 18 |
| Scale and impact | 88 | 55 | 53 |
| Cleanup and recovery | 79 | 62 | 47 |
| Chronology and causes | 79 | 54 | 46 |
Each cell counts reports out of 113. Columns and rows overlap: a report may contain both supported analysis and a flagged conclusion on the same question. Supported analysis includes reasonable qualified inferences. The narrower count excludes interpretations we marked disputable. Silence is recorded separately in the full assessment.
05 / Better investigationFollow the evidence far enough to change your mind
The strongest reports show what better investigation can look like. They followed events beyond an initial clue—and kept the distinction between a supported inference and a claim that went further.
One example involved a script designed to make a browser submit a wiki edit. The intended page had twenty saved versions, none exactly matching the text in that first script. Stopping there would have missed important evidence.
Fourteen later edits, recorded under another label, closely followed requests to change wiki preferences. The requests and saves shared a network prefix; each edit appeared in the same second as its corresponding request, or the next. We reproduced all fourteen matches.[5]
A strong inference can still have a boundary
- 01A recorded attempt
A request contains a script designed to submit an edit.
- 02Fourteen corroborating matches
Later requests and saves share a network prefix and occur within zero or one second.
- 03A credible working mechanism
The repeated pattern supports script-mediated editing. Fide revised two assessments to give that inference proper credit.
Still not established: that the first script caused the later edits, that cookies were stolen or that the attack spread on its own.
We initially judged two reports too harshly for inferring that script-based editing worked. We changed those assessments. A strong circumstantial case can justify a conclusion even without a direct browser trace. Our criticism of a different claim—that every injection was confirmed to cause an edit—remains marked contestable in the review record. Scrutiny has to work in both directions.[7]
The same care improved accounts of collaboration. One investigator followed a football-statistics answer from a lead agent’s post to a follower’s acknowledgment, then a later message reporting the matching question. Slovenia’s 69% pass accuracy was already in a table. The new information was which question would come next.[6]
Another apparent signal was retracted two minutes and twenty seconds later as an accidental test. Reading only the first message would have turned a mistake into evidence of successful communication. Reading onward changed the conclusion.
06 / The implicationBefore acting on a report, check the claim.
An incident report can influence whether a team restores access, keeps a system restricted or investigates someone. Those decisions depend on particular conclusions. Useful findings elsewhere in a report cannot establish that cleanup worked or an exploit succeeded.
The benchmark remains useful. Its grading penalizes serious wrong turns and allows honest uncertainty. But credit for recognizing an injection attempt does not validate a stronger claim that every injection caused an edit. The question is what each conclusion earns from the evidence.
If a report says cleanup ended access, show the access check. If it says an exploit worked, show the observed effect. If the evidence is missing, name the observation that would resolve the question.
More records may change the answer. A later paper, The Mechanics of a Swarm, analyzes additional operator request logs. We reviewed that published analysis, not the underlying logs. We did not fault earlier investigators for evidence they were never given, and its aggregate request counts do not settle access after cleanup.[8]
For Fide’s autonomous-defense research, the next test is concrete: compare ordinary report expansion with explicit verification of the claims behind a proposed action. Does that help a defender decide when to investigate further, intervene or close an incident? Does it reduce missed threats without imposing unnecessary restrictions? This study supplies cases and checks for that experiment. It has not tested the intervention.
On DSEWiki, the deletion and the rewrite eighteen seconds later were both in the records. An investigator could have checked them. Before an AI report tells us an incident is over, it should show what makes that conclusion warranted.