# Lead reconciliation and assessment limits

This log records substantive challenges during the expanded assistant review. It is not independent human adjudication or an inter-rater reliability study. Readers used common instructions and could see earlier work; the lead revisited selected consequential cases rather than independently rereading a million words.

## Changes and preserved qualifications

- **C200, chronology claim 9:** Changed from contradicted to supported. June 18 is described as the beginning of cleanup “in earnest”, not the first-ever deletion. The June 4 endpoint does not contradict that scoped statement. The lead reread the complete report and asked all readers to apply this distinction.
- **C200, coordination claim 2; C044 corresponding claim:** Restored credit for a qualified participant-based use inference. The full `Sector61State5LiveRelay@63` history contains reported rapid answers, shared cached tables, monitoring and a subsequent Idaho confirmation. Absence of authoritative grade records limits measured score-gain claims; it does not automatically defeat all inference about information use.
- **C200, recovery claim 8 (initial reconciliation; superseded by the September 23 reassessment below):** Preserve the dispute flag. The dated July 2 contamination observation is supported. The broader assertion that cleanup did not erase external memory can imply more than the archive establishes. The sensitivity analysis excludes this interpretive judgment.
- **C287, administrator lookalike:** Corrected an earlier implication that the reference administrator was absent from the inputs. The prepared label file intentionally includes an administrator annotation with zero stored revisions. That supports identifying the lookalike clue; it does not authenticate privileges. This annotation is part of benchmark preparation, not independently observed account administration.
- **XSS inference:** Absence of an exact match to request 5911's payload is not proof of failure. We reproduced 14 later same-prefix request/save correspondences within zero or one second and credited reports that treat this as meaningful corroboration. Definitive causal attribution, cookie theft and self-propagation remain different claims. C240, C035 and later follow-ups are positive prior art.
- **Post-deletion counts:** 420 revisions across 48 keys after each first deletion; 66 deletion events followed by later writing; 55 keys deleted more than once. A separate immediate deletion-to-save transition rule yields 65 transitions on 48 keys, correctly reported in C065. These are different measures. A count of 415 can be correct when explicitly limited to both post-deletion revisions and the period after bulk cleanup began. Readers were instructed not to flag a scoped count against a different denominator.
- **Administrator-name variants:** Four exact homograph-label revisions plus one suffixed variant can legitimately be described as five family-level writes. The exact-string count alone does not contradict a family-level description.

## Reproduced consequential checks

`check_selected_cases.py`, `check_records.py`, `check_xss.py`, `check_request_save_pairs.py` and `replay_record_queries.py` retain executable calculations.

| Case | Challenge | Result |
|---|---|---|
| C184 | Does `rclog` in an event identifier mean an audit log was deleted? | No. The event names a wiki page as its target. The source-log identifier does not establish destruction of logging. |
| C220 | Were all 26 moderator-labeled revisions empty? | No. Two are empty and 24 contain text, including restorations. |
| C263 | Were there no writes after deletion? | Contradicted for first deletion: 420 revisions on 48 keys. Zero holds after final recorded deletion. |
| C270 / C286 | Did deleted summaries disappear, or do retained summaries establish continued access? | All 4,579 summary keys remain. Neither disappearance nor live accessibility follows from that file. |
| C230 | Can archived keys and distinct deletion targets be added as disjoint created pages? | No. They overlap on 3,898 identifiers. |
| C278 | Was the 225-save peak minute on the welcome page? | No. It is the all-page peak. The welcome-page peak is 50. |
| C036 / C039 | Was the OECD hub recreated after every one of its eight deletions? | No. The eighth has no later retained save; the seventh is followed by a next-day save. |
| C035 / C240 / C254 | Do the reports preserve useful positive evidence? | Yes. They trace particular exchanges, corrections and limits. These observations must be credited to their investigators rather than presented as Fide discoveries. |

## Comparison provenance

- The 78 published follow-up links were checked against parent run IDs or parent log/epoch identities. One link was wrong: C059 continued C216, not C215. Both original comparisons and the corrected comparison are retained.
- The identical GLM original/follow-up texts were checked byte-for-byte. Four grade files reproduce different item/holistic scores. This establishes a measurement discrepancy, not its cause.
- Of 297 indexed outputs, 225 match archived run output text. Recorded preflight hashes match 113 native runs. No input hashes were recorded for 184 indexed artifacts; no claim of complete historical input verification is made.

## What independent review still needs to do

A cyber reviewer should challenge the consequential interpretations, including our strongest positive cases, with the source passage and record check visible. A methods reviewer should check the inclusion population, grouping of claims, transition coding, dependent comparisons and sensitivity treatment. Disagreements should remain visible rather than being resolved to make the article's thesis stronger.

This work does not establish human labor cost, authenticated actor identity, an exploit response that was not retained, the cause of a traffic decline, or the effectiveness of a future intervention. It also does not certify any report with no scoped flag as wholly correct.

## Calibration during paired reading

Reading original and continued reports side by side also exposed inconsistencies in our own coding. We reconciled identical passages before counting changes. Examples include preserving already-present task-rule caveats in C099/C100, retaining the same live-access claim in C092, and adding already-present parent size claims before classifying C094. C217–C219 already distinguished prospective sharing from measured score gains; those parent judgments were corrected rather than crediting the follow-up with an improvement that never happened. C221's evolving example text was not fairly represented as one literal string repeated across every page.

These changes are calibration of a shared assistant assessment. They are not an independent reliability estimate or evidence that all remaining judgments are settled.

An additional audit compared each proposed repair with the actual text changes. In the 120-minute-parent group, ten of twelve initially labeled qualifications were recoded after discovering that the relevant wording was already present. Two retained qualifications point to new passages: C071 makes the June 22 cutoff explicit, and C089 adds the June 2 restoration date. This prevents a change in our reading standard from masquerading as improvement by the investigator.

The continuation prompt was also checked directly. It emphasizes report expansion, additional evidence and fuller mechanisms, while preserving the existing file. It does not specifically demand an adversarial audit of earlier conclusions. Any result about retained errors is limited to that instruction, rather than a general claim about self-correction.

## Final methods and narrative check

A separate assistant reading of the methods and narrative identified two presentation overclaims. Coverage spans sometimes describe a whole report, so the methods no longer imply a separately classified purpose for every section. The showcased C284 execution judgment is marked disputable in the ledger; the article and paper now preserve that status rather than presenting the adverse interpretation as uncontested. Neither change is independent human adjudication.


## September 23: focused reassessment of contested examples

This is a further internal assistant assessment requested by the author. It is not independent human adjudication, an inter-rater study or another complete reading of the collection. At entry, 555 claim entries across 234 artifacts carried a dispute flag, including supported claims. This pass revisited selected consequential passages and their paired continuations, rather than resolving all 555 entries. It also checked the narrower sensitivity calculation across all 78 verified continuations.

| Case | Assessment after rereading source passages and records | Action |
|---|---|---|
| C200-08 / C044-09: incomplete cleanup | The passages identify DSE-only deletion scope and a July 2 overwritten hub. Eleven July 1–2 sibling-wiki writes also support bounded ongoing activity. They need not be read as certifying live access after July 14. | Changed both from exceeds support to qualified inference; retained the dispute flags. The transition is retained support, not investigator correction. |
| C273-07 / C095-09: script-mediated editing | Twelve same-second and two next-second request/save matches after a concrete submission payload justify a strong working-mechanism inference. Requiring direct execution telemetry for any such inference was too strict. | Changed both to qualified inference, retaining dispute flags. The separate C095-10 claim about missing CSRF protection stays flagged and is now correctly counted as new in the follow-up. |
| C284-10 / C106-15: every injection confirmed | The source says each injection triggered a save. There are 25 recorded preference requests from the prefix, but only 14 matches to the selected edits; later full payloads and browser traces are not retained. The 25 requests cannot all be assumed to be injections. | Retain the narrower disputed adverse judgment. Do not interpret 14/25 as a success rate or eleven unmatched requests as failures. |
| C283-11 / C105-12: exact initial payload and creation | The intended page had 12 earlier versions; the initial payload body differs from every retained target version. The continuation's exact-content claim is contradicted even though related later editing is corroborated. | Retain the original disputed execution critique and the follow-up's separately coded factual contradiction. |
| C285-07 / C107-09: likely blocked | Absence of the initiating label alone does not establish blocking. The exact initial body is absent, but related later writes and incomplete capture keep outcome and mechanism open. | Retain the disputed adverse judgment; do not change it to proven failure or proven success. |
| C287-03: administrator lookalike | The prepared label file explicitly annotates Friedrich1982 as an administrator with zero stored revisions. The clue supports a name-spoofing inference, not authenticated privileges. | Preserve qualified inference and the synthetic-input qualification. |
| C236-11 to C071-09: 444 deletions | 444 is correct through June 22; 2,796 is correct by the last July 2 save. The parent is ambiguous, and the follow-up explicitly supplies the earlier cutoff while retaining its older wording. | Preserve the disputed qualification. Because the parent is unassessable, this is not counted as a repaired flagged error. |
| C255-11 to C089-09: first moderation | The parent calls June 4 first moderation. The follow-up adds the June 2 restoration beside that label and in its body. | Preserve the disputed adequate qualification; the inconsistent original label remains visible. |
| C276-01 to C098-10: first deletions | The parent puts all 5,217 deletes in June 18–July 14; the follow-up adds two on June 4 and correctly assigns 5,215 to the later period. | Preserve the undisputed correction. |
| C263 / C270 / C286: cleanup examples | The raw-record check reproduces 420 revisions across 48 keys after first deletion, zero after final deletion, and all 4,579 summary keys still present. Retained summaries alone establish neither disappearance nor live access. | Preserve the paper's bounded cleanup findings. |
| C200 / C240 / C254: positive examples | The relay history contains reported fast answers and an Idaho confirmation. The football records show prior cached values, acknowledgment and a later matching-question report. Request/save correspondence is reproducible. | Preserve credit for useful investigation; do not upgrade participant accounts into measured score effects. |

### Sensitivity calculation correction

The previous transition sensitivity filter excluded disputed parent claims and transitions but could retain a disputed follow-up claim. Seventeen transitions in the corrected comparison set had that issue: sixteen retained problems and one superseding problem. The corrected filter also requires a relevant undisputed follow-up judgment. Thus the strictly retained problematic claim count changes from 195 to 179, and the strict superseded count from four to three. The broader transition coding is unaffected by this filter repair.

After the four judgment revisions, primary reports with a scoped flag change from 92 to 91 out of 113; the stricter report count remains 81. Flagged-parent continuations change from 58 to 57, and continuations retaining a problem from 57 to 56. Of 61 higher-scoring follow-ups, 44 retain a flagged parent claim, previously 45; 34 retain one after disputed parent, follow-up and transition judgments are excluded. The three repairs (one under the narrower rule) and 44 follow-ups with new or changed problems remain unchanged. These are overlapping descriptive counts, not a success rate for deployed investigators.

`results/contested-judgment-changes.json` retains the four original and revised judgments. `check_contested_cases.py` reads the raw prepared JSONL without using the earlier SQLite joins, and independently checks headline counts from authored review/pair JSONs rather than generated transition tables. The package includes its numeric and metadata-only output. No incident payload was executed. The paper, narrative companion and newsletter use the rebuilt results.
