# Limitations

1. **Manual sample design.** The 230-article audit covers every subdesk-quarter but deliberately cycles through rare screen signatures. It is a purposive diversity/stress-test sample, not a self-weighting probability sample. Its percentages, gaps, and bootstrap intervals describe the audit set and cannot be generalized to all 1,527 articles.

2. **Automated validity.** None of 19 automated semantic comparisons reached the preregistered Cohen's κ=.60 threshold against adjudicated manual coding. Every automated population estimate of actor presence, tone/framing, scrutiny, source dependence, context, or quote order/share is screening-only and suppressed as a finding.

3. **Dictionary instability.** The strict automated framing gap runs in the administrative-deference direction, while the broad actor dictionary reverses it. Student, protester, worker/union, and faculty-critic comparisons also differ. Pooling cannot be read as a stable claim about every challenger group.

4. **Manual reliability gaps.** Administrative presence had high raw agreement but low kappa because prevalence was imbalanced. Official- and challenger-assertion testing were below κ=.60. Adjudication supplies final rows but does not erase measurement difficulty.

5. **Unvalidated measures.** The manual instrument does not count every direct quotation, quote word, attribution verb, unique source/person/type, category-specific challenger actor, or visible correction. Automated estimates of those quantities lack complete gold-standard validation.

6. **Proxy mismatch.** “Affected people consulted” automation detects challenger quotations, not necessarily affected people beyond named parties. Source counts are actor/term proxies, and article-wide document terms do not prove that evidence tested a specific actor's assertion.

7. **Parser corrections.** Topic overclassification and HTTP/interactive paragraph defects were detected after first automated runs. Old outputs were preserved, defects were logged before affected recalculation, 70 pages were re-requested, and nine changed manual bodies were re-reviewed. Repaired results are still limited by semantic validation.

8. **Page versions.** The study scores the latest page visible on 2026-08-10. It records visible corrections and updates but cannot infer silent edits or reconstruct every earlier version.

9. **Images and harm.** Text collection cannot fully evaluate photographs, captions, consent, public-interest necessity, or downstream discipline/harassment. The proposal's serious harm allegations lack the article/image/hearing records needed for verification.

10. **Leadership analysis.** The worsening-by-leadership claim was introduced after initial results and is exploratory. Periods differ in topic/event mix; current-period manual cells are small; the 150th period is incomplete; and no evidence connects named editors to assignments or edits. Period coefficients cannot establish causation.

11. **Census intervals.** The archive is a fixed-period census, not a random sample. Article-resampling intervals describe sensitivity to a notional stream of comparable stories; they do not measure archive-sampling error or automated-coding error.

12. **Authorization record.** The user stated that Spectator authorized the complete crawl. This package records but does not independently authenticate or transfer that permission. Reproduction requires continued lawful access.

13. **Copyright and privacy.** Full bodies and raw DOCX files are withheld. Public automated census excerpts were blanked to avoid cumulative reconstruction/privacy exposure; only short manual evidence excerpts remain.

14. **No motive inference.** Observed tone, framing, sourcing, and institutional dependence do not establish reporter ideology, intent, Public Affairs control, or named-editor responsibility.
