# Methodology and deviation log

## Design record

The original confirmatory design was frozen on 2026-08-10 before the first collection attempt. That attempt stopped at a third-party search-index frame because Spectator's published Terms restrict automated extraction. After the user stated that Spectator had authorized a complete University News crawl and expressly instructed the team to proceed, a prospective University News addendum was frozen at 2026-08-10 18:12 UTC, before any University News treatment outcome was calculated. Its SHA-256 is `26bd2ae9dce549a6d488dd3946051841cba8ed3381b374e92f6da7ae03e72e9b`.

The addendum replaces the discovery-only corpus definition, not the original file. Both preregistrations remain public. The directional hypothesis is that administrators/trustees receive more deferential treatment than pooled student, protester, worker/union, and faculty-critic actors. Competing explanations include topic and event severity, breaking deadlines, article type and length, source availability/safety, legal attribution, announcement timing, subdesk, period, and repeated authors. Null and contrary results were allowed from the outset.

## Input documents

The three supplied DOCX files were preserved byte-for-byte and inspected through the OOXML package, accepted/deleted revision views, comments and anchors, embedded links, and rendered pages. The proposal had no comments or tracked revisions. The revised draft had nine unresolved structural comments and no tracked changes. EV1 had 120 insertion and 55 deletion elements, 12 unresolved comments, and an inserted response letter; element counts are not counts of substantive edits. The revised draft is a clean rewrite rather than simply the accepted EV1 redline.

Document assertions were leads, not ground truth. Potentially publishable claims were separated into the 60-row `claim-ledger.csv`, with status, supporting and contrary evidence, safe wording, caveats, and last-check date. Serious allegations about identifiable people, photographs, discipline, anonymity, “doxxing,” boycott activity, and Public Affairs influence were not repeated as facts without reliable corroboration.

## Archive enumeration and collection

The official `/news/` page identifies University News as three subdesks: Administration, Academics, and Student Life. Collection followed their visible archive pagination through every page and retained canonical article URLs dated 2023-10-01 through the 2026-08-10 cutoff. The deduplicated frame contains 1,527 URLs: 671 Administration, 266 Academics, and 590 Student Life. There were no cross-subdesk duplicates.

Every URL was retrieved successfully. Request starts were separated by at least one second across retries. Five workers overlapped network latency without increasing the global request-start rate. The crawler used ordinary public HTML and parsed the page's embedded `Fusion.globalContent` object; it did not use a login, hidden endpoint, paywall bypass, CAPTCHA bypass, or access-control circumvention. After the independently detected interactive-template parser repair described below, the final private JSONL has 1,527 unique nonempty bodies, no missing or extra URLs, and SHA-256 `b697367b2f86dbd04d11c5a25a0d4113a4f94606febf18b2979e6bc8d628f37b`.

The public evidence bundle excludes bodies, page HTML, photographs, and the input DOCX files. It includes public metadata, normalized-body hashes, coding, and excerpts of at most 25 words.

## Eligibility and units

The census is every article linked by the official Administration, Academics, or Student Life archives in the window. Opinion, City News, Sports, Spectrum, and other desks are excluded by design. All archive articles remain in `articles.csv`; routine items form the neutral comparator. Paired actor contrasts require both an administration/trustee actor and at least one challenger actor. Stories without both groups remain in article-level descriptive rates.

An article unit is one canonical URL, scored as visible at cutoff. An actor unit is one category in one article: `administration_trustees`, `students_orgs`, `protesters`, `workers_unions`, `faculty_critics`, `government_officials`, or `other`. The four challenger categories are reported separately before pooling. A quotation unit is one direct quotation with an attributable speaker; slogans and nested quotations are excluded.

## Automated screen

The public rule-based script applies fixed strict and broad actor dictionaries; masks typographic quotation text for reporter-voice estimates; scores fixed favorable, adverse, delegitimating, legitimating, and conflict terms in actor windows; classifies headline/lede frames; extracts direct quotations and attribution-verb classes; and screens official origin, release reliance, assertion testing, document use, affected-person sourcing, power/history/consequence context, follow-up, visible updates/corrections, article type, and topics.

These fields are measurement aids, not judgments of motive or fairness. Whole-article sentiment is not used. “Claimed” and similar verbs are not automatically unfair; factual adverse events are not automatically negative tone; opposing quotations alone do not count as independent verification. Source and quotation attribution are conservative proxies and can miss delayed or name-only attribution. The actor file contains only categories detected by the strict or broad rules; an omitted category is absent under both dictionaries, not a human-coded zero.

The primary automated estimands use within-article differences:

1. administration minus pooled-challenger scrutiny (positive means more scrutiny of administration);
2. administration minus pooled-challenger marked reporter framing (negative is consistent with administrative deference);
3. administration minus challenger first-quote probability and direct-quote-word share (positive indicates official priority, but is not alone called deference).

Article-clustered paired bootstraps use 5,000 fixed-seed replications. Adjusted linear/linear-probability models use subdesk, primary topic, log word count, publication month, article-type proxy, and author count, with first-author clustered sandwich standard errors. Month absorbs leadership period; leadership coefficients are never interpreted causally. Secondary families receive Benjamini–Hochberg adjustment. Dictionary, threshold, quotation masking, length/type, peak-event, correction, subdesk, and challenger-category robustness checks are preserved.

## Manual validation

Before examining manual outcomes, a deterministic 230-article audit (15.06% of the census) was selected with seed text `university-news-manual-audit-2026-08-10-v1`. Allocation was proportional across subdesk-quarter cells. Within cells, a round-robin over screen signatures intentionally maximized diversity across relevant-topic, actor, official-origin, marked-language, and length screens; SHA-256 rank broke ties. This makes the audit useful for error discovery but not self-weighting or representative of prevalence. Manual rates and intervals describe only the audit set.

Two coders independently reviewed 140 full articles each, with 50 articles overlapping and no cross-inspection. They coded the complete manual codebook: eligibility/type, actor presence and tone, headline/lede frame, actor-specific skepticism, assertion testing, official reliance, documents, affected sources, context, follow-up, first actor/quote, and foreseeable identification risk. A third reviewer adjudicated all 50 overlap articles after the independent files were frozen.

Raw agreement and Cohen's kappa are reported for nominal variables; ordered variables use quadratic-weighted kappa. Automated-versus-adjudicated validation reports agreement, Cohen's kappa, sensitivity, specificity, false positives, and false negatives. Per the frozen addendum, automated semantic population estimates are suppressed when the relevant automated-versus-manual kappa is below 0.60. Mechanical metadata/counts are exempt. Automated quotation-word share, attribution verbs, unique-source counts, image risk, and category-specific presence lack a complete manual gold standard and remain screening/descriptive measures.

## Corrections, missingness, and updates

The latest accessible version at cutoff was scored. Visible corrections, editor's notes, and update labels are retained; silent edits are not inferred. Missing, not applicable, no substantive claim, unclear, and not coded are distinct. All 1,527 pages had body text, so full-text missingness is zero; semantic fields can still be blank when an actor or material claim is absent.

## Post-outcome parser correction

At 2026-08-10 18:28:42 UTC, face-validity review found that the first automated topic proxy searched full bodies, making incidental history label 1,013 of 1,527 stories `protest_palestine`. Before recalculation, the defect, likely consequence, first-run script hash, and affected outputs were logged. Topic scoring was restricted to headline, description, and first two body paragraphs, reducing that label to 540. The pre-fix results are preserved in `automated-screening-results-prefixed-topic.json`. Actor presence, unadjusted treatment outcomes, quotations, and metadata were unchanged. Every topic-adjusted model after the repair is labeled post-outcome exploratory.

Independent review then found that recovered interactive templates had lost paragraph structure and that HTTP-recovered descriptions held subdesk labels. At 19:02 UTC, before the affected recalculation, 69 interactive pages and one concatenated standard story were re-requested at the same rate limit; 204 standard descriptions were rebuilt from structured story text; update/correction labels were excluded from description; and nine changed manual-sample bodies were re-reviewed. The pre-repair corpus and screening JSON remain preserved. `corpus-parser-repair-log.md` records old/new hashes, exact scope, expected consequences, and the five manual rows whose codes changed.

## Ethical and reporting framework

The complete Society of Professional Journalists Code of Ethics is used as a nonbinding external framework: Seek Truth and Report It; Minimize Harm; Act Independently; and Be Accountable and Transparent. Tone, framing, source imbalance, institutional dependence, factual accuracy, and ethical criticism are reported as different constructs. The study does not infer reporter intent, editor ideology, or named-editor causation.

## Post-request leadership-period analysis

After the census and initial validation results existed, the user added a thesis that coverage was best during the April–May 2024 encampment peak and became progressively more pro-administration, especially under the 150th-board leadership of editor in chief Tsehai Alfred and University News co-editor Spencer Davis. This was not preregistered and is labeled exploratory.

The fixed sequence is: encampment/NYPD peak (2024-04-17–2024-05-31); later 148th board (2024-06-01–2024-12-10); 149th board (2024-12-11–2025-12-09); and the right-censored 150th board (2025-12-10–2026-08-10). For each stage, `analyze_leadership_exploratory.py` reports the automated screening gaps and the diversity-sample manual gaps with article-resampling intervals, plus current-minus-encampment contrasts and topic composition. A progressive-worsening claim requires substantively related, reliable measures to move monotonically in the pro-administration direction. No automated semantic measure passed validation; manual period cells are small and nonrepresentative. Period labels organize time only and cannot identify editorial causation.

## Reproduction

`scripts/university_automated_analysis.py` reproduces the census screen from an authorized private corpus; `scripts/prepare_university_adjudication.py` validates independent coding and computes intercoder reliability; `scripts/finalize_university_manual.py` merges the adjudication; `scripts/analyze_university_manual.py` summarizes the audit set; `scripts/analyze_university_validation.py` tests automation; and `scripts/build_university_results.py` creates website JSON. `scripts/check_package.py` checks schemas, references, excerpts, counts, hashes, and copyright exclusions. Exact body-dependent reproduction requires a lawful local corpus whose hash matches the manifest.
