# Preregistration addendum: authorized University News census

**Frozen:** 2026-08-10 18:12 UTC, before calculation of University News treatment outcomes  
**Study window:** 2023-10-01 through 2026-08-10  
**Relationship to original preregistration:** This prospective addendum replaces the original discovery-only corpus definition and access-barrier deviation because the user subsequently stated that Spectator authorized a crawl of the complete University News archive. The original preregistration remains preserved and is not rewritten.

## Research question and competing explanations

**Directional hypothesis.** In University News coverage, Columbia/Barnard administrators and trustees receive more deferential treatment than students, student organizations, protesters, workers/unions, and faculty critics. Observable implications include less negative or marked reporter framing, less testing of official assertions, earlier/more quotation, and greater untested reliance on official releases or statements.

**Competing explanations.** Any raw gap may be produced by event severity, the factual posture of a dispute, breaking-news deadlines, article length/type, subdesk, source availability or safety, topic, period, repeated authors, legal attribution conventions, or administrators' control of announcement timing. University News may instead treat administrators more skeptically, show no meaningful gap, or vary by outcome and topic.

## Corpus and eligibility

The census frame is every canonical article linked by Spectator's official `News | Administration`, `News | Academics`, or `News | Student Life` archives with a URL date from 2023-10-01 through 2026-08-10 inclusive. The archive contains 1,527 unique URLs: 671 Administration, 266 Academics, and 590 Student Life. Opinion, City News, Sports, Spectrum, and other sections are outside this addendum even if a University News staff member contributed.

All 1,527 rows remain in the public article index. Article-level treatment comparisons require at least one relevant actor. The primary within-article actor comparison requires both an administration/trustee actor and at least one pooled challenger actor. Routine stories without challengers remain a neutral comparator and inform announcement/release-dependence estimates, but they do not enter paired actor-treatment contrasts.

## Units

- **Article unit:** one canonical University News URL, using the latest page version accessible at the 2026-08-10 cutoff.
- **Actor unit:** one actor category within an article. Primary categories are `administration_trustees`, `students_orgs`, `protesters`, `workers_unions`, and `faculty_critics`; government and other actors are retained as contextual categories. The primary challenger comparison pools the four challenger categories only after category-specific estimates are reported.
- **Quotation unit:** one direct quotation attributed to a source. Nested quotations and unattributed slogans are excluded from actor quote-share calculations.

## Automated variables fixed before scoring

Automated results are screening/measurement aids and never evidence of motive. Dictionaries and rules are stored in the analysis script.

1. **Actor presence:** case-insensitive entity/dictionary matching in headline, description, and body, with role-sensitive exclusions for sports teams, locations, course titles, and quoted institutional names.
2. **Headline and lede actor frame:** reporter-written actor windows in the title, description, and first two body paragraphs. Clearly attributed allegations and quotation content are separated from the reporter's voice.
3. **Actor-window evaluative language:** counts of fixed favorable, adverse, delegitimizing, and legitimating terms within actor windows, normalized per 1,000 actor-window words. Negation and quotation masking are applied before scoring.
4. **Direct quotations and placement:** attributed-quotation counts, quoted-word counts, first substantively mentioned actor, and first quoted actor.
5. **Attribution verbs:** neutral, epistemically marked, concession/fault, and emotive/adversarial classes from the frozen codebook.
6. **Official-origin/release reliance:** an official email, statement, release, announcement, spokesperson, or University webpage supplies the event/central facts in the headline, description, or first two paragraphs. A high-reliance flag additionally requires no visible independent record test and no affected/challenger source in the first half of the story.
7. **Assertion testing:** comparison to documents, data, archived pages, filings, policy text, prior conduct, independent experts, or directly contradicting evidence that bears on a material actor claim. Opposing quotations alone are coded separately as response/context, not independent testing.
8. **Documents/records/data, affected-person sourcing, power/history/consequence context, follow-up orientation, and visible correction/update notes.** These follow the original codebook.
9. **Controls:** subdesk, topic, publication month/leadership period, log word count, article type proxy, number of authors, and first author.

Whole-article sentiment is not a primary outcome. A transparent actor-window lexicon estimate will be compared with manual actor-specific judgment and suppressed as a population estimate if reliability is below the threshold below.

## Primary and secondary estimands

**Primary estimand 1: paired scrutiny gap.** Within articles containing claims by both sides, administration assertion-testing/scrutiny score minus challenger score. Positive values mean more scrutiny of administration; negative values are consistent with deference to administration.

**Primary estimand 2: paired reporter-framing gap.** Administration adverse/marked actor-window rate minus challenger adverse/marked rate. Negative values are consistent with the hypothesis; positive values are contrary.

**Primary estimand 3: quotation priority gap.** Administration-minus-challenger difference in first-quote probability and direct-quote-word share within articles containing both actor groups. Positive values may indicate official sourcing priority but are not alone labeled deference.

**Secondary estimands:** official-origin/high-release-reliance rates; independent-document testing; affected-person sourcing; source-type diversity; follow-up; corrections/updates; category-specific challenger comparisons; and topic/subdesk/leadership-period heterogeneity.

## Manual validation and reliability

A deterministic, stratified audit of at least 230 articles (15.1% of the census) will be selected before manual outcomes are examined. Strata cross subdesk, quarter, relevant-topic screen, actor-presence screen, release-origin screen, article-length band, and automated marked-language flag. Routine comparators remain eligible. A fixed SHA-256 ranking seed will break ties.

Manual review codes article eligibility/type and every original actor/source/framing variable that can be judged from accessible text. At least 50 audited articles will be independently double-coded without cross-inspection before adjudication. Report raw agreement and Cohen's kappa for categorical variables; use weighted kappa or intraclass correlation for ordered/count measures. Automated population estimates requiring semantic judgment are suppressed when relevant manual agreement or automated-versus-manual kappa is below 0.60. Mechanical metadata/count estimates may still be reported with validation error rates.

## Statistical tests

- Report counts, denominators, means/proportions, administration-minus-challenger effect sizes, and 95% confidence intervals.
- Use article-clustered bootstrap intervals for actor-unit contrasts and paired bootstrap intervals for within-article gaps.
- Fit adjusted linear/logistic models when outcome variation and effective sample size are adequate, controlling for subdesk, topic, log word count, month/leadership period, article type proxy, and number of authors. Use first-author clustered robust standard errors; sensitivity analyses will use article clustering and author fixed effects where estimable.
- Report unadjusted and adjusted results. Apply Benjamini-Hochberg false-discovery control to secondary outcome families, not to the three declared primary estimands.
- Do not interpret leadership-period coefficients causally or attribute a pattern to a named editor.

## Corrections and updated articles

Score the latest accessible version at cutoff. Preserve visible correction, editor's-note, and update text. Do not infer silent edits. Sensitivity analyses exclude visibly corrected articles and treat visible updates as a control.

## Missing data and failures

`INACCESSIBLE`, `NOT_APPLICABLE`, `NO_SUBSTANTIVE_CLAIM`, `UNCLEAR`, and `NOT_CODED` remain distinct. Missing full text is never coded as neutral, zero sourcing, or no verification. Report missingness by subdesk, date, topic screen, and actor screen. The crawl is a census only if every official archive URL has a successful record; otherwise call it a near-census and state the exact failures.

## Robustness checks

1. Strict versus broad actor dictionaries.
2. Strict reporter-voice marked-language threshold versus adverse-plus-conflict threshold.
3. Exclude quotation text entirely from evaluative windows.
4. Compare all actor-present articles with the within-article paired subset.
5. Exclude breaking briefs, obituaries, event recaps, and stories under 300 words.
6. Exclude the April 2024 encampment/NYPD peak and the 2025 federal-funding peak.
7. Separate Administration, Academics, and Student Life.
8. Separate each challenger category before pooling.
9. Exclude visibly corrected/updated stories.
10. Re-estimate after manual adjudication and after removing automated cases with low confidence.

## Criteria for changing this addendum

Changes are permitted only for a documented access failure, parsing defect, reliability failure, or impossible model assumption. Every change must be timestamped before the affected outcome is recalculated, retain the original specification, state its likely directional consequence, and label post-outcome changes exploratory.

## Interpretation

Evidence will be described as “consistent with” the hypothesis only when the direction is stable across substantively related outcomes and robustness checks. Mixed results will be labeled mixed. A difference in sentiment, quotation order, or release reliance alone will not be called proof of bias.
