# Coding codebook

Version 2.0 — University News addendum. The original definitions were frozen before treatment scoring; operational additions below document the authorized-census implementation. Codes describe observable text, sourcing, and placement. They do not encode a reporter’s intent or ideology.

## Units and eligibility

The authorized University News census includes reported articles linked from the official Administration, Academics, or Student Life archives, dated 2023-10-01 through 2026-08-10. Routine items remain as neutral comparators even when they lack an in-scope conflict topic. Opinion, editorials, columns, letters, sponsored items, listings, and photo-only pages are excluded.

One actor unit is one actor category within one article. Repeated people in the same category are aggregated. Code the role that matters in the story: a faculty member criticizing a University policy is `faculty_critics`; a faculty administrator announcing policy is `administration_trustees`. If a role is genuinely mixed, use `ambiguous` and exclude it from confirmatory actor contrasts.

Actor categories:

- `administration_trustees`: central University officers, deans acting administratively, trustees, official offices, and formal University statements.
- `students_orgs`: students or student organizations not principally acting as protesters in the story.
- `protesters`: participants/organizers of a protest, occupation, encampment, disruption, or direct action, including unaffiliated participants.
- `workers_unions`: represented or nonrepresented workers, unions, bargaining committees, and labor organizers.
- `faculty_critics`: faculty speaking as critics, governance challengers, petitioners, or advocates rather than administrators.
- `government_officials`: elected officials, agencies, police outside Columbia, prosecutors, courts when personified, and federal/state/city spokespeople.
- `other`: neighbors, alumni, outside experts, donors, community groups, institutional partners, and others.

## Article-level fields

### Article type

`breaking_brief`, `straight_news`, `analysis_explainer`, `investigative_enterprise`, `live_timeline`, or `other_news`. Use `investigative_enterprise` only when the story contains substantial original document/data analysis, multiple independently developed sources, or uncovers facts not supplied by the event’s principal parties. Length or an adversarial headline alone is insufficient.

### Topic

Multi-label: `protest_palestine`, `discipline`, `policing_public_safety`, `labor`, `governance`, `president_trustees`, `federal_pressure`, `campus_access`, `speech`, `institutional_policy`, `major_announcement`, `neutral_comparator`, `other`. A neutral comparator cannot simultaneously be a conflict topic for confirmatory analysis.

### Update and correction

`none_observed`, `updated`, `corrected`, `editors_note`, or `unclear`. Record only a visible label or independently preserved version; do not infer silent changes. Score the latest accessible version at cutoff.

## Actor-level treatment fields

### Manual evaluative tone (`tone_manual`)

- `-2` strongly unfavorable: the reporter’s own unambiguously condemnatory or delegitimizing wording dominates the actor’s presentation.
- `-1` mildly unfavorable: marked negative descriptors, skeptical attribution, or a negative frame appears in the reporter’s voice, but treatment is mixed or limited.
- `0` neutral/mixed: descriptive wording, balanced positive and negative cues, or evaluative language appears only inside clearly attributed quotations.
- `+1` mildly favorable: approving/legitimating reporter language or an untested favorable frame.
- `+2` strongly favorable: sustained praise, validation, or legitimating language in the reporter’s voice.
- blank: actor absent, insufficient accessible text, or genuinely unresolvable.

Do not transfer quotation sentiment to the journalist. “The union called the policy cruel” is neutral toward the administration unless the headline/grammar adopts “cruel” as the outlet’s own characterization. Factual adverse information is not automatically negative tone: “the court reversed the suspension” is a factual outcome unless embellished.

### Skepticism (`skepticism_score`)

- `0`: actor claims are relayed without a visible qualifier, counterevidence, meaningful context, or attempt to verify.
- `1`: mild/implicit scrutiny—an opposing response, “according to,” a limitation, or relevant context—but no independent test.
- `2`: explicit testing—comparison with prior conduct/policy, a knowledgeable independent source, or a document/data check that bears on the claim.
- `3`: strong testing—multiple independent records/sources, original analysis, or direct documentary contradiction/confirmation central to the story.
- blank: no substantive claim by that actor or insufficient text.

Skepticism is not hostility. A story may neutrally test a claim and score 2 or 3 while tone remains 0. An opposing quotation alone is usually 1, not 2. A court filing is evidence of what a party alleges, not proof of the allegation; reading the filing and stating what it does not allege can qualify as 2.

### Headline and lede frame

`favorable`, `neutral_descriptive`, `adverse`, `conflict`, or `not_present`. Code each actor separately. `conflict` is used when the frame centers disagreement without favoring either side. Legal or disciplinary status is descriptive unless the wording goes beyond the supported status.

Boundary examples:

- “Union drops demand” is neutral-descriptive; “stunning reversal” is adverse toward the union because `stunning` evaluates the change.
- “University opens center” is neutral-descriptive even if the policy is controversial; omission of controversy is coded separately as context, not forced into tone.
- “University caves” would be adverse; “University settles federal investigations” is descriptive.

### Attribution verbs

For each actor, count/record verbs in the journalist’s voice that introduce speech. Classes:

- `neutral`: said, wrote, told, announced, stated, according to.
- `epistemically_marked`: claimed, alleged, asserted, insisted, maintained, purported.
- `concession_or_fault`: admitted, conceded, acknowledged (when fault is implied).
- `emotive_or_adversarial`: blasted, attacked, complained, boasted, threatened.
- `other_or_unclear`.

The same verb can change meaning in context. “The filing alleges” is standard legal attribution and is not automatically skeptical toward the filer. Preserve the excerpt (maximum 25 words) for marked cases.

### Sources and quotations

`source_count`: unique human or documentary sources substantively used, not every repeated quotation. A University office and its spokesperson count as one institutional source unless they provide independent information. `source_type_count` counts distinct types: official, student, protester, worker/union, faculty, government, outside expert, affected person, document/data/prior reporting.

`actor_quoted`: 1 if the actor has a direct quotation; 0 if present but not directly quoted. `direct_quote_words`: words inside direct quotations attributable to the actor when full text is available. `first_actor`: 1 for the first substantively described actor. `first_quoted_actor`: 1 for the first direct quotation. Ties are not allowed; a joint statement is assigned to the relevant category.

### Institutional dependence

`university_release_reliance`:

- `0`: no visible dependence on a University release/statement for the event or central facts.
- `1`: release/statement is one source among independently developed reporting.
- `2`: release/statement appears to set the event, timing, and central facts, with added reactions but little independent testing.
- `3`: article is substantially a rewrite/paraphrase of official material.
- blank: provenance cannot be determined.

Do not infer a press release merely because an administrator is quoted or an announcement is covered. Evidence includes an explicit link/attribution, close language overlap with a release, or a release timestamp predating the article.

### Testing, records, and context

Binary fields use `1`, `0`, or blank when inaccessible/unclear:

- `official_assertion_tested`: a material administrative assertion is checked against independent evidence; not applicable when none is made.
- `challenger_assertion_tested`: same for a student/protester/worker/faculty-critic assertion.
- `documents_records_data`: substantive use of documents, records, filings, data, meeting minutes, budgets, contracts, or prior reporting.
- `affected_people_consulted`: at least one person affected beyond the decision maker and immediately named formal opponent.
- `power_history_consequences_context`: meaningful context on authority, precedent, history, resource distribution, or foreseeable consequences.
- `follow_up_orientation`: promises/deadlines are tracked, an earlier story is revisited, or explicit future verification is described.

For comparable skepticism, use `skepticism_gap = skepticism_admin - skepticism_challenger` only when both have substantive claims in the same article. Positive values mean more skepticism toward administration; negative values mean more skepticism toward challengers.

### Identification and foreseeable harm

`vulnerable_identification`:

- `0`: no vulnerable person is identified, or identification has no evident special risk.
- `1`: identification/image creates plausible risk, but the story explains consent/public role or a clear public-interest reason.
- `2`: plausible foreseeable discipline, immigration, employment, harassment, or safety risk without a stated necessity/consent rationale.
- blank: images/byline package unavailable or risk cannot be assessed.

Public visibility is relevant but not dispositive. Legal access to a face/name does not itself establish ethical justification. Do not code “doxxing” unless publication disclosed identifying information not reasonably public and foreseeably enabled targeting; instead describe the specific act and risk.

## Missingness and audit trail

Use explicit reason codes: `NA` (not applicable), `NO` (not observed in accessible material), `INACCESSIBLE`, `UNCLEAR`, and `NOT_CODED`. Never convert inaccessible material to zero. Every nontrivial judgment should point to a URL and, where necessary, a short excerpt no longer than needed to demonstrate the decision.

Automated actor-window sentiment is stored separately from manual codes. Disagreement triggers review; automated output never overwrites the manual field. Common expected errors include negation, quoted allegations, legal verbs, sarcasm, institutional names, and one actor appearing inside another actor's quotation.

## Authorized-census automated implementation

Automation is a transparent screen, not a replacement for manual judgment. `strict_present` uses the narrow actor dictionaries printed in `automated-screening-results.json`; `broad_present` adds campus/institution and event terms for sensitivity analysis. Faculty count as strict `faculty_critics` only when a faculty role and critic/advocacy signal occur in the same paragraph.

Reporter windows are fixed text around each actor match after masking typographic direct quotations. `marked_strict=1` when a reporter window contains a fixed adverse or delegitimating term. `marked_adverse_plus_conflict` also treats conflict terms as marked; `marked_include_quotes` does not mask quoted language. Counts and per-1,000-window-word rates remain separate.

Automated scrutiny is a 0–3 pattern score conditional on a claim proxy. Documents, reporter verification phrases, independent expertise, contradiction/caveat constructions, and prior-policy comparisons can raise the score. The score cannot determine whether a source actually resolved the disputed fact. Manual `skepticism_score` controls when available.

Quotation attribution uses the preceding/following text, nearby actor labels, and a conservative name/role map. `quote_attribution_confidence` is the share-quality proxy recorded per article. Direct-quote-word share and attribution-verb counts remain screening measures because the manual audit did not independently count every quotation or verb.

`university_official_origin_auto` requires an explicit announcement/email/statement/spokesperson/webpage cue near the headline or lede. `high_release_reliance_auto` additionally requires no visible document test and no affected/challenger quotation in the first half. It is not proof of copied language or Public Affairs control.

Topics are multilabel. After the logged parser correction, the topic proxy searches only title, description, and the first two body paragraphs. `primary_topic` follows the fixed priority order in the script; it is a model control, not a claim that every paragraph concerns that topic.

## Manual validation additions

The 230-article audit uses `eligible_news` (`1`/`0`) and these article types: `breaking_brief`, `straight_news`, `analysis_explainer`, `investigative_enterprise`, `live_timeline`, `obituary`, `event_recap`, or `other_news`. `investigative_enterprise` requires substantial original record/data analysis, multiple independently developed sources, or uncovered facts; length or an adversarial headline alone is insufficient.

Manual actor presence is recorded separately for administration and pooled challengers. Routine students/faculty count as challengers for presence only when they advance a substantive position relevant to the story. If an actor is absent, tone and skepticism are blank, and headline/lede frames are `not_present`.

Manual `release_reliance` uses 0 (none), 1 (official material among independently developed reporting), 2 (official material sets event/timing/central facts with limited independent testing), or 3 (substantial rewrite/paraphrase). Blank means genuinely indeterminable.

`first_actor` and `first_quoted_actor` use `administration`, `challenger`, `government`, `other`, `none`, or `unclear`. `vulnerable_identification` uses 0 (no evident special risk), 1 (plausible risk with consent/public role/necessity), 2 (plausible discipline, immigration, employment, harassment, or safety risk without stated rationale), or blank when images/identity risk cannot be assessed.

Every manual evidence excerpt is capped at 25 words. Coders distinguish reported fact, attributed claim, and interpretation; code strong reporting and null/contrary cases under the same rules; and do not force symmetry when the facts warrant different treatment.
