Methodology

How Normalize the News decides that an article was revised, and how we label the result.

How we detect a revision

The Revision Tracker is fully automated. No person chooses which articles to watch, decides which changes are worth showing, or edits the result. The same deterministic process runs on every tracked article, so the same inputs always produce the same output.

  1. Discover. We read the public RSS feeds of the outlets we track to learn which articles have been published.
  2. Archive. We ask the Internet Archive to capture each tracked article over time (via its Save Page Now service). The snapshots are stored by the Archive (a neutral, independent third party), not by us. That means anyone can re-check our work against the same public record.
  3. Confirm. A change is only considered once an article has at least two distinct-content captures in the Archive. We compare the two most recent captures. This filters out transient rendering issues, A/B tests, and loading artifacts.
  4. Noise filter. A fixed rule ignores whitespace, boilerplate, and any change below a set character threshold, so trivial churn is never published.
  5. Classify & publish automatically. A mechanical rule (below) tags the change Substantive or Minor, and it is published immediately with links to both public Archive captures. There is no human review step; the trade-off is that we favor transparency and neutrality over editorial curation, and we show our evidence so you can judge every call yourself.

Substantive vs. Minor

Every revision we publish is tagged with one of two labels:

Substantive A change to facts, figures, quotes, attributions, named parties, numbers, dates, or the framing of what happened.

Minor A change to wording, grammar, formatting, headline length, or paragraph order that does not alter the reported facts.

The label is applied by a fixed rule, not a judgment call: a change is “Substantive” if it removes a paragraph or alters at least a set number of characters; otherwise it is “Minor.” These labels describe the size and nature of the textual change only. We do not assess or state a publication’s reason for making a change.

What we do and do not do

We report facts: what the text said before, what it says now, and when each version was captured. We do not draw conclusions about intent, and we do not speculate about why a change was made. Readers can review the side-by-side text and draw their own conclusions.

How a story’s title is chosen

We do not write headlines. When several outlets cover the same event, the story page is titled with one of those outlets’ own published headlines, reproduced word for word, and the page names which outlet it came from so you can click through and check it against the source.

Which one is chosen mechanically, by the same fixed rule every time. Each candidate headline is scored for loaded language: words that characterize a person rather than report what they did, verdict-carrying verbs and puffery, the writer stepping out to comment on the story, conversational openers, exclamation and question marks, copy addressed to you rather than about the event, a second sentence, and quotation marks around a run of one to three words. The lowest score wins, and ties go to whoever published first, so when no headline in the set trips a single signal, which is the common case, the title is simply the earliest one filed.

The rule reads only the words in the headline. It does not know which outlet wrote them, so no outlet is favored or penalized, and the rule moves a title toward a tabloid as readily as away from one when the tabloid files the plainer headline. It is a fixed list of words and patterns rather than a model, so it does not drift, and a reader who thinks a particular call is wrong can point at the specific line that caused it.

The obvious alternative, titling a story with whichever outlet filed first, sounds neutral and is not. On a breaking story the fastest filer is reliably the one writing for attention, and that headline would then stand as the story’s name over every other account of the same event. Stories published before this rule existed keep the title they were given; on those the page says the title came from one of the outlets listed rather than naming one.

How the coverage ledger works

The Coverage-Presence Ledger reports, for one story, which outlets we hold a matching article from. The denominator is a fixed, published set: the outlets recorded in our outlet universe for the NYC local-politics beat, listed by name at the top of the ledger. It is not every NYC outlet, and it is never redefined per story after the fact.

Which articles belong to a story is decided automatically. Each article’s headline and summary are converted to a numeric vector and compared against other articles in a rolling recent window; close enough, and they are grouped. No person decides membership. We do not publish a single similarity number as the site-wide rule, because the setting has been tightened since launch and stories grouped under the earlier, looser setting keep the grouping they were given. Quoting today’s number as though it governed every story on the site would be false.

Because matching is automated, an outlet listed as no match is a statement about our index and nothing more. A piece that scored below the grouping threshold is listed exactly the same way as one that was never written. We do not claim to know which happened, and we do not speculate about why an outlet has or has not published.

Some outlets block automated readers, which is their right and which we respect. From those we hold no articles at all, so we can say nothing about what they covered. Reporting them as absent would be a claim about a newsroom drawn from an empty index, so instead they are reported as not readable by us, named on the page, and excluded from both the covered and the not-covered count. An outlet falls into that bucket when we have ingested nothing at all from it within the recent window the page states.

What we mean by provenance

Provenance is not a rating and not a score. It is the set of signals that already exist around a piece of reporting, surfaced in one place so a reader can look at them directly. The Provenance page reports three of them per outlet, and they come from two different kinds of place, which is worth keeping straight.

Two are counts from our own database. Articles tracked is how many URLs from that outlet the Revision Tracker is watching. With an archive capture is how many of those the Internet Archive has captured at least once, which is what makes a later edit provable rather than asserted. Revisions published is how many edits we have reported and still stand behind. A revision we published and then withdrew is excluded, because a withdrawn finding is not a finding.

One is a lookup against a dated file. Two public programs, the Journalism Trust Initiative and the Trust Project, publish lists of the newsrooms that have joined them. We transcribed those lists by hand on a stated date and match an outlet to them by domain. This is deliberately not a live query: we would rather show a reader a standing with a date attached than a standing that silently changes.

The distinction that matters most on that page is between an outlet we have no record of and an outlet we do have a record of that shows no membership. Only the second one is a lookup anybody performed. Neither is a criticism. Both registries skew national and international, neither is a prerequisite for good local reporting, and joining one is a paperwork exercise a small newsroom may reasonably never get to.

Three further signals appear only in the live check on that page, where you paste a link and we fetch that page and look at it right then: whether the article’s media carries a Content Credential (C2PA), and whether the outlet put an AI-disclosure or a sponsored-content label on the page. Those results are computed on the spot, labelled unreviewed, and never stored, so they are not reported per outlet anywhere. Our Content Credential reader in particular is not finished, and until it is, its answer is a fact about our reader rather than about anyone’s photograph.

How the headline check works

The check on Headline-vs-Body is a word-overlap heuristic and nothing more. No language model is involved. It strips the headline down to its key terms, drops stopwords and filler, splits the article body into sentences, and picks the single sentence covering the most of those terms. The fraction covered is the score.

The score maps to a status through cutoffs published in the code: at or above 0.7 the body sentence is reported as supporting the headline, below 0.35 the claim is reported as not found in the body, and everything between is partial. Confidence is a separate number answering a separate question: how much signal the heuristic had, not how supported the claim is. A headline yielding fewer than three key terms, or a best sentence matching fewer than three of them, is capped at low confidence, and a low-confidence result has its status withheld. You still see the matched sentence, because the sentence is the part you can check yourself.

Word overlap is not entailment. A sentence can repeat every word in a headline and still contradict it, and a body can support a headline in words the headline never used. That gap is why the matched sentence is always shown and why the result is phrased as what the body text does or doesn’t say, never as what the outlet meant. We do not label a headline misleading, and we make no claim about anyone’s intent.

The check is run by you, on a link you paste, and NTN stores no part of it. That also means NTN publishes no headline findings: the four examples on that page were written by us to show the four states the interface can produce, including the withheld one. They are not results, no real outlet appears in them, and they are labelled as invented where they sit.

How the quote-omission diff works

The diff on Quote-omission is deterministic string alignment. No language model is involved and nothing is fetched at read time. A published quote is split into fragments wherever the outlet used an ellipsis, each fragment is located in the full source passage by exact search after case, curly quotes and whitespace are folded, and the search only ever moves forward, so the fragments stay in the order the outlet used them. What sits between two located fragments is what was cut.

If any fragment cannot be located, we show no comparison at all rather than a partial or guessed one. There is no fuzzy fallback. Sentence counts on an omitted span are an approximation from splitting on sentence-ending punctuation, which is why they are phrased as approximate where they appear.

Trimming a quote is usually legitimate. Outlets cut for length, to remove filler, or to stay on topic, and most cuts change nothing about what was said. We do not allege intent and we do not call a cut misleading. We show what sits between the quoted portions and link the source so the judgment is yours.

The comparison only works against sources we can quote in full, so it is currently limited to public-domain U.S. government and federal-court material: White House remarks and briefing transcripts, the Congressional Record, and federal court opinions. In each example on that page the full passage, the named official or judge, and the source link are real and verbatim. The published quote is not: it is a reconstruction we wrote to show how an outlet might trim that real passage, the outlet name beside it is a neutral placeholder, and no real publication is depicted as having run it. That is stated on the page itself as well.

Images and attribution

Where a story shows a photo, it is the outlet’s own social-share image, the picture the outlet publishes in its page metadata (the og:image) precisely so that link previews can display it. We load that image directly from the outlet, at thumbnail size, always labelled with the source (“Photo: [outlet]”), on a card that links back to the original article.

We do not copy, re-host, crop, or store anyone’s photograph, and we do not lift images out of an article’s body. When an outlet publishes no share image or when its share image is an obvious wire-service or stock photo (Getty, AP, Reuters, and the like), we show our own generated cover art instead of that image. The result is that every visual on the site is either the outlet’s own attributed, linked share image or original artwork we made.

If you are a publisher and would prefer we not display your share image, tell us and we will remove it promptly. See Corrections for how to reach us.