Misinformation Never Closed Its Door
Reputable news sites block AI crawlers at 60 per cent and misinformation sites at 9.1 per cent, which leaves retrieval reading whatever had the least reason to close the door.
In October 2025 three researchers counted who on the web shuts the door on AI crawlers, and had the good sense to count the other side of the room as well. Across a sample of reputable news outlets, 60.0 per cent disallow at least one AI crawler. Across misinformation sites, 9.1 per cent do. Same file, same four lines of plain text, read against two populations: on one side nearly two thirds barring the way, on the other nine in ten leaving it open. Steinacker-Olsztyn, Gosain and Dao put the finding in their title, as a question: Is Misinformation More Open?
The door closed quickly, and only on one side of the room. In September 2023 the figure for reputable outlets stood at 23 per cent; by May 2025 it was approaching 60. Twenty months. Over the same stretch the other population barely moved, because a site that lives on traffic and amplification has no reason to turn away a reader, least of all one that never asks for a subscription and never leaves the page.
One file makes the mechanism concrete. Consumer Reports has been testing products since 1936, and its robots.txt today carries five full prohibitions: GPTBot, Google-Extended, CCBot, meta-externalagent, meta-webindexer. ClaudeBot and PerplexityBot are not named anywhere in it, so they walk in. Ninety years of comparative testing are shut to OpenAI, to Google, and to the public corpus that almost every smaller lab trains on, while remaining open to two others. Nobody declared a principle here. Five separate decisions were taken about five counterparties, each with its own history of negotiations, lawsuits and legal exposure.
Consent decides the corpus. When a model with browsing turned on answers your question, it is reading the list of sites that agreed to be read, and agreement is not distributed at random. It tracks having something to lose, having a legal department, and above all having a second channel in which the same text can be sold at a price.
The Wall Street Journal's robots.txt shows what that second channel looks like once someone writes it down. The default rule denies everything to everyone. Beneath it sits a named list of agents that are allowed through, and OpenAI's three sit in it: ChatGPT-User, GPTBot, OAI-SearchBot. Anthropic's crawler is absent from that list, so is Perplexity's, so is Common Crawl's, and the default catches them. In May 2024 News Corp signed a licensing agreement with OpenAI reported at more than $250 million over five years. The permission file is the commercial agreement, rendered machine-readable, and the paid lane is what emptied the free one.
Any operations team knows this failure of measurement: the inventory that counts only the machines answering the scanner. The number you get is true, and it is true of a population that is not the one you meant to measure. What makes this case worse is that the non-response correlates with the very variable you care about, reliability, and it correlates in the wrong direction. The sample is oriented rather than noisy. It also rests on a layer that is purely declarative, since every machine announces its own name and nobody verifies it.
Then again, the outlets that block have every reason to block. A robots.txt file is the only leverage available to a publisher who cannot sue everyone, and the researchers never suggest that the ones barring the door are in the wrong. Defending your own work is the sensible move. The systemic result simply does not depend on anyone's intentions: it emerges from a sum of individually reasonable choices, taken by the parties with the most to protect, inside a mechanism that records no reason for refusal and can count only whoever stayed.
A second stage sits on top of this one, and Deana Burke named it well on Boys Club: post hoc citation. The model writes the answer first, then goes looking for a source that supports it. Taken alone that is procedural laziness. Taken on top of a pool already filtered by consent, it turns into a confirmation machine, because the answer gets composed and the evidence gets fetched from the one shelf still open, which belongs to whoever had nothing to lose by staying open. Consent decides the corpus, and the citation arrives afterwards to ratify it.
Clicking the links a machine hands you accomplishes little if all you check is that the page loads. Look at the domain, then ask the question the interface will never prompt: which sources are absent from this answer, and what do the absent ones have in common. Most of the time they have a legal department and a licensing contract. The quality of a generated answer now turns less on how good the model is and more on who had a reason, that particular morning, to leave the door open.