NYT vs. OpenAI and Microsoft: Paywall Bypass Claims and the Fight Over News Data

Unredacted filings in The New York Times copyright lawsuit against OpenAI and Microsoft allege the companies bypassed paywalls, scraped articles at scale and stripped copyright notices while privately acknowledging harm to publishers.
The Times sued on December 27, 2023, accusing OpenAI and Microsoft of using millions of its articles without permission to train generative models, in a case that centers on unlicensed training under U.S. copyright law. Reuters
The new material surfaced in filings unsealed ahead of September 17, 2026. According to the Times brief described in press reporting, a top Microsoft executive privately described the AI training practices as "theft." OpenAI leadership used similar internal language, saying its models posed an "existential threat" to publishers and journalists whose work trained them. TechCrunch
The filings allege OpenAI and Microsoft obtained Times content by bypassing paywalls undetected, building training datasets through mass scraping, and deliberately stripping copyright notices from training data.
The broader context here is why that distinction may matter to the court. Copying openly available web text is treated differently from circumventing access controls and removing copyright management information, which touches on willfulness and on DMCA-related questions in addition to fair use.
Much of the new information comes from the Times' own brief, not from the underlying exhibits, which remain sealed. The allegations summarize internal documents and testimony as characterized by one party in contested litigation.
What the internal communications allege
The brief quotes OpenAI Head of ChatGPT Nick Turley writing internally that publishers face an "existential threat" from products like the chatbot, which are "largely substitutive" and will get more substitutive as they improve. OpenAI President Greg Brockman is quoted separately describing the models as "excellent at news."
Microsoft Chief Executive Satya Nadella was deposed in 2026. The brief states he testified that anything paywalled should be licensed by anyone who wants to use it for grounding or training. It further states he testified he would have invoked Microsoft's right to require OpenAI to retrain its models if he had learned OpenAI scraped and trained on paywalled information.
Grounding, the use of retrieved source material at answer time to constrain generation, is distinct from pretraining on data beforehand. Think of pretraining as studying textbooks in advance, while grounding is like checking a source during an open-book test. Nadella's stated position groups them together for paywalled content. Both require a license.
Grounding, traffic and the doom loop
The filing cites Microsoft's own data to put a number on substitution. That data indicates its Copilot "answer engine" caused click-through rates for the New York Times domain to fall by as much as 93% compared with traditional Bing search.
Microsoft Director of Applied Science Brent Hecht addressed that decline in a January 2024 internal presentation. He described the traffic decline as a "doom loop" that would hurt model performance and the entire web. If answer engines satisfy queries without outbound clicks, publishers lose the traffic that funds original reporting. Over time the supply of fresh, high-quality tokens for future training and retrieval degrades.
The exhibits remain sealed, so methodology, baselines and causation behind the 93% figure have not been publicly tested.
In my view, the source of the number still matters for the case. Because the metric comes from Microsoft's own measurement, as characterized by the Times, it is likely to carry weight in the summary judgment record.
Where the cases stand
Judge Sidney Stein is presiding over competing requests for summary judgment in the New York Times case against OpenAI and Microsoft. The docket, identified as The New York Times Company v. Microsoft Corporation, reflects active motion practice into September 2026, including a notice of OpenAI's March 10, 2026 demonstratives regarding its response in opposition to plaintiffs' motion for protective order, and an order directing plaintiffs to file their reply in further support of their motion for sanctions on or before September 10, 2026, subject to a word limit. CourtListener
The dispute extends beyond one publisher. The Seattle Times and Newsday sued OpenAI and Microsoft alleging copyright infringement, with allegations that the companies scraped the newspapers' websites, including content behind paywalls. Reuters
In early September 2026 the Trump administration filed a brief in defense of OpenAI's unlicensed use of copyrighted material to train large language models. That filing does not decide the Times case, but it signals how the executive branch wants the court to weigh training use, innovation and publisher harm.
The broader context here is familiar to anyone who built on search and social distribution. Publishers traded access for traffic, then watched the terms change as aggregation moved up the stack. Training corpora repeated the pattern at larger scale, and inference-time answers are now repeating it again by keeping users inside the model response.
In my view, the lasting question for technologists is not whether models need high-quality news tokens. They do, for pretraining, for fine-tuning and for grounded generation. The question is what licensing and attribution plumbing keeps that supply sustainable. Getting around paywalls, bulk scraping and stripped notices are weak foundations for production systems. Licensed grounding with click-through, citation and revenue share takes longer to negotiate but is easier to defend and maintain. It is also worth noting that if courts narrow fair use for paywalled training data, retraining duties and retrieval controls become operational risks, not only legal theories. Teams need to know which sources sit in their corpora, how they were acquired, and how they would be removed.
Over the long arc, there is still reason for optimism. My children grew up moving from encyclopedias to search to chatbots with little sense of loss, only convenience. The task now is to preserve the economic loop that pays for the reporting they were summarizing. Models that license, cite and send users outward will have fresher data to work with than models that do not.


