Technology

The Seattle Times and Newsday Sue OpenAI and Microsoft Over AI Training Data

Martin HollowayPublished 59m ago6 min readBased on 11 sources
Reading level
The Seattle Times and Newsday Sue OpenAI and Microsoft Over AI Training Data
Photo by Brett Sayles on Pexels

The Seattle Times and Newsday filed a federal lawsuit against OpenAI and Microsoft on September 4, 2026, accusing the companies of scraping the newspapers' websites for AI training data without permission or compensation Reuters.

The 38-page complaint was filed in the U.S. District Court for the Southern District of New York, in Manhattan Newsday. It alleges that OpenAI and Microsoft's AI models violate the news organizations' copyrights. The Seattle Times and Newsday are seeking a court order requiring the defendants to destroy copies of their works, as well as any training datasets or AI models that incorporated them Reuters.

The suit follows a similar copyright infringement case The New York Times filed against OpenAI and Microsoft in 2023 The Seattle Times. That earlier litigation has become a bellwether for the broader legal confrontation between publishers and AI developers over training data. The New York Times has spent more than $28 million on its lawsuit since 2023 The Seattle Times.

The legal landscape has expanded considerably since the Times filed. In April 2024, eight U.S. newspapers sued OpenAI and Microsoft for copyright infringement Newsday. In June 2024, the Center for Investigative Reporting, a news nonprofit, filed its own suit claiming OpenAI used its journalism content without permission and without offering compensation Newsday. A federal judge subsequently ruled that The New York Times and other newspapers can proceed with their copyright lawsuit against OpenAI and Microsoft Newsday.

More recently, newspapers involved in the litigation urged a judge to sanction OpenAI, alleging the ChatGPT maker is hiding evidence important to what could be a landmark copyright infringement trial Newsday.

The Seattle Times and Newsday case enters a legal corridor that is already well-traveled. The destruction remedy they seek — models and datasets, not just cached copies of articles — goes to the core question these cases turn on: whether ingesting copyrighted material for model training constitutes fair use, or whether the resulting models themselves are infringing works whose continued deployment cannot be separated from the training data they absorbed.

The financial stakes are sobering. The New York Times' $28 million in legal costs since 2023 is a data point that matters for smaller and regional publishers weighing whether to file or join consolidated actions. The Seattle Times and Newsday are not small outlets, but neither do they have the Times' resources. The cost of litigating against two of the best-funded technology companies in the world, over a period that could extend years, is a material constraint on who can afford to defend their copyrights through the courts.

The sanctions motion from the existing newspaper plaintiffs adds another dimension. Allegations of evidence concealment, if sustained, could affect the discovery process — the pre-trial phase where each side must share relevant evidence — for all subsequent plaintiffs, including The Seattle Times and Newsday, by establishing patterns of conduct that new filings can reference. The Southern District of New York, where this latest suit was filed, is also where the Times case is proceeding, which means the same bench and local rules will govern procedural questions.

The remedy sought by the plaintiffs deserves attention. An order compelling destruction of trained models would be unprecedented in scale and would raise practical questions about model provenance, derivative works, and the feasibility of isolating specific training contributions within a large language model's learned weights — the numerical parameters a model develops during training. Courts have not yet grappled with the technical reality that training data cannot be cleanly excised from a deployed model the way a file can be deleted from a server. How the Southern District handles this question, if the case reaches the remedies phase, will set a precedent that extends well beyond the parties involved.

For now, the filing adds two more prominent news organizations to a growing plaintiff pool that now includes individual newspapers, regional chains, and investigative nonprofits. The legal theory across these cases is consistent: unlicensed scraping of published journalism for commercial AI training constitutes copyright infringement, and the fair use defense that OpenAI and Microsoft are expected to raise does not apply when the output competes with or substitutes for the original works.

OpenAI and Microsoft will have the opportunity to respond to the complaint. The case will proceed through motions to dismiss and discovery, stages at which the earlier-filed cases have already established some of the procedural and substantive contours the parties will navigate.