Technology

USA Today Sues OpenAI Over News Data for AI Training

Martin HollowayPublished 29m ago4 min readBased on 9 sources
Reading level
USA Today Sues OpenAI Over News Data for AI Training
Photo by imgix on Unsplash

USA Today Co. is suing OpenAI, alleging the company copied hundreds of thousands of its articles to train AI models without authorization. The publisher seeks more than $250 million in damages The Verge.

The price tag is explicit. The core claim is simple. OpenAI never asked for permission to use the content, according to the complaint.

The suit covers more than the flagship title, including The Tennessean, Indy Star, The Columbus Dispatch and The Oklahoman. The allegation is not about a single crawl or a single output. It describes systematic intake of a large, continuously updated news collection into the training systems for general-purpose models.

USA Today Co. is not litigating alone. The New York Times sued OpenAI and Microsoft in 2023, accusing the companies of using millions of newspaper articles. Microsoft is OpenAI's largest financial backer Reuters.

A separate action from The Seattle Times and Newsday also names OpenAI and Microsoft as defendants for alleged copyright infringement. It was filed in the U.S. District Court for the Southern District of New York and alleges the companies scraped the newspapers' websites Reuters.

The docket keeps expanding. More than 30 U.S. local newspaper groups that together own about 400 titles have sued OpenAI and Microsoft over unauthorized scraping and use of their content Press Gazette. An earlier eight-newspaper action accused OpenAI, the maker of ChatGPT, of copyright infringement, saying the technology companies took millions of copyrighted news articles without permission or payment.

The dispute extends beyond daily news. Encyclopedia Britannica alleged OpenAI unlawfully copied nearly 100,000 of its articles to train GPT large language models. Major publishers have separately alleged Meta used millions of books and articles without permission to train its Llama AI model. Authors suing Microsoft over a model allegedly trained on pirated books sought a court order blocking infringement and statutory damages of up to $150,000 for each work allegedly misused.

Across these cases, the technical questions are consistent. How the text collection was acquired. Whether web-scale crawling, organizing the data, breaking text into tokens, the small chunks a model processes, and updating weights, the internal settings, counts as reproduction or transformation under copyright law. Where the line sits between ingesting data during training and repeating it from memory at answer time. Those questions turn not on model design alone but on provenance, logging and control, in other words on records showing where data came from and how it was handled.

The broader context here is that courts are being asked to define rules for the data supply chain behind foundation models, the large base models that support many AI tools. Publishers argue for prior consent and compensation. Model developers have generally pointed to fair use, the legal rule that allows limited use of copyrighted work, and to the transformative nature of training.

In my view, that either-or framing hides the more durable engineering and business problem.

For context on how we got here, the web was built for indexing and retrieval, with robots.txt, a file that tells automated crawlers what they may copy, and terms of service as imperfect proxies for permission. Large-scale training broke those assumptions. It requires machine-readable rights signals, durable provenance records and licensing systems that work at crawl scale, not through one-off negotiations.

Looking at what this means for builders and buyers, the near-term risk is less about any single verdict than about fragmentation. Different publishers, different jurisdictions and different licensing demands will apply. Enterprise teams already manage data governance, model cards that document how a model was built, and indemnification clauses that assign legal risk. That discipline will now have to extend upstream to training data sourcing and downstream to output filtering.

The long-term picture here remains hopeful. Structured access to high-quality editorial, reference and book content would improve models. It would reduce hallucination, when models invent false facts, improve citation and create a revenue path for original reporting and research. My own children grew up moving from printed encyclopedias to Wikipedia to chatbots, and each shift made information easier to get while making its source harder to see. Fixing that provenance layer is where the real opportunity lies.