Technology

AI Companies Are Destroying Rare Books to Feed Their Models

Martin HollowayPublished 2w ago5 min readBased on 5 sources
Reading level
AI Companies Are Destroying Rare Books to Feed Their Models
Photo by KoolShooters on Pexels

Amazon is cutting the spines off rare books and scanning them to obtain training data for its AI models, according to an investigation by 404 Media published August 17, 2026 TechCrunch.

The investigation used a tracking device concealed inside a rare book, which ultimately arrived at an Amazon facility in Las Vegas identified as VGT3. The facility marks itself with a logo of a dinosaur clutching a book in its claws TechCrunch. Amazon told 404 Media in a statement that it "purchases books through commercial channels to improve the products and services customers use" TechCrunch.

The practice is driven by a specific technical need. Rare books that are out of print or otherwise unavailable on the internet are valuable training data for large language models precisely because they contain text the models have not already ingested. Texts published before 2022 are especially prized: there is no chance they were written by an LLM, which means they are free of the synthetic content now flooding the web, commonly called "AI slop" 404 Media.

The motivation connects to a problem called model collapse. When a model trains on too much AI-generated text, its output quality degrades — errors compound and the range of what the model produces narrows. As the web fills with machine-written content, training on it repeatedly makes the problem worse. Pre-2022 printed books, by definition, sit outside that feedback loop.

Amazon is not the only AI firm whose destructive book-scanning practices have come to light. Court filings reported by the Washington Post in January 2026 revealed that Anthropic destructively scanned millions of books, including by buying, scanning, and disposing of them, to build its AI models Washington Post. A lawsuit filed in summer 2025 first brought Anthropic's mass destruction of print books into public view Ars Technica.

The supply chain is broader than any single company. AI firms are quietly bulk-buying rare books and destroying them, and secondhand booksellers have begun pushing back. In early August 2026, Australian secondhand booksellers raised alarm over what they described as the "horrific" destruction of rare titles flowing into the AI supply chain The Guardian. Booksellers in Australia and elsewhere believe their stock may have been caught up in a pipeline that sees old books scanned and then destroyed The Guardian.

What emerges across these reports is a supply chain with several identifiable stages. AI companies acquire books through commercial channels, including bulk purchases from secondhand sellers. The books are transported to scanning facilities where their spines are cut to enable fast, high-volume digitization. The physical copies are then destroyed. The resulting text corpus feeds model training. Each stage is invisible to the original sellers unless they take active measures, as 404 Media did with its tracking device.

The irony is sharp enough to require no embellishment. Amazon began as an online bookseller. It now destroys books to fuel a different kind of commerce entirely. Anthropic, whose stated mission centers on responsible AI development, was outed for the same practice through litigation. And booksellers, the original custodians of the physical inventory now being consumed, are often unaware their stock is entering a destructive pipeline at all.

The broader context here is the intensifying scarcity of clean training data. LLM developers have already ingested the bulk of the open web. What remains accessible and high-quality, particularly text guaranteed to be human-authored, is increasingly locked behind paywalls, embedded in physical artifacts, or both. Destroying rare books to extract that text is the logical, if troubling, endpoint of the current approach to data acquisition. Non-destructive scanning exists and is well established in archival practice. It is slower and more expensive.

Worth flagging is the asymmetry of the transaction. A secondhand bookseller sells a volume for what the used-book market will bear. The buyer extracts text that may contribute to a model worth billions. The physical artifact, potentially irreplaceable, is gone. The seller has no knowledge of the downstream use and no share in the value created. Whether this asymmetry rises to the level of a regulatory concern is a question policymakers have not yet visibly engaged with, though the pattern of litigation suggests courts will reach it first.

The AI training data pipeline is, at root, a resource extraction operation. We have watched this industry cycle through successive resource constraints: compute, then talent, then attention. Clean, human-authored text is the current bottleneck, and firms are solving it the way extractive industries always do, by acquiring the raw material as cheaply as possible and consuming it. That the raw material happens to be rare printed books, cultural artifacts that may not exist in any other form, gives the story a particular edge. It also, in my view, raises a question the industry has yet to answer with any seriousness: what happens to the artifacts, and the information ecology they represent, when the data has been extracted and the physical copies are gone?

Model collapse provides the technical pressure. Profit provides the commercial one. What is missing is any counterpressure, regulatory or normative, that would require preservation of the source material. Until that exists, the destruction will continue at whatever rate the market for training data demands.