Search Everything in One Place

Explore the web, images, videos, news, and more – all in one place.

News

AI companies are still buying up old books by the pallet – then shredding them

AI Companies Are Still Buying Up Old Books by the Pallet – Then Shredding Them
AI Companies Are Still Buying Up Old Books by the Pallet – Then Shredding Them

AI companies are bulk-buying used books, scanning them, then destroying the originals to build clean training datasets - and a court ruling made it legal.

A used bookseller gets an order for 400 volumes. History, botany, regional law, German-language economics. The subjects share nothing in common except one detail: every title has an ISBN. That’s the fingerprint. AI companies - or their intermediaries, shielded by NDAs - are systematically buying physical books from secondhand markets, slicing off their spines for high-speed scanning, and pulping the originals. The reason is brutally simple. The internet is now so saturated with AI-generated text that pre-2022 printed books represent the last large, uncontaminated reservoir of human-authored knowledge, a dynamic that reflects the same appetite for data driving ventures like the Stargate Project.

The Data Wall Nobody Warned You About

Training AI on AI output is like photocopying a photocopy - each generation degrades until the result is mush.

Researchers call it model collapse.” Feed a language model its own synthetic output and quality deteriorates fast. Meanwhile, some writers are deliberately poisoning web content to corrupt scraping pipelines. Old printed books sidestep all of this. Edited by humans, published through traditional workflows, and existing entirely offline, they remain untouched by bots or sabotage strategies. BookData.ai, a firm marketing book-sourced datasets, describes them as “the highest-density source of structured, coherent human thought,” according to its website. For readers curious how these datasets ultimately surface in consumer tools, a look at AI-Powered Websites illustrates the downstream results.

The industrial mechanics are precise:

  • Standard pallets hold 800–1,200 books; buyers scale from pilot orders to 10,000+ volumes per batch, with destructive scanners processing 80–120 pages per minute after spines are cut and originals pulped afterward
  • Buyers deploy AutoBuy flags on ISBN lists through platforms like Alibris and Biblio - picture a Spotify playlist auto-adding tracks, except the vinyl gets shredded at the end
  • Anthropic‘s “Project Panama” spent tens of millions on this exact pipeline, using contractor Datamation to scan books for Claude’s training data, with booksellers identifying these buyers by abnormal volume, subject-agnostic orders, and total indifference to pricing, according to court documents reported by the Washington Post

One Court Ruling Changed Everything

A federal judge called the buy-scan-destroy pipeline “clearly transformative” - and other labs immediately took notes.

In Bartz v. Anthropic, Judge William Alsup ruled that purchasing a book, scanning it, destroying the physical copy, and keeping only the digital version for AI training qualified as fair use under U.S. copyright law. What the ruling means in practice is straightforward: buy a book, scan it, destroy it, keep the data. Because the digital copy “replaced” the physical one and was never redistributed as a book substitute, the use was deemed transformative. Anthropic separately settled for a reported $1.5 billion covering roughly 500,000 works, according to Dataconomy - preserving every trained model built on the data.

That ruling handed the entire industry a template. Other labs and AI service companies now explicitly cite it.

Not every player chose the shredder route. Harvard University, partnering with Google and Microsoft, released nearly one million public-domain digitized books in 254 languages - no destruction required. Microsoft’s Burton Davis called starting with public-domain data “prudent,” noting that libraries hold “significant amounts of interesting cultural, historical and language data” missing from online commentary, according to BNN Bloomberg. The existence of that parallel track makes clear that labs have cleaner options. They just don’t always take them.

The EU’s Digital Single Market Directive complicates things further, allowing rights holders to opt out of text and data mining. For older books whose authors are dead or unreachable, though, that opt-out mechanism is essentially decorative.

Every out-of-print volume that disappears into a scanner may be the last copy accessible to anyone outside an AI lab’s servers. Rare, foreign-language, and low-circulation titles - many never digitized elsewhere - vanish from physical existence after a single scan. Who controls humanity’s textual heritage is no longer a question anyone can afford to leave unanswered, especially as the push to build AI Data Centers accelerates the industry’s resource acquisition at every level.

From the coolest cars to the must-have gadgets, GadgetReview’s daily newsletter keeps you in the know. Subscribe - it’s fun, fast, and free.
Read full story on Gadget Review

Related News

More stories you might be interested in.

Norway found a remote kill switch in a Chinese bus. Your car has one too.
The Auto Wire·22 hours ago

Norway found a remote kill switch in a Chinese bus. Your car has one too.

Last summer, engineers working for Oslo's public transit agency drove two electric buses into an abandoned mine and started trying to hack their own fleet. They weren't hunting for foreign spies. They were testing something far more basic: whether the same digital channel that lets a manufacturer beam a software fix to a bus could also work as a kill switch. The answer, they found, was yes. Garage Deals: Nowell Leather’s Hand-Stitched EDC Gear...

Top