AI companies are facing a data problem of their own making: the open web is clogged with AI-generated text, and training new models on that sludge invites "model collapse," where machines degrade by eating their own output. One company is filling the need for human-written content by buying old printed books in bulk. As 404 Media reports, a book-database firm called ISBNdb pitches AI labs on pre-2022 print books as pristine training data. "The world's best AI training data is sitting on a shelf," its site says.
ISBNdb will sell scans of 1,000 to 1 million books and promises a "Strict NDA on every engagement" - because, it admits, "'AI company destroys two million books' is not a headline that generates sympathy." The company slices off a book's spine and feeds the loose pages through a machine, the same thing Anthropic did to millions of books. A judge ruled it legal because the paper originals were destroyed.
Booksellers are ambivalent. One told 404 Media his weekly sales jumped from about 20 books to hundreds since April. The buyers show "a total disregard for the price," he said - "a tell for AI because they have just so much money."
Previously:
• Free Software Foundation demands freedom from Anthropic, not money
• New York Times sues OpenAI and Microsoft, claiming copyright infringement
The post AI companies are buying up old books to escape their own AI slop appeared first on Boing Boing.