News: 0184567108

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop (404media.co)

(Tuesday July 21, 2026 @05:00PM (BeauHD) from the model-collapse-avoidance dept.)


An anonymous reader quotes a report from 404 Media:

> As AI companies search for more training data to improve their models, one company is [1]offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing. "The world's best AI training data is sitting on a shelf," ISBNdb, a company that produces what it claims is "the world's largest book database," and that offers high-volume book acquisition services for AI companies, says on [2]its site . "Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative."

>

> In [3]one article on its site , ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don't include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in "model collapse," a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models.

>

> "Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage [...] "Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools." [...] ISBNdb advertises that it can keep the identity of AI companies secret. "Strict NDA [non-disclosure agreement] on every engagement," ISBNdb's site says. "Every project begins with a legally binding non-disclosure agreement. Your identity, strategy, and acquisition targets are never disclosed." ISBNdb notes that AI companies may not want to be caught destroying printed books during the scanning process. "The optics problem is real," ISBNdb's site says. "'AI company destroys two million books' is not a headline that generates sympathy."



[1] https://www.404media.co/ai-companies-are-buying-tons-of-old-books-because-theyre-free-of-ai-slop/

[2] https://isbndb.com/

[3] https://isbndb.com/blog/ai-training-data-poisoning/



Rainbows End (Score:5, Informative)

by TwistedGreen ( 80055 )

Yeah, it's happening.

Not the best idea (Score:2)

by CEC-P ( 10248912 )

Every modern AI engine has trouble with versioning and outdated info being replaced with newer info. I've had multiple say "oh, that's the new interface. It's here now" about once a week. Scan books from like 1950 and you'll probably get "There actually is no subatomic particle smaller than the proton" and "Actually, women shouldn't work on office environments or their ovaries will explode."

Re: (Score:1)

by Anonymous Coward

The idea is to have both. And the training is (partially) split into knowledge (or generally base training) and style training in the post-training step. So you can have a LLM describe a modern computer UI without saying delve into and "why it matters" a hundred times if you train the style on AI style free writing.

Re: (Score:2)

by karmawarrior ( 311177 )

That doesn't make much sense either, the further back, the difference in the very way English is spoken/written.

Honestly this, again, means I get to raise a point that I never got a good reply to at the beginning of this LLM BS: what is the incentive for creating new, clean, correct, works at this point? Because "People always want to argue about shit on Reddit" isn't going to get us new information on the majority of topics with the possible exception of politics. And let's be honest, that's not a thing wh

Re: (Score:1)

by Bentbob ( 1081243 )

If I had a huge pile of money to burn, and wanted to be more of a AI goblin than a tech bro, I'd scan only golden and silver age Science Fiction novels, novellas, short stories, etc, and feed it into a LLM... to do something. Maybe try to get it licensed and implemented into work software (or would it be "platforms"? big tech love their "platforms" these days)

Isaac Asimov's "The Feeling of Power" has felt so much more realistic these days. A couple of decades ago I would have gone that the scenario was very

Re: (Score:2)

by gweihir ( 88907 )

Yep. In fact a pretty bad idea. But I think they are getting desperate.

Fact Check (Score:2)

by SlashbotAgent ( 6477336 )

Companies are offering books for sale.

Book sellers are selling books.

The sellers seem to have no idea if the buyers are using it for AI training, just that the sales volumes have increased.

Anecdote: I see an increase in book stores. AI also see in increase in youths reading/collecting physical books at a higher rate than recent years. YMMV

Re: Pre-2019 for AI, pre-2010 for Woke (Score:1, Flamebait)

by greytree ( 7124971 )

There was woke before 2010, but we called that nonsense Political Correctness and had a good laugh at it.

Then they made it law.

$3000 per book (Score:4, Interesting)

by silentbozo ( 542534 )

[1]https://techcrunch.com/2026/07... [techcrunch.com]

"The payout will deliver $3,000 per work across an estimated 500,000 works, shared among the authors and publishers who hold rights to them. While the settlement is believed to be the largest in the history of U.S. copyright law, many authors and creators still donâ(TM)t view it as a win.

Thatâ(TM)s because of how the legal question was resolved. Alsup sided with Anthropic on the core issue. He ruled that training an AI model on copyrighted text counts as fair use â" a decision widely seen as a turning point for the AI industry. But the ruling didnâ(TM)t excuse how Anthropic obtained the books in the first place. Anthropic had built its training library from two sources: books it purchased and scanned (fine), and books it downloaded from pirate sites like Library Genesis and Pirate Library Mirror. Alsup found the second method illegal on its own terms and said that piracy question could go to trial; Anthropic agreed to a settlement soon after to avoid a trial and whatever damages a jury might have awarded."

So... it's legal to scan books you own and then use them to train LLMs, but it's not legal to use scans that someone else made (I'll assume in this case, they didn't own the books in question.) Hence... the perverse incentive to buy and re-scan books that might already have been scanned... and the cheapest way of doing it is to chop the spine off.

"Internal Anthropic documents about its plan to scan millions of books, revealed in the copyright lawsuit, donâ(TM)t make clear why the company wanted to destroy the books in the process. A deposition of Tom Harvey, who Anthropic hired to lead the project and who previously helped create Google Books, shows that one company Anthropic contracted to scan the books was Datamation, which offers both âoehigh volume destructive and non-destructive book scanningâ services. In a destructive book scanning process, the spine of the book is cut so the pages can be fed into a scanning machine, which is faster and cheaper than non-destructive book scanning.

Regardless of its original intentions, the federal judge in the copyright lawsuit from authors against Anthropic, William Alsup, found that Anthropicâ(TM)s creation of digital copies of the books was legal specifically because the books were destroyed.

âoeHere, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy,â Alsup wrote in his ruling. âoeThe print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company.â "

Kind of fucked up that the scan can't be shared (or donated). I imagine in most cases, good copies of these books no longer exist in libraries or in the Library of Congress. With current law, all books that are covered under copyright will eventually fall into the public domain, but this is meaningless unless copies exist for people to redistribute once that limit is reached. Essentially companies are exploiting the monopoly benefit extended through copyright without allowing society to benefit from the material falling into the public domain, which is the implicit contract to using state power to enforce copyright.

Ironically, destroying physical copies in order to comply with the 1:1 rule makes the remaining copies that much more valuable.

[1] https://techcrunch.com/2026/07/20/anthropics-landmark-1-5b-copyright-settlement-is-approved/

Destroying more books than the Nazis (Score:2)

by xack ( 5304745 )

All because you can't eat your own slop. Tobacco company executives don't smoke either.

Publishers prefer destruction (Score:2)

by greytree ( 7124971 )

Publishers prefer book destruction to reasonable copyright terms.

That's why we, the people, must remove their copyright privilege after five years.

Copyright EXISTS to encourage creation not destruction.

Dear ChatGPT, I feel so lazy today... (Score:3)

by ffkom ( 3519199 )

... what ailment may have befallen me?

ChatGPT: According to my knowledge of the latest advancements in [1]Humorism [wikipedia.org], an over-abundance of Black Bile is probably the cause. Thy shall seek to get some cupping done. I learned that from a book.

[1] https://en.wikipedia.org/wiki/Humorism

A BIG assumption that those books aren't slop too (Score:1)

by SmaryJerry ( 2759091 )

There is AI slot and also hand written slop. A book being old does not make it peer reviewed, although it is for sure curated by the publisher. There are also countless non-fiction books that were written at a time when the 'ground truth' was completely different or not as settled. For example: cocaine was prescribed in the past.

Re: Model collapse is a transient phenomenon (Score:2)

by presidenteloco ( 659168 )

When AIs get unarguably smarter and more knowledgeable than even erudite humans and editors, and human domain specialists, as will inevitably happen, then model collapse should no longer be a thing, as long as there is diversity of models feeding off EACH OTHERs' outputs rather than all feeding off their own outputs. A lack of diversity of models could lead to semantic "inbreeding" and model collapse, but a diversity of superhuman intelligent and knowledgeable AI systems should produce superhuman-quality o

Re: (Score:2)

by gweihir ( 88907 )

You apparently have no understood what model collapse is. Here is a hint: It cannot be avoided. The first few papers about it already proved that.

Oh the humanity! (Score:2)

by toxonix ( 1793960 )

Someone's going to feed it a hardcopy of "AI For Dummies" and THEN we'll be sorry. The singularity will be upon us.

Re: (Score:2)

by 93 Escort Wagon ( 326346 )

> Someone's going to feed it a hardcopy of "AI For Dummies" and THEN we'll be sorry. The singularity will be upon us.

Maybe we'll get lucky and that'll make Sam Altman's head explode!

Looks like they are getting desperate (Score:2)

by gweihir ( 88907 )

That stunt only works once. Everything to keep the hype going a bit longer, I guess.

Booths for two or more.