this post was submitted on 28 Jul 2026
625 points (98.3% liked)

Technology

86729 readers
4742 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
 

Artificial intelligence labs are in a new arms race to buy up millions of rare books, slicing them open, scanning the pages and pulping the remains — sparking concerns that the last remaining copies of out-of-print texts are being destroyed on an industrial scale.

ISBNdb notes that “print books from the pre-LLM era are structurally guaranteed to be free of this contamination”.

“Millions of the most valuable books have never been digitised. They exist only in physical form, scattered across library shelves, used bookstores, and out-of-print catalogues. We get them to you at scale.”

you are viewing a single comment's thread
view the rest of the comments
[–] unwarlikeExtortion@lemmy.ml 15 points 23 hours ago (1 children)

To be honest, I'd be MUCH less against this practice if this shredding at the very least included saving the original scans at at least 300 dpi for everyone to access, freely.

Ideally, they wouldn't cut them by the spine, but even giving the despined ones to an archive/library would be just above the unpermissible line.

Despining, scanning and public access archiving is borderline permissible.

[–] JackbyDev@programming.dev 9 points 22 hours ago (1 children)

Also, if every AI company does this, then it's more books gone. If they cooperated they'd only need to do it once. Good scans of all books in the public domain would be so useful for us all as society.

[–] Armok_the_bunny@lemmy.world 2 points 6 hours ago (1 children)

While that would be nice, unfortunately it would almost certainly run into the very same legal issues that are also causing them to pulp the books after scanning them.

[–] JackbyDev@programming.dev 2 points 2 hours ago (1 children)

What issues specifically? Not disagreeing, trying to understand what you mean.

[–] Armok_the_bunny@lemmy.world 2 points 2 hours ago (1 children)

IIRC destruction of the physical copy to prevent resale while still possessing a copy is required for book scanning to be free use, or otherwise not encounter problems with copyright law. Proceeding to share/distribute those scans, or especially trying to charge money for them, would encounter those exact same copyright barriers.

[–] JackbyDev@programming.dev 2 points 1 hour ago

Oh, maybe. Even then, for things in the public domain this shouldn't matter. I know that's probably not the majority of what they're scanning though.