this post was submitted on 18 Aug 2026
495 points (97.7% liked)

Technology

87434 readers
3650 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


founded 3 years ago
MODERATORS
 

The symbol for Amazon's VGT3, the Las Vegas facility where it scans book for AI training data.

Amazon is buying massive quantities of books, scanning them for AI training data, and destroying them in the process.

A 404 Media investigation was able to reveal Amazon’s book buying operation, which hasn’t been previously reported, by placing a tracking device in a rare book we suspected would be acquired by an AI company for training data, and following it around the country to its final destination.

That final destination was an Amazon warehouse in Las Vegas, Nevada. Amazon employees who work at this location say all they do is receive massive shipments of printed books which they then cut the bindings off in order to scan the books more quickly. The printed book is destroyed in the process. The logo of the Amazon team that works at this warehouse, called VGT3, is a dinosaur, brandishing its teeth and with a book in its hands.

“Amazon purchases books through commercial channels to help develop and improve the products and services our customers use,” an Amazon spokesperson told me in a statement.

The world’s AI companies are constantly looking for, and spending extreme resources to locate, more material to train their AI models. With books, that sometimes means destroying them in the process, something that large parts of the public have spoken up against, and which we can now confirm Amazon is doing.

In July, I published a story about booksellers who reported a historical spike in sales starting in the past year. They suspected this spike in sales was due to AI companies acquiring any books they can in search of new training data. Printed books are valuable as training data because a lot of the text they contain is not readily available on the internet, which AI companies have already scraped. The data is also conveniently organized and, if the book was printed before 2022, is guaranteed to be free of AI-generated text, which can make any AI model that is trained on it worse via a recursive process called “model collapse.”

📖 Do you know work at a facility where you scan books? I would love to hear from you. Using a non-work device, you can message me securely on Signal at @emanuel.404. Otherwise, send me an email at emanuel@404media.co.

These booksellers suspected AI companies were behind these large bulk purchases because of the high number of books they were buying, the seemingly random choice of books, and the fact that these buyers, unlike libraries and universities, did not seem price sensitive at all. But booksellers couldn’t say for certain who was behind the large purchases because the marketplaces where they sell their books keep the buyers anonymous. When an order comes in, a bookseller ships the sold books to a warehouse operated by the marketplaces, where books are sorted and then sent to the buyer.

In July, one bookseller told me they received a very large order of around 1,000 books on Biblio, one of these marketplaces. The seller agreed to put an Apple AirTag provided by 404 Media in one of the books included in this order so we could see where the book was going. And by extension, which company, AI or otherwise, was behind this massive order. 404 Media granted the bookseller anonymity because they worried sharing this information would harm their business. Biblio did not respond to a request for comment.

Another article to provide some clarity on the types of books: https://finance.biggo.com/news/37a2899b-5f1a-4571-9bb5-1156c0ac4605

top 50 comments
sorted by: hot top controversial new old
[–] justsomeguy@lemmy.world 68 points 4 days ago (3 children)

Alright AI companies I can only hate you so much. Stop adding reasons to the list.

[–] REDACTED@infosec.pub 29 points 4 days ago (1 children)

The fact that their logo is a dinosaur destroying a book..

[–] Zarobi@aussie.zone 25 points 4 days ago (3 children)

What's the bet the logo is A.I. generated as well?

[–] luciferofastora@feddit.org 1 points 1 day ago

Look at the left cover. No, not the middle one; this book has three.

[–] Grail@multiverse.soulism.net 21 points 4 days ago

100%. The dinosaur isn't actually eating it, because AI always gets the prompt wrong

[–] Ganbat@lemmy.dbzer0.com 2 points 3 days ago

The book appears to have two spines...

[–] Test_Tickles@lemmy.world 18 points 4 days ago (2 children)

Look up Google and Project Ocean. Google already did this 2 decades ago. They even went to great lengths to build machines that would very slowly and gently turn pages and non-destructively scan books.
They already have a digital library of 25 million books just sitting there, ready to be instantly and non-destructively copied infinitely.

[–] ThomasWilliams@lemmy.world 1 points 2 days ago

But a court ruling said they had to destroy the books otherwise they would be liable for copyright infringement

[–] Mirshe@lemmy.world 4 points 4 days ago (1 children)

No no, that's too slow and slow is expensive. Time is money and we have to move faster and break more things in order to disrupt the market. /s

I mean yea that's literally why they are destroying books.

[–] Solrac@lemmy.world 3 points 3 days ago

You don't hate them enough. The only response is to do onto them, as they do to us, as they do to these books

[–] BillCheddar@lemmy.world 3 points 2 days ago

They also want the rare books scanned so they can break any/all book ciphers, including those based on rare finds.

[–] hard_zero1@discuss.tchncs.de 30 points 4 days ago (4 children)

If this was done by trustworthy organizations as an effort to digitize and preserve all the books, I would appreciate it. Maybe an association of libraries should do that and, to get the cost back, sell the digital versions / ebooks to the AI companies for an additional price). Then, at least, we would not loose the contents of those rare books to the AI conpanies and prevent them from obtaining a monopoly on the data. And each book would only get destroyed once.

But libraries/bookshops are probably not allowed to sell digital versions, and AI companies are not allowed to use borrowed ebooks.

[–] Duamerthrax@lemmy.world 16 points 4 days ago (1 children)

That exists already. Archive.org has a digitization service, but then archive.org would put a public copy up and the AIbros wouldn't be the sole owners of that training data.

https://digitization.archive.org/

There's also methods to digitize books without destroying the book and for the purpose of AI training that should be more then sufficient. AIbros are just so arrogant to think that other people haven't solved the problem already or that their time is too valuable to be slowed by proper methods.

https://www.diybookscanner.org/

load more comments (1 replies)
[–] distal@lemmy.ml 1 points 2 days ago

No, even this is a very bad idea. Ideally, no books are destroyed. A close relative of mine works as an archivist. They have tons of examples of how digitising and then destroying can create extra burdens. The data on hard drives or flash devices is too fragile.

[–] frongt@lemmy.zip 9 points 4 days ago

Yeah. Libraries got in trouble during COVID for relaxing their ebook borrowing rules. AI companies don't give a shit and are happy to settle lawsuits for a fraction of the money they take in from investors.

[–] P1nkman@lemmy.world 6 points 4 days ago

... not allowed

You think that would stop them?

[–] NewNewAugustEast@lemmy.zip 7 points 3 days ago (4 children)

Can they please give me an example of a rare book?

[–] zarkanian@sh.itjust.works 1 points 2 days ago (1 children)

Are you skeptical that the books being destroyed are actually rare? Or are you skeptical of the existence of rare books?

[–] NewNewAugustEast@lemmy.zip 1 points 2 days ago (1 children)

Skeptical that the books being destroyed are rare.

[–] distal@lemmy.ml 2 points 2 days ago

Pretty sure information that is hard to come by on the internet includes rare books.

load more comments (3 replies)
[–] probable_possum@leminal.space 15 points 4 days ago* (last edited 4 days ago) (14 children)

Are we talking about 14th century handcrafted masterpieces or 2000s university textbooks? Do they destroy cultural heritage items or dusty low value sold-by-weight books?

I need to know if I have to feel angry and agitated or indifferent.

[–] Test_Tickles@lemmy.world 16 points 4 days ago

Textbooks are not "rare", nor are they something that universities and libraries would be price sensitive about. The books are being bought from booksellers who make a living buying and selling rare books. So, while I doubt that they are all going to "masterpieces", they are going to be books valuable enough to support an industry of people and expensive enough that universities and libraries would be price sensitive about them.

[–] zarkanian@sh.itjust.works 1 points 2 days ago

The answer is yes in all cases. Well, almost all. They don't want books printed after 2022, because those might have AI-generated text in them. Anything else is fair game.

[–] frongt@lemmy.zip 6 points 4 days ago

The latter. These are books that have been sitting on the shelf in a warehouse for years. They're not sold by weight, but they are stuff like random technical manuals for stuff very few people care about.

My only hope is that an eventual lawsuit forces the AI companies to release the digitized versions to an archive or library.

[–] skisnow@lemmy.ca 5 points 4 days ago (1 children)

We're almost certainly talking about literally everything a bot scraping catalogues decides it doesn't have. I highly doubt it'll be discriminating.

load more comments (1 replies)
load more comments (10 replies)
[–] queermunist@lemmy.ml 11 points 4 days ago (3 children)

I don't understand the concept of a "rare" book. Every book should be infinitely reproducible, the fact that they aren't is a crime.

[–] Ebby@lemmy.ssba.com 18 points 4 days ago* (last edited 3 days ago) (3 children)

There are many ways a book can be rare. First editions, signed copies, and books with little demand. Can't fire up the printers and make those again.

In the case of little demand, I have a book "Two thousand leagues under the seas" not the more common "Twenty thousand leagues under the sea". It's rare in that it's very difficult to find more literal translation than the version we are accustomed with. I suspect search engines simply think I've made a typo. Either way, why print it if there is little demand or can't find the book?

[–] Leomas@lemmy.world 3 points 3 days ago (1 children)

I agree with your point overall, but why is the book two thousand leagues under the sea more literal, when the original is called "Vingt Mille Lieues sous les mers", is a league 10 times as much, as a Lieue? Genuinely curious (and hard to google)

load more comments (1 replies)
[–] zaphod@sopuli.xyz 1 points 3 days ago

The title doesn't make sense, or is it an extremely shortened version and it's a pun on the fact that it shortened the journey? Do you happen to know who the translator was?

load more comments (1 replies)
[–] Todd_cross@lemmy.dbzer0.com 7 points 3 days ago (3 children)

A book's value is not just held in its text. They're also valuable as historical objects.

load more comments (3 replies)
[–] ThomasWilliams@lemmy.world 1 points 2 days ago (1 children)

They only print a few dozen copies and then Google destroys the only one left.

It's not a difficult concept.

[–] queermunist@lemmy.ml 1 points 2 days ago

It's not like we're printing books from old school printers where they had to assemble the letters by hand on a printing plate for each page, or they need scribes to hand transcribe books with a quill. It's all computer. There's no reason for any book to be rare.

[–] BilSabab@lemmy.world 4 points 3 days ago (1 children)

local AI companies are the reason Ukrainian book publishing industry experienced an uptick in sales lately. Bros stock up entire catalog at once and no one complains because its a hefty paycheck.

load more comments (1 replies)
[–] altkey@lemmy.dbzer0.com 11 points 4 days ago (3 children)

The symbol for Amazon's VGT3, the Las Vegas facility where it scans book for AI training data.

I thought it's 404media's obviously satirical preview picture to dab on aibros, but the truth is even more weird

load more comments (3 replies)
[–] mr_tyler_durden@lemmy.world 0 points 2 days ago (1 children)

ISBNs or STFU.

“Rare” means absolutely nothing without context. This is trash reporting, you sent the books, you know what they are, by not publishing the titles/ISBN it proves you know you don’t have a valid case. I’m betting these were mostly/all used books.

The pearl clutching about this is absurd. So what if they buy 1 of every book to scan?

[–] themachinestops@lemmy.dbzer0.com 2 points 2 days ago (1 children)

I also wanted to know what the rare books are and this is what I Found:

https://finance.biggo.com/news/37a2899b-5f1a-4571-9bb5-1156c0ac4605

I also heard that they are also scanning old outdated medical books.

So are not really aiming for rare, they are just gathering at random and getting their hands on any old book they can find. Some rare books might or might not be in the books they scan.

[–] luciferofastora@feddit.org 1 points 1 day ago

I also heard that they are also scanning old outdated medical books.

What the fuck

Why would you want to train on outdated information? What the hell are these people smoking?

load more comments
view more: next ›