An employee described how the Amazon VGT3 warehouse destroys university and government books for AI training
A new interview by 404 Media with an employee at the Amazon VGT3 warehouse in Las Vegas reveals that the books being destroyed include volumes from university libraries (University of London) and government documents; employees are unsure of the purpose, and the destroyed books cannot be reassembled.
A new interview by 404 Media with an anonymous employee at the VGT3 warehouse in the Amazon warehouse complex in northeastern Las Vegas provides details about the scale of the operation: the books being processed included shipments from university libraries, including University of London, as well as government documents marked as presented to Parliament on behalf of the British Queen. According to the employee, many workers are unsure of the actual purpose of the operation, the process changes daily and appears disorganized. After the pages are scanned, the books are cut out of their bindings and the loose sheets are dumped, mixed together, into large containers called shuttles, making it impossible to reassemble them.
The VGT3 warehouse is located in the same building as LAS8, an Amazon facility for printing books to order (print-on-demand), and both are part of the broader Amazon warehouse complex in Las Vegas. LAS8/VGT3 was the destination of a shipment of approximately 1 000 books ordered through the Biblio marketplace, which 404 Media tracked using an AirTag placed in one of the books, from a California airport through Milwaukee and a distribution center operated by Trifinity to a truck heading to Nevada. Amazon told 404 Media that it “purchases books through commercial channels to improve the products and services customers use".
According to earlier reporting, rare and unavailable books are valuable to AI companies primarily because their text is not available on the internet and books published before 2022 are guaranteed not to contain AI-generated text, reducing the risk of so-called model collapse (a degradation in model quality after training on AI-generated text). According to TechCrunch, Anthropic also previously resorted to pirated books to train its models.
Why it matters
Details from the interview show that the scale of book purchasing and destruction for training AI models also encompasses materials from public institutions (university libraries, government documents), beyond commercial titles, and that the process takes place without transparent communication of its purpose even to the warehouse workers themselves. For publishers, libraries and booksellers, this means a risk that their collections may irreversibly end up as training data without their knowledge; for companies developing AI models, it raises further questions about the transparency and legitimacy of data sources.
What was added since the original report
Verified updates
-
The books come from university collections, including University of London; The materials being scanned include government documents; Employees are not entirely sure of the purpose of the AI training operation; The destroyed books are mixed together and remain impossible to reassemble; The LAS8 warehouse (print-on-demand) is part of the same Las Vegas complex
- The books come from university collections, including University of London
- The materials being scanned include government documents
- Employees are not entirely sure of the purpose of the AI training operation
- The destroyed books are mixed together and remain impossible to reassemble
- The LAS8 warehouse (print-on-demand) is part of the same Las Vegas complex
-
Rare and out-of-print books are specifically targeted for their value in training; Texts published before 2022 are not AI-generated (a key characteristic for training without the risk of degradation); Anthropic also uses illegally pirated books to train its models; Model collapse risk arises when training LLMs on AI-generated content; The internet as a primary source of training data has been substantially exhausted
- Rare and out-of-print books are specifically targeted for their value in training
- Texts published before 2022 are not AI-generated (a key characteristic for training without the risk of degradation)
- Anthropic also uses illegally pirated books to train its models
- Model collapse risk arises when training LLMs on AI-generated content
- The internet as a primary source of training data has been substantially exhausted
Two audiences, two different impacts
What this means
For individuals
Sellers and owners of rare or library books should be aware that their titles may end up cut apart and destroyed in a warehouse used to scan data for AI training, without them knowing who the actual buyer is.
For a business
Companies purchasing or processing training data for AI models face a growing risk of reputational and legal issues surrounding the destructive processing of physical books, including materials from university libraries and government documents, without a clearly communicated purpose to their own employees.
Risks and complianceCheck the original
Event sources
confirmed by 2 independent sources · 2 publishers, 2 independent. We count feeds from the same owner only once.