Amazon built its name by selling books. Now, reporting from Las Vegas suggests some rare volumes may be ending their lives as raw material for the company’s AI ambitions.
The alarm began last month, when 404 Media reported that AI companies were purchasing used books in bulk for training data. The latest reporting pushed the concern further: the outlet placed a tracker inside a shipment of rare books and traced it to Amazon’s VGT3 facility in Las Vegas.1
Workers at that location said the operation is brutally straightforward. Amazon receives large shipments of printed books, removes their bindings and scans them faster once the pages are loose — a process in which “the printed book is destroyed.”1 The warehouse’s own symbol, according to the reporting, is a dinosaur clutching a book: an unintentionally stark image for a dispute over whether hard-to-replace texts are being treated as disposable data.2
Amazon’s position is that it is sourcing material through ordinary means. The company told 404 Media it “purchases books through commercial channels to improve the products and services customers use.”2 But critics see a deeper collision between the voracious text demands of large language models and the finite supply of out-of-print books.
The stakes extend beyond nostalgia. Older, obscure texts can be especially useful to AI developers precisely because they may not already be available online — and because books published before 2022 predate the flood of machine-written material. As AI systems increasingly train on synthetic output, researchers and industry observers warn of “model collapse,” in which repeated exposure to generated text can degrade a model’s results.2
The reporting does not establish how many books Amazon has processed or whether every tracked volume was scanned. It does, however, crystallize the central tension: the same physical books long valued as cultural objects are becoming prized inputs in the race to build smarter machines.