Amazon Reportedly Buys and Destroys Rare Books for AI Training
An investigation has linked large-scale purchases of rare and out-of-print books to an Amazon facility where physical copies are reportedly cut apart and scanned to obtain training data for AI systems.

Amazon is facing criticism over reports that it has been purchasing large quantities of rare, out-of-print and secondhand books for artificial-intelligence training. An investigation tracked a shipment of books to an Amazon facility in Las Vegas, where employees reportedly process large volumes of printed material for scanning. The reports have raised concerns among booksellers, collectors and publishers about the destruction of valuable physical works.
According to the investigation, books arriving at the facility are reportedly cut along their bindings so that their pages can be scanned more efficiently. The physical copies are consequently destroyed during the process. Amazon has said that it purchases books through commercial channels to help develop and improve its products and services, but the company has not publicly provided a complete explanation of the specific operation described in the investigation.
The controversy comes amid a wider rush by AI companies to obtain large quantities of high-quality human-written material. Rare and older books can be particularly valuable because they contain text produced before the widespread emergence of generative AI, potentially giving developers cleaner datasets. Similar reports have linked other AI-related companies and intermediaries to large orders of older books, including thousands of individual titles.
For booksellers and collectors, the practice raises concerns that historically significant or difficult-to-replace books could be treated primarily as raw data rather than cultural objects. Australian booksellers, for example, have raised alarms after discovering that some older titles may have entered an AI-related supply chain in which books were scanned and subsequently destroyed. The issue has therefore become both a technology story and a cultural-preservation debate.
The situation also highlights the growing tension between the enormous demand for AI training data and the value of human-created intellectual works. As AI companies build increasingly capable models, they are looking for vast quantities of reliable text, while publishers, authors and collectors continue to debate how copyrighted and historical material should be used. If the reported Amazon operation expands, it could intensify calls for greater transparency about where AI training data comes from and what happens to physical works after they are digitized.



