Amazon is acquiring significant numbers of printed books, including rare volumes, to digitize them for artificial intelligence training. The process, uncovered by a 404 Media investigation, involves scanning the books at a dedicated Amazon facility in Las Vegas, Nevada, after which the physical copies are destroyed. Employees at this facility, identified as VGT3, reportedly remove the bindings of books to expedite the scanning process. The team's logo, a dinosaur holding a book, has been noted.

The investigation began when 404 Media, suspecting that AI companies were driving a recent surge in book sales, placed a tracking device in a rare book shipment. This shipment was traced to the Amazon warehouse in Las Vegas. The publication noted that the books involved were rare, meaning few copies were in circulation, either due to limited original print runs or their presence in less common languages.

Amazon acknowledged its book acquisition practices in a statement to 404 Media, stating, "Amazon purchases books through commercial channels to help develop and improve the products and services our customers use". This practice aligns with the broader trend of AI companies seeking vast amounts of text data to train their large language models, especially as easily accessible online data becomes more saturated.

The method of destroying physical books after scanning has drawn parallels to similar operations by other AI companies. Anthropic, for instance, was revealed to have a "Project Panama" that involved acquiring books, removing their spines, and scanning them for training data. A lawsuit filed by authors against Anthropic highlighted this practice, with a judge ruling that the scanning qualified as fair use, partly because the original copies were destroyed. This destruction prevents the creation of duplicate copies that could compete with publishers.

Booksellers have observed a historical spike in book sales over the past year, suspecting AI companies as the primary buyers. Some speculate that AI firms are attempting to systematically scan every printed book by utilizing ISBN numbers, which are unique identifiers for published works. The rarity of some of the books acquired means they may hold unique information not readily available elsewhere, making them potentially valuable for AI training.

The practice of acquiring and destroying books for AI training is controversial. Critics argue that it removes potentially irreplaceable knowledge from public access and concentrates it within the proprietary data sets of corporations. The destruction of original copies also raises ethical questions for booksellers who value the recirculation and preservation of literature. While Amazon has not detailed specific AI models trained with this data, it is understood that such data is crucial for developing and refining AI capabilities. The company's use of scanned book data for its "Nova models" has been mentioned.

The trend suggests a resource-intensive and legally complex approach by AI companies to secure high-quality training data. As easily scrapable web data becomes less novel, the acquisition of physical media, even at the cost of its destruction, represents a significant effort in the ongoing race for advanced AI development.