Secondhand booksellers across the UK and Ireland are experiencing an unusual surge in bulk orders, prompting speculation that artificial intelligence companies are acquiring vast quantities of books for data acquisition purposes. The pattern of orders, often comprising disparate and seemingly random titles, has led many booksellers to suspect a connection to the growing demand for training data by AI firms.
This phenomenon is not isolated to the British Isles. Booksellers in the United States, Australia, and continental Europe have also reported similar increases in peculiar bulk purchases. The trend appears to be linked to the practice of AI companies acquiring physical books, scanning their contents to train large language models (LLMs), and subsequently discarding the original copies.
Anthropic, a prominent AI company and developer of the Claude chatbot, has been identified as one such entity. Reports indicate that Anthropic spent tens of millions of dollars on acquiring millions of books. The company's project, internally codenamed "Project Panama," involved destructively scanning these books. This process included removing bindings, slicing pages, and scanning them for data. Anthropic stated that sourcing books is a common method for training LLMs across the AI industry. The company also noted that it does not acquire rare or antiquarian books for this purpose.
The unusual nature of the orders is highlighted by their lack of thematic coherence, a departure from typical book collector behaviour. For instance, one bookseller noted an order that included an Estonian translation of a John le Carré novel, a specific imprint of Anne Brontë's Agnes Grey, and a military history magazine. Another bookseller described receiving orders for books on subjects as varied as 18th-century African agricultural implements and biographies of 1950s racing car drivers. These orders often originate from buyers in the US, Canada, and continental Europe, as well as within the UK.
The practice of AI companies acquiring and destroying books for training data has raised ethical and cultural questions. Some observers argue that this method, while potentially offering a source of "clean" data free from AI-generated text, leads to the irreversible loss of physical literary works. The focus on books published before 2022 is partly driven by a desire to avoid text that may have been generated by AI itself, a phenomenon that could lead to "model collapse" where AI models trained on AI-generated text become less capable over time. Additionally, concerns exist about data poisoning, where malicious actors might introduce subtle alterations into digital text that could corrupt AI models. Books predating these techniques offer a more stable data source.
Legal precedents have emerged regarding the use of books for AI training. A judge ruled that Anthropic's method of purchasing physical books, scanning them, and then destroying the originals constituted "fair use" under US copyright law. This ruling has potentially paved the way for other AI laboratories to adopt similar practices. The practice is also seen as a way to circumvent copyright issues by leveraging the "first-sale doctrine," which allows a buyer to do what they wish with an item after its initial purchase.
Intermediaries, such as ISBNdb, have reportedly begun offering services to AI companies for anonymous bulk book purchases, ranging from thousands to a million copies per order. These services often promote strict non-disclosure agreements and may frame the process as "digital preservation" to mitigate potential public backlash. While there is no definitive proof that every unusual order is from an AI company or that every purchased book is destroyed, the pattern of behaviour and the emergence of these specialised services suggest a coordinated effort within the AI industry to source and process physical books for training data.
