Used-book retailers around the globe are receiving massive, anonymous bulk orders of low-value non-fiction titles that bookstore owners suspect are being harvested to train artificial intelligence models. According to RTÉ, Galway bookseller Tomás Kenny received an order in May for several thousand books featuring eclectic, outdated subjects, including three-decade-old driving test manuals and financial guides from the Celtic Tiger era.
The Mechanics of Scattergun Book Orders
The orders typically arrive through third-party intermediaries that conceal the ultimate identity of the buyer, according to statements given by Tomás Kenny to RTÉ’s News at One. While bulk purchases are common for library supplies or institutional clients, the composition of these recent orders lacks any thematic coherence. Kenny described the selections as “scattergun,” consisting of poor-quality, out-of-date non-fiction titles that few human readers would seek out simultaneously.
Independent booksellers in Sweden and Germany have reported observing identical purchasing patterns on social media platforms. With the expansion of digital retail reducing the total number of independent book vendors worldwide, these scattered transactions have emerged across distinct international markets, pointing toward an automated or centralized acquisition strategy.
Copyright Scrutiny and AI Training Data
The suspicion that these physical books are headed for artificial intelligence ingestion coincides with mounting legal scrutiny over how tech companies source training data. Last month, a US federal judge granted final approval for an historic €1.3 billion copyright class-action settlement involving Anthropic, according to reporting highlighted by RTÉ. That litigation accused the company of using more than 500,000 pirated books to train its Claude AI model, with eligible authors and publishers slated to receive roughly €2,600 per qualifying title.

While booksellers like Kenny emphasize they lack concrete forensic evidence tying the bulk purchases directly to technology firms, the industry-wide phenomenon highlights a growing demand for diverse, long-form text datasets as developers seek legal workarounds or supplementary material to train large language models.
Worth a look