The trail from a July bulk order ends in a Las Vegas warehouse, where copies that may be impossible to replace are stripped, scanned and thrown away — and Amazon will not say what for.

A bookseller slipped an Apple AirTag into one volume from a July order of about 1,000 books placed through the Biblio marketplace. The tracker stopped in the VGT3 section of Amazon's LAS8 facility in the northeast of Las Vegas — the first time, as 404 Media reported earlier this week, that anyone has publicly traced where large-scale book buying tied to AI training ends up.

Employees there say that taking in massive shipments of printed books and cutting the bindings off them is all the warehouse does, 404 Media reported: stripped of its spine, a book feeds through a scanner far faster, and does not survive the process. The logo of the team that works there is a dinosaur baring its teeth, a book in its hands. What goes through the scanner comes off shelves the public could otherwise reach.

Who was doing the buying had been guesswork for about a year. Booksellers watched unusually large orders arrive from purchasers who showed no sensitivity to price, often stayed anonymous, and appeared to pick titles at random. AI companies were the obvious suspects, but proving it was another matter.

Asked by Ars Technica about the findings, Amazon declined to address them and repeated only the general statement it had already given: that it buys books from commercial suppliers as part of building and improving the products and services its customers use. The statement makes no mention of AI training.

The broader appetite for paper is easy enough to account for. Large language models have already worked through most of the text that is readily available online, which pushes companies toward less accessible material. Much of what printed pages carry is not easy to find on the internet at all, leaving those volumes a largely untouched reservoir of training text. Print produced before 2022 also comes with a very low chance that another AI wrote it, reducing the risk of model collapse — the fall-off in quality that sets in when a model is trained over and over on output produced by AI.

Amazon is building what it considers frontier models, which need large volumes of unique text to stay level with Google, OpenAI and Anthropic. Scarce, hard-to-find titles could hand it an edge there, since rivals Anthropic and xAI have publicly stated that they do not train on rare or antique books.

Amazon did not invent the approach, and Anthropic's stated abstention is narrower than it sounds — it covers rare and antique books, not ordinary ones. An authors' lawsuit exposed an in-house Anthropic programme, Project Panama, that bought books on marketplaces, cut off their spines and digitized them by the million.

Destructive scanning has already been through a courtroom: a federal judge found that using lawfully bought books to train AI can count as fair use, an outcome that went partly in Anthropic's favour and an argument Meta and OpenAI have also advanced. The reasoning leaned in part on the destruction itself — because the printed originals were destroyed during scanning, they were not copied and resold.