Generated by Codex with GPT 5.6 Sol XHigh

Techmeme surfaced Emanuel Maiberg’s August 17 investigation, “We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility”. The story turns a murky argument about AI training data into a physical supply-chain investigation: 404 Media placed a tracker in a rare-book shipment and followed it to Amazon’s VGT3 facility in Las Vegas, where Amazon scans books for AI training and destroys the paper copies.

The finding matters because it identifies a major buyer behind a pattern that booksellers had already noticed. Dealers in several countries had reported unusually large orders for old, obscure, and specialized books, often routed through intermediaries that would not name the final customer. Earlier reporting could show that AI labs wanted printed books and that destructive scanning existed, but it could not reliably connect a particular order to a particular company. Following one shipment does not reveal the scale of Amazon’s entire operation, yet it replaces speculation about at least one part of the market with a documented route and destination.

Why AI companies want the physical books

Books published before the recent wave of generative AI are attractive training material because they are long-form, edited, and overwhelmingly human-written. They also contain technical, historical, and cultural knowledge that may never have been digitized or may sit behind fragmented licensing arrangements. As synthetic text spreads across the web, an old print collection becomes a relatively clean reservoir of language and information.

The fastest way to turn that reservoir into machine-readable text is often destructive scanning. A vendor removes the binding, separates or trims the pages, runs them through high-speed scanners, performs optical character recognition, and discards the remains. The process is cheaper and faster than photographing a bound volume page by page. It also treats the book as a temporary container for text rather than as an artifact whose paper, binding, annotations, ownership marks, and printing history may carry value of their own.

That distinction is easy to ignore when a title is common and replaceable. It becomes serious when “rare” means scarce, out of print, locally significant, or one of only a small number of surviving copies. A digital scan may preserve the words while removing the physical copy from circulation and placing the replacement inside a private corporate dataset. The knowledge survives in one sense, but public access and the object’s historical evidence may not.

The law creates an unusual incentive

The practice also sits inside a legal structure that can reward one-for-one destruction. In the June 2025 fair-use order in Bartz v. Anthropic, a federal district judge treated Anthropic’s scanning of lawfully purchased print books—and disposal of the originals—as a permissible internal format change. The same order sharply distinguished those purchases from millions of books obtained through pirate libraries.

That ruling concerned Anthropic, not Amazon, and a district-court decision does not settle every future AI copyright case. Still, it helps explain why buying, cutting, scanning, and discarding can look safer than copying an unauthorized ebook or keeping both a print and digital copy. The physical destruction is not merely an unfortunate side effect of efficiency; under this reasoning, it helps support the claim that the company replaced one owned copy with another instead of multiplying copies for distribution.

The result is a mismatch between copyright logic and preservation. Copyright asks whether the copy was lawfully acquired, how it was transformed, and whether the new use harms a protected market. A library or archivist asks different questions: How many copies remain? Does this edition contain unique evidence? Will researchers or the public still be able to consult it? A transaction can look acceptable under the first framework while causing an irreversible loss under the second.

What the investigation proves—and what it does not

The tracker establishes that a shipment of rare books reached an Amazon facility used for AI scanning. It does not establish how many scarce titles Amazon has destroyed, whether every book entering the facility is cut apart, or whether “rare” always means irreplaceable. Those distinctions matter. Public concern should rest on inventories and preservation practices, not on the assumption that every old book is unique.

But the unanswered questions are now operational rather than hypothetical. Amazon can identify scarce editions before destructive scanning, use nondestructive methods for them, or transfer physical copies to libraries after making a scan where rights and logistics permit. Suppliers can disclose that books may be destroyed and screen orders against rarity data. The cost would be higher, but that is precisely the point: an AI dataset should not become cheaper by silently transferring the risk of cultural loss to everyone else.

The investigation’s durable lesson is that AI governance begins well before model training. It includes procurement contracts, warehouse workflows, scanning vendors, rarity checks, and decisions about who retains access to the resulting archive. Once the binding has been cut and the pages discarded, a policy written later cannot restore the object.