A tracking device hidden inside a rare book confirmed on August 17, 2026 what many in the publishing world had long suspected: Amazon is purchasing out-of-print volumes by the truckload, slicing off their spines, and feeding the scanned pages into its AI training pipeline. The facility receiving those books — a Las Vegas warehouse identified by its unusual logo of a dinosaur clutching a book — is designated VGT3, and it has now become the symbol of an industry-wide reckoning over where AI training data comes from and what gets destroyed in the process.
Background You Need
The AI training data crisis has been building for years, though its urgency has sharpened dramatically as the largest language models approach the ceiling of what the public internet can offer. Early LLMs were trained on enormous crawls of web text — Wikipedia, Reddit, Common Crawl, digitized news archives. That supply is not infinite, and the models have already consumed most of the high-quality portion. What remains on the internet is increasingly AI-generated content, which creates a compounding problem researchers call model collapse: when a model trains on outputs from other models, small errors and distortions amplify across generations, degrading quality in ways that are difficult to reverse. Human-written text, especially text produced before the LLM era, has become the scarcer and more valuable ingredient.
That scarcity has pushed AI companies toward books — specifically physical books that were never digitized at scale. Rare, out-of-print volumes are particularly attractive because they represent guaranteed pre-2022 human writing, untouched by the generative AI content flood. They also can't be licensed through standard channels because most don't have active rights holders willing to negotiate deals. This combination of qualities — scarce, human-authored, legally complicated — has made them a target.
Amazon is not the first company to recognize this. Literary Hub reports that Anthropic was already pursuing what it internally called "Project Panama" — a secret effort to destructively scan books for training data — details of which emerged during a copyright lawsuit brought by authors against the company earlier this year. That disclosure prompted The Washington Post to report in January 2026 that Anthropic was "at work on a secret project to destructively scan all the books in the world." Multiple reports now indicate that Anthropic, Meta, Google, Microsoft, and OpenAI are all facing active lawsuits alleging similar unauthorized use of copyrighted text.
What's New
The Amazon story is distinguished by how it was confirmed. 404 Media, the outlet that broke the story, placed a tracker inside a rare book and sold it through commercial channels. The device ended up at VGT3, the Las Vegas facility — the same one whose dinosaur-clutching-a-book branding now reads as something between irony and corporate declaration of intent. The methodology matters: this is not a whistleblower account or an inference from circumstantial evidence. It is a confirmed chain of custody, from bookseller to Amazon's AI data operation.
Amazon's response, provided to 404 Media, was carefully worded: the company said it "purchases books through commercial channels to improve the products and services customers use." It did not deny destructive scanning. That phrasing matters because, as Literary Hub notes, Amazon had previously denied practicing destructive scanning specifically — making this non-denial a meaningful shift in posture. When a company's prior denial was specific and its current statement is general, the gap between them is the story.
The scale of the operation is not fully known, but multiple reports indicate the purchases are not occasional or experimental. The VGT3 facility appears to be a dedicated operation, not a side function of a fulfillment center. The dinosaur logo — presumably an internal branding decision someone thought clever — suggests the team working there has enough self-awareness to recognize what they're doing: a predator consuming something that can't fight back.
What makes rare books worth the logistics cost is precisely the pre-2022 cutoff. Any book published and digitized after large language models became prevalent risks containing text that was itself AI-assisted or AI-generated. Rare physical books from earlier decades carry no such risk. For a training dataset, that provenance guarantee is worth the cost of acquisition, spine removal, scanning, OCR processing, and physical disposal. The books don't survive the process.
The Pushback
Our read is that the model collapse concern driving this behavior is real and not well understood outside research circles. The feedback loop — AI generates text, that text gets scraped into training sets, new models train on it — has already started. Some researchers have published evidence of degraded output quality in models trained on successively more synthetic data. The industry's response has been to race toward the remaining stocks of verified human writing, and rare books are a logical endpoint of that race. That this is happening is not surprising; that it requires a hidden AirTag to confirm it is.
The legal exposure is significant and underappreciated. Copyright in the United States does not have an exception for AI training, and while that question is still being litigated across multiple active cases, the industry's bet appears to be that training data from physical books — especially older ones where rights are fragmented or unclear — is practically harder to trace than scraped web content. Rare books leave paper trails in the antiquarian market, but those trails are not the same as the digital logs that make web scraping easier to litigate. Whether that calculus holds up in court remains open. What the 404 Media investigation demonstrated is that the physical trail is traceable too — you just need to plant a tracker.
Ars Technica's coverage of the AirTag investigation adds texture to the supply chain: the books travel through commercial channels before reaching VGT3, meaning Amazon is purchasing through the same rare book markets that librarians, collectors, and researchers use. That's not a neutral fact for the rare book trade, which operates on the assumption that books, however worn, survive their transactions.
Sources
techcrunch.com Amazon, which started off selling books, is destroying rare texts to train AI Literary Hub » Now Amazon is destroying rare books to train its AI. Hidden Airtag reveals Amazon is trashing rare books to train AI - Ars TechnicaBased on
https://techcrunch.com/2026/08/17/amazon-once-an-online-bookseller-is-destroying-rare-books-to-train-ai-models/— techcrunch.comThis article is an original, AI-assisted summary and analysis. Credit for the underlying reporting or footage belongs to the source above.

Written by the vybecoding.ai editorial team
Published on August 17, 2026