An AirTag Followed 1,000 Books to Amazon. What Reporters Found Inside Was an AI-Era Nightmare for Bibliophiles
An AirTag Followed 1,000 Books to Amazon. What Reporters Found Inside Was an AI-Era Nightmare for Bibliophiles

An AirTag Followed 1,000 Books to Amazon. What Reporters Found Inside Was an AI-Era Nightmare for Bibliophiles

Share story

Advertisement

It began with something almost absurdly simple.

An Apple AirTag.

A used book.

And a bookseller who wanted to know who was suddenly buying enormous quantities of obscure physical books.

For months, booksellers had noticed unusual orders appearing through used- and rare-book marketplaces.

Hundreds of titles.

Sometimes thousands.

Often with little obvious connection between subjects.

The buyers did not appear especially concerned about condition.

They were not necessarily searching for beautiful first editions or collectible bindings.

What seemed to matter was the text.

404 Media reporter Emanuel Maiberg had been investigating suspicions that artificial-intelligence companies were behind at least some of these bulk purchases, hunting for printed material that had never been digitized and could provide enormous amounts of high-quality human-written text for AI training.

Then, in July 2026, a bookseller received an unusually large order of roughly 1,000 books through Biblio, an online marketplace specializing in used, rare and out-of-print books.

The seller agreed to an experiment.

404 Media supplied an AirTag.

The tracker was hidden inside one of the books.

The shipment left.

And reporters watched.

The tracked book moved from California through Wisconsin, spent roughly two weeks near Kenosha, traveled west through Colorado and eventually stopped at an Amazon facility in Las Vegas.

More specifically, it reached the northern portion of Amazon's LAS8 complex, where employees identified a separate operation known internally as VGT3.

What workers described happening there sounded almost designed to horrify book collectors.

Books arrive in bulk.

Barcodes and ISBNs are checked.

Bindings are cut away.

Loose pages are fed rapidly through industrial scanners.

After scanning, the pages are thrown together into large containers.

The physical book cannot be reconstructed.

The text survives digitally.

The object does not.

404 Media concluded that Amazon was destructively scanning books as part of an AI-training-data operation.

Amazon did not specifically confirm the AI-training purpose in its public statement, but it acknowledged buying books commercially to develop and improve its products and services.

The story is real.

But some viral versions go beyond what the investigation actually established.

Amazon was not shown destroying books to “dispose of evidence.”

The shipment contained books sold through legitimate commercial channels, not secretly stolen library holdings.

Not every volume moving through VGT3 has been proven to be rare or irreplaceable.

And although evidence gathered by 404 Media strongly connects the facility to AI data work, Amazon itself did not publicly identify the specific model or product receiving the scans.

The more accurate story is already extraordinary enough:

A journalist used an AirTag to trace an anonymous bulk purchase of roughly 1,000 used and rare books to an Amazon facility where employees say books are physically cut apart, rapidly digitized and discarded—and where 404 Media says the resulting data is being used for artificial intelligence.

The Investigation Started With Strange Book Orders

The AirTag investigation did not emerge from nowhere.

Before tracing Amazon, Maiberg had already documented a strange change in the secondhand-book market.

Booksellers reported unusually large orders for older books.

The purchases sometimes involved obscure nonfiction, specialized subjects and titles difficult to find online.

One company pitching book-acquisition services to AI developers had described physical books as an unusually valuable source of training material because they contain edited, structured, human-authored information that cannot necessarily be recovered through ordinary web crawling.

That logic is becoming increasingly important in AI development.

The internet is enormous.

But the internet is not the same thing as all human knowledge.

Millions of books have never been fully digitized.

Others exist digitally only behind paywalls, in private databases or in forms that cannot simply be scraped.

Older physical books therefore represent something extremely valuable:

human-written text that may not already exist inside everybody else's training dataset.

Why Older Books Are So Valuable to AI Companies

Large language models improve partly through exposure to enormous amounts of text.

But quantity is not the only issue.

Quality matters.

Books often provide characteristics that web pages do not consistently offer:

  • Long-form arguments
  • Carefully edited prose
  • Specialized vocabulary
  • Narrative structure
  • Historical information
  • Technical expertise
  • Rare subject matter
  • Material absent from the public web

There is another increasingly valuable characteristic.

Older books overwhelmingly predate the current explosion of generative-AI content.

That makes them attractive as high-provenance human data.

The distinction matters because the modern internet increasingly contains text created or heavily rewritten by AI systems.

Researchers have shown that indiscriminately training successive generations of models on recursively generated synthetic data can produce degradation known as model collapse, in which models progressively lose information about the original human-data distribution.

The Nature study behind that concept warned that preserving access to original human-generated data becomes increasingly valuable as AI-generated content spreads across the web.

A book printed in 1986 does not have a ChatGPT-generated chapter hidden halfway through it.

A technical manual from 1994 was not written by Claude.

A local-history book published in 1972 is unlikely to have been generated by a modern language model.

That provenance has value.

“Pre-2022” Does Not Literally Mean Every Word Is Guaranteed Human

Some descriptions of the story claim that books published before 2022 are completely guaranteed to contain no machine-generated text.

That is slightly too absolute.

Computer-generated language existed long before ChatGPT.

Templates, automated reporting systems and experimental text-generation systems existed for years.

What changed after 2022 was scale.

Generative AI suddenly became capable of producing enormous amounts of fluent prose cheaply and quickly.

The useful distinction therefore is not:

Before 2022 = mathematically guaranteed human.

It is:

Older print material overwhelmingly predates the mass proliferation of modern LLM-generated text and therefore offers much cleaner provenance than today's open web.

That alone makes physical books extraordinarily attractive training material.

The AirTag Took a Cross-Country Journey

According to reporting reconstructing the 404 Media experiment, the tracked book first left California by air.

It reached the Milwaukee area.

Then it remained for approximately two weeks at a distribution site near Kenosha, Wisconsin.

Afterward, it moved west by truck.

The tracker registered a stop around Grand Junction, Colorado.

Eventually, its signal settled at Amazon's LAS8 facility in North Las Vegas.

That initially created confusion.

LAS8 is associated with Amazon's print-on-demand business.

A company printing books seemed like a strange destination for a shipment suspected of being destined for destructive scanning.

Then the reporting narrowed the location.

The AirTag had reached the northern portion of the facility.

Employees described another operation there:

VGT3.

What Is VGT3?

VGT3 appears to be an internal Amazon operational designation rather than the public name of an AI laboratory.

404 Media's follow-up reporting placed the operation inside the same larger facility as LAS8 and interviewed an employee who had worked there.

The detail that made the facility instantly memorable was its logo.

A dinosaur.

Holding a book.

The imagery became darkly symbolic once the operation inside was known.

Workers described a book-processing pipeline involving bulk intake, sorting, cutting and high-speed scanning.

But describing VGT3 as a mysterious “high-security Amazon AI laboratory” would exaggerate what the evidence shows.

It is better understood as a warehouse-scale digitization operation associated by 404 Media with AI training work.

The important part is not whether it looks like a secret laboratory.

The important part is what the machinery does.

First, Amazon Checks What Books It Has

The employee interviewed by 404 Media described boxes and large bulk containers of books arriving at the operation.

Workers placed volumes into totes before sending them toward other processing stations.

The employee said another area scanned barcodes to determine what Amazon needed and identify duplicates.

This detail supports a theory booksellers had already developed.

The purchases did not necessarily look like traditional book collecting.

They appeared potentially systematic.

Booksellers suspected AI companies might be working their way through ISBN databases, trying to acquire textual coverage rather than collectible objects.

404 Media noted that some large purchases reportedly excluded extremely rare books lacking ISBNs, further contributing to that theory.

It resembles dataset construction.

Not library building in the ordinary human sense.

Then the Bindings Come Off

Once selected for scanning, books move toward cutting stations.

The VGT3 employee described machines with guarded cutting areas where workers place books and activate a blade.

The binding is removed.

The pages become loose sheets.

Why destroy the binding?

Speed.

Scanning a bound book non-destructively can be slow.

Pages have to be turned.

Curvature near the spine must be corrected.

Fragile volumes require careful handling.

By slicing away the binding, pages can be fed through automated scanners much like sheets moving through a high-speed document-processing machine.

The efficiency difference can be enormous.

But it changes the nature of the project.

This is not preservation scanning.

It is destructive digitization.

Sam Altman’s ChatGPT Water Claim, Fact-Checked
Sam Altman says one California almond uses as much water as 38,000 ChatGPT queries. The broad point holds—but the exact math needs context.

The Scanners Can Process Pages Extremely Quickly

The employee told 404 Media that VGT3 appeared to contain roughly 20 to 25 scanning machines.

The machines were compared visually with devices that rapidly count banknotes.

Loose pages could be fed through while images appeared on computer screens.

At warehouse scale, this approach makes sense computationally.

Imagine manually turning 300 pages in one book.

Now imagine doing that with one million books.

At one book per minute—which is already unrealistically fast—one million books would require almost two years of nonstop processing on a single station.

High-speed destructive scanning turns the problem into industrial automation.

The book is converted from an object designed for human reading into raw information for machines.

And After Scanning, the Book Is Gone

This is the detail causing the strongest emotional reaction.

The worker described seeing scanned pages thrown into large containers alongside other loose paper.

Once the binding is removed and hundreds of pages are mixed with pages from other books, reconstructing the original becomes effectively impossible.

The knowledge exists digitally.

The physical copy no longer functions as a book.

For ordinary mass-market paperbacks, many people may consider that little different from recycling damaged books.

For genuinely scarce or out-of-print material, the issue becomes far more sensitive.

A surviving physical copy can contain value beyond text:

Marginalia.

Paper.

Binding.

Printing history.

Ownership marks.

Typographic details.

Illustrations.

Insertions.

Bookplates.

Annotations.

Even physical wear can matter to scholars.

A text file does not preserve all of that.

Were the Books Actually Rare?

Some were certainly sold through channels dealing in rare and out-of-print books.

The tracked shipment came through Biblio, a marketplace known for used, collectible, rare and difficult-to-find titles.

The employee interviewed by 404 Media described used books apparently originating from libraries, international shipments, foreign-language volumes and unusual government documents. Some appeared obscure enough that the employee suspected rarity.

But it would be inaccurate to describe every book processed at VGT3 as a priceless rare volume.

The worker said the facility handled all kinds of books, including brand-new sealed copies.

So the responsible claim is:

The operation processes large quantities of books, including used, obscure, foreign-language and potentially rare material.

That is different from claiming Amazon is systematically destroying priceless museum artifacts.

The Facility Was Receiving Books From Outside the United States

The employee recalled books appearing to come from overseas, including boxes connected with London institutions and large quantities of Japanese-language material.

German and Russian books were also mentioned.

That makes sense for AI training.

If the objective is model capability rather than resale, language diversity is extremely valuable.

Modern foundation models are expected to understand:

Japanese.

German.

Russian.

French.

Spanish.

Scientific terminology.

Legal writing.

Historical documents.

And countless specialized subdomains.

A random-looking international library can therefore be more valuable to an AI developer than a collection of current bestsellers.

What Did Amazon Say?

Amazon's response was remarkably short.

The company told 404 Media that it purchases books through commercial channels to help develop and improve products and services used by customers.

That statement confirms an important part of the story:

Amazon is buying the books.

It did not accuse 404 Media of fabricating the purchases.

It did not say the shipment accidentally entered the facility.

But Amazon's statement did not publicly specify:

  • Which AI model receives the resulting data
  • Whether the scans train Amazon Nova
  • How many books have been processed
  • How long VGT3 has operated
  • Whether similar facilities exist elsewhere
  • How titles are selected
  • What happens to every digital scan
  • Whether all scanned material is used in training

Those unanswered questions matter.

Did Amazon Confirm the Books Are Used for AI Training?

Not explicitly in the public statement quoted by 404 Media.

404 Media nevertheless identifies the operation as an AI-training facility based on its reporting, sources, tracking evidence and employee information.

Secondary outlets sometimes collapse those two things into one sentence:

“Amazon admitted it destroys rare books to train AI.”

That is stronger than the company's actual words.

A better formulation is:

404 Media reports that VGT3 destructively scans books for AI training; Amazon confirmed commercial book purchases for product development but did not publicly name the specific AI system or training program.

That distinction is worth preserving.

Amazon Is Not the First AI Company Caught Destructively Scanning Books

This is where the story becomes much bigger than Amazon.

The practice was already documented on an astonishing scale at Anthropic.

Court filings in Bartz v. Anthropic revealed that Anthropic purchased millions of physical books, removed their bindings, cut pages to size and scanned them into machine-readable files.

The paper originals were discarded.

According to court documents later unsealed, the operation was internally associated with Project Panama, an attempt to build a massive book corpus for AI development.

The documents revealed how valuable AI companies considered physical books once the easily available web had already been heavily harvested.

Amazon therefore appears to be participating in a broader industry movement rather than inventing a completely new technique.

The Anthropic case matters enormously because it addressed almost exactly the legal question raised by Amazon's apparent operation:

Can a company buy a copyrighted physical book, destroy it, scan it and keep a digital replacement?

In June 2025, U.S. District Judge William Alsup ruled on that question in Bartz v. Anthropic.

Anthropic had legally purchased millions of print books.

Its contractors stripped the bindings.

Pages were scanned.

The paper copies were discarded.

Digital copies entered Anthropic's internal library.

The court held that, on the facts before it, converting lawfully purchased print copies into internal digital replacements qualified as fair use.

One important reason was that Anthropic destroyed the source copy rather than retaining both a physical and digital version, and there was no evidence it distributed the digital library externally.

The Court Treated Pirated Books Very Differently

Anthropic had another source of books.

Pirate libraries.

The company had obtained millions of digital books from unauthorized repositories.

The court refused to treat that acquisition the same way.

Judge Alsup concluded that building a permanent central library from pirated copies could not simply be excused because the material might later be used for transformative AI training.

That created a major practical distinction:

Buy a physical copy lawfully and destructively replace it with an internal digital copy: potentially fair use under the facts of that case.

Download unauthorized copies from pirate libraries: a substantially different legal problem.

This helps explain why AI companies might be willing to spend enormous amounts acquiring physical books.

Buying paper can create a cleaner legal path than downloading an unauthorized PDF.

Does Buying the Book Automatically Give Amazon the Right to Train AI on It?

Not automatically.

Copyright law is more complicated than ownership of the physical object.

When you buy a book, you own that particular copy.

You can normally:

Read it.

Resell it.

Give it away.

Throw it out.

Cut it apart.

The first-sale doctrine gives owners substantial control over the particular physical copy they purchased.

But copyright in the text itself remains with the copyright holder.

Making reproductions creates a separate legal question.

That is where fair use enters.

The Anthropic ruling is highly relevant, but it does not create a universal rule declaring every imaginable AI-training use of every purchased book lawful.

Fair-use analysis is fact-specific.

The U.S. Copyright Office's major report on generative-AI training similarly emphasizes that different stages and purposes of copying may require separate legal analysis.

The Office also noted that licensing markets for AI training are developing and that the availability of viable licensing can matter when courts evaluate potential market harm.

This matters because the debate is evolving.

A company may argue:

We purchased the book.

The scan is internal.

Training is transformative.

The author may respond:

You created a valuable commercial digital asset from my work without paying for a digital or training license.

Those questions are not fully settled across every jurisdiction and use case.

The Amazon Investigation Is Not Evidence of “Destroying the Evidence”

One of the most dramatic claims circulating around the story is that Amazon destroys the books after scanning in order to eliminate evidence.

There is no evidence of that motive.

The physical destruction appears integral to the scanning process itself.

The spine is cut so pages can pass quickly through high-speed scanners.

Once that happens, the book has already been destroyed as a bound object.

Throwing the pages together afterward is a consequence of destructive scanning.

There may also be legal advantages to replacing one purchased physical copy with one internal digital copy, as the Anthropic ruling demonstrates.

But saying the company destroys books “to hide evidence” imputes a motive the reporting does not establish.

That phrase should be avoided.

Why Not Scan the Books Non-Destructively?

Technically, Amazon could.

Libraries and archives digitize rare materials without slicing off their bindings.

Specialized book scanners photograph pages from above.

Glass platens can hold books carefully.

Software corrects curvature.

Human operators can turn pages.

But all of those methods are slower.

More expensive.

More labor-intensive.

https://www.revlox.com/artificial-intelligence/the-ethics-of-ai-art-who-owns-creativity-in-the-digital-age/And poorly suited to millions of ordinary volumes.

If the goal is cultural preservation, those costs make sense.

If the goal is extracting text at industrial scale, destructive scanning is vastly more efficient.

This difference reflects a deeper philosophical conflict.

A librarian sees an artifact.

A dataset engineer may see 120,000 tokens.

What Is Lost When the Text Survives but the Book Does Not?

For many books, perhaps very little of unique cultural value.

A damaged mass-market paperback printed in hundreds of thousands of copies can be replaced easily.

But physical books sometimes contain information that OCR cannot capture.

Consider:

Handwritten notes.

Printing errors.

Inserted letters.

Owner signatures.

Library stamps.

Paper composition.

Binding construction.

Typography.

Illustration quality.

Evidence of censorship.

Wear patterns.

Historical provenance.

Researchers may care about those things independently of the printed words.

This is why blanket industrial destruction makes archivists uncomfortable.

A company focused on text extraction may have no practical mechanism for recognizing which individual copy possesses unexpected historical value.

Rare Does Not Necessarily Mean Expensive

The word “rare” also deserves clarification.

A book can be scarce without being financially valuable.

A technical proceedings volume from 1974 may have almost no collector demand.

A local agricultural report from 1961 may sell for $10.

An obscure regional history may sit unsold for years.

But if only a handful of copies remain, destroying one still reduces the surviving record.

AI companies may actually prefer precisely these obscure volumes because their text is less likely to exist online.

That creates an unusual inversion.

The books humans value least commercially may be especially valuable computationally.

AI Data Centers’ 3.4-Trillion-Gallon Water Footprint
A Ceres report links data-center electricity in seven U.S. states to 3.4 trillion gallons of freshwater withdrawals. Here’s what that number really means.

The Irony Is Hard to Ignore: Amazon Started as a Bookstore

Amazon's origins make the story culturally irresistible.

Jeff Bezos founded the company in the 1990s as an online bookseller.

Books were ideal for early e-commerce because millions of standardized titles could be catalogued using ISBNs.

Amazon then expanded.

Music.

Electronics.

Cloud computing.

Streaming.

Logistics.

Advertising.

Artificial intelligence.

Three decades later, an AirTag hidden inside a bulk book order apparently led journalists to an Amazon operation where physical books are converted into machine-readable data and discarded.

The symbolism almost writes itself.

The company that helped move books onto the internet is now reported to be moving books into AI.

The Dinosaur Logo Makes the Symbolism Even Stranger

VGT3's reported logo shows a dinosaur holding a book.

Under ordinary circumstances that might seem playful.

In the context of a warehouse where books are physically sliced apart, it became instantly controversial.

The imagery suggests consumption.

A dinosaur devouring knowledge.

That interpretation may not have been what its designers intended.

But once the operation became public, the logo became part of the story.

It is difficult to imagine a better visual metaphor for the anxiety surrounding AI training:

Human culture goes in.

A machine consumes it.

Something new comes out.

The original disappears.

AI Companies Have a Data Problem

Underneath the outrage lies a technical reality.

Frontier AI development requires enormous datasets.

The open internet provided an extraordinary initial resource.

But that resource has limitations.

Much of the best material is duplicated.

Some is low quality.

Some is copyrighted.

Some is behind paywalls.

And increasingly, some is synthetic.

Model developers therefore want datasets with:

High-quality prose.

Known provenance.

Specialized subject matter.

Multiple languages.

Long-form coherence.

Little contamination from previous AI generations.

Books satisfy nearly every requirement.

That makes their systematic acquisition almost inevitable.

The Internet Is Becoming Contaminated With AI’s Own Output

The word “contaminated” does not mean every piece of AI-generated text is bad.

The technical problem is statistical.

If later models train indiscriminately on content produced by earlier models, errors and simplifications can recursively propagate.

The 2024 Nature study on model collapse showed that models trained repeatedly on generated data can progressively lose information about the tails of the original distribution.

Researchers concluded that preserving access to genuine original data becomes increasingly important as generated material spreads.

Physical books effectively form a giant offline archive created before that feedback loop became widespread.

AI companies know it.

This Could Make Forgotten Books More Valuable Than Ever

For decades, digitization was often framed as a way to save forgotten texts.

Now the relationship may reverse.

A physical book that has never been digitized can become commercially desirable precisely because it is digitally absent.

That could produce surprising consequences for secondhand markets.

Obscure technical books.

Foreign-language volumes.

Local histories.

Old reference works.

Dissertations.

Proceedings.

Government reports.

Specialized manuals.

These may suddenly attract buyers who have no intention of reading them individually.

Their value lies in what they contribute to a corpus.

Booksellers May Have No Idea Who the Real Buyer Is

Marketplaces often mediate transactions.

The bookseller receives an order.

Packages the books.

Ships them.

Payment arrives.

But the ultimate destination may not be obvious.

That was part of what made the AirTag experiment necessary.

The seller suspected an AI connection but did not know which company was responsible.

The tracker converted suspicion into a physical destination.

That raises an ethical question distinct from copyright law:

Should sellers know when buyers intend to destroy the books?

Some booksellers may not care.

Others absolutely would.

Legality and Ethics Are Different Questions

Suppose Amazon lawfully buys a used book.

Suppose destructive scanning proves legally permissible.

That still does not resolve the cultural question.

Is there something wrong with destroying books that might otherwise remain available to readers?

For an ordinary paperback with millions of copies, probably few people would object strongly.

For an obscure academic text with only a handful of surviving copies, the answer becomes less obvious.

The legal system primarily protects copyright interests.

Cultural preservation involves different values.

Scarcity.

Access.

Historical continuity.

Material heritage.

Those do not always align neatly with copyright law.

There Is Also an Author-Compensation Question

When Amazon buys a used book, the author typically receives nothing from that resale.

That is normal.

Used-book markets have always worked that way.

But AI introduces another layer.

A single purchased copy could theoretically be digitized and incorporated into a model used by millions of customers.

The author may see no additional payment.

AI companies argue that training is transformative computational analysis.

Authors and publishers argue that their works provide valuable input into commercial systems.

The U.S. Copyright Office has recognized that emerging licensing markets may play an increasingly important role in resolving precisely this tension.

The Anthropic Decision Does Not End the Debate

The 2025 Bartz ruling was hugely important.

It found Anthropic's LLM training use transformative and treated conversion of lawfully purchased print books into internal digital replacements as fair use under the facts before the court.

But it was a federal district court ruling.

It did not create a Supreme Court rule binding every AI copyright dispute.

Different cases involve different facts.

Different markets.

Different outputs.

Different acquisition methods.

Different alleged harms.

The law remains contested.

What 404 Media’s AirTag Actually Proved

The investigation's strongest achievement is remarkably concrete.

Before the AirTag, booksellers had suspicions.

After the AirTag, they had a destination.

The tracked volume from an approximately 1,000-book order arrived at Amazon's Las Vegas facility.

Workers independently described bulk book intake and destructive scanning there.

404 Media then interviewed an employee who described cutting machines, scanners and discarded pages in detail.

That is much stronger evidence than an anonymous rumor about “AI companies buying books.”

What the Investigation Did Not Prove

Several questions remain unresolved publicly.

It did not establish:

  • The total number of books Amazon has scanned
  • The complete list of suppliers
  • Exactly which Amazon models use the text
  • Whether every scanned title enters model training
  • The percentage of books that are genuinely rare
  • Whether Amazon possesses unique or irreplaceable copies
  • How many scanning locations exist
  • Whether authors or publishers receive any licensing payment
  • The exact retention policy for scanned files

Those are natural next questions.

The reporting has identified the machine.

We still do not know the full size of the system around it.

The Follow-Up Employee Interview Made the Story Harder to Dismiss

After the original AirTag report, 404 Media published a detailed interview with someone who had worked at VGT3.

The employee described the operation from inside.

Books were unpacked.

Sorted.

Moved toward barcode processing.

Bindings were cut.

Loose sheets were scanned.

Scanned pages ended up mixed together in large containers.

The employee also described seeing international books, library materials and obscure documents.

That testimony strongly reinforces the original investigation.

This was not simply an AirTag happening to ping near an Amazon warehouse.

There was a corresponding book-digitization operation inside.

Are Libraries Selling Books That End Up in These Pipelines?

The VGT3 employee said some material appeared to have originated from library liquidation or deaccession channels.

Libraries routinely remove books.

Duplicates.

Outdated editions.

Low-circulation titles.

Damaged copies.

Materials outside collection priorities.

Those books may be sold or recycled.

There is nothing inherently improper about that.

But once AI companies become large-scale buyers, institutions may begin reconsidering where deaccessioned material goes—particularly if the books are scarce.

The downstream buyer suddenly matters.

The Hidden Pollution Behind the AI Data Center Boom
Texas AI data centers are turning to gas turbines and diesel generators. Explore the air pollution, climate risks, permit gaps, and cleaner alternatives.

Could Books Be Scanned Without Destroying Them?

Absolutely.

Non-destructive digitization is routine in libraries.

But it is slower and more expensive.

That means the core issue is not technological necessity.

It is optimization.

Amazon's reported pipeline appears optimized for throughput.

Cutting allows pages to move through automatic feeders rapidly.

From a dataset-engineering perspective, this is efficient.

From a preservation perspective, it is brutal.

Both descriptions can be simultaneously true.

Why Not Just License Existing E-Books?

Because the digital-rights landscape is fragmented.

Different publishers hold different rights.

E-book licenses may contain restrictions.

Some old books have no commercial e-book edition.

Others may have uncertain rights ownership.

Large-scale negotiations are expensive.

Physical books, meanwhile, can be purchased through established markets.

The Anthropic litigation showed that at least one court viewed destructive conversion of lawfully purchased copies favorably under fair-use analysis.

That creates a strong incentive to buy paper.

The Physical Book Has Become a Data Container

The strangest conceptual change may be this:

To a human reader, the book is the product.

To an AI-training pipeline, the physical book is merely a temporary container for information.

Once the information has been extracted, the container has fulfilled its purpose.

That is why the pages can go into a recycling bin.

The perspective is almost the reverse of traditional collecting.

Collectors preserve the object even after memorizing the words.

AI pipelines preserve the words and discard the object.

Is This the Destruction of Human Knowledge?

Not literally.

The scans preserve the textual information.

And unless the processed volume is the final surviving copy, other physical examples continue to exist.

Calling the operation the “destruction of human knowledge” is therefore rhetorically powerful but technically inaccurate.

What is destroyed is the physical artifact.

The intellectual content is being preserved precisely because Amazon wants to use it.

The more serious concern is whether systematic acquisition could reduce public access to scarce physical copies while transferring their contents into proprietary corporate datasets.

That is a much more precise criticism.

A Private AI Dataset Is Not the Same as a Public Digital Library

This difference is crucial.

Imagine a university scans a rare book and makes it searchable for scholars worldwide.

The physical copy is preserved.

The digital copy expands public access.

Now imagine a private corporation buys the volume, destroys it and places the text inside a proprietary training corpus.

That may produce useful technology.

But it does not necessarily expand public access to the book itself.

Users may gain indirect access to patterns learned from the text.

They do not necessarily gain the ability to read, verify or cite the original work.

That changes the cultural calculation.

The Story Is Really About Who Gets to Build the Next Library

Traditional libraries are built for retrieval.

You ask for a book.

The library gives you the book.

AI training libraries work differently.

The model ingests the books during development.

Users later ask questions.

The model generates answers influenced by statistical patterns learned across the corpus.

The source disappears into the system.

That makes provenance difficult to observe.

The modern race for books is therefore not simply about digitization.

It is about converting human-written culture into model capability.

The Ethics of AI Art: Who Owns Creativity in the Digital Age?
Explore the ethics of AI art, including copyright, ownership, training data, artist consent, originality, creative labor, and who owns creativity in the digital age.

Frequently Asked Questions About Amazon, VGT3 and the AirTag Investigation

Did an AirTag really track books to Amazon?

Yes.

404 Media supplied an Apple AirTag that a bookseller hid in one book from an approximately 1,000-volume order placed through Biblio.

The tracker ultimately reached Amazon's LAS8 complex in Las Vegas.

Who conducted the investigation?

404 Media journalist Emanuel Maiberg reported the investigation.

When was the original investigation published?

404 Media published the AirTag investigation on August 17, 2026.

How many books were in the tracked order?

The bookseller described an order of approximately 1,000 books.

Where were the books purchased?

The order was made through Biblio, a marketplace for used, rare and out-of-print books.

Where did the AirTag end up?

At Amazon's LAS8 facility in the Las Vegas area, specifically the part of the building associated with an operation called VGT3.

What is VGT3?

VGT3 is an Amazon operation inside the larger Las Vegas facility where employees described receiving, cutting and scanning physical books. 404 Media identifies it as an AI-training-data facility.

Does VGT3 really cut books apart?

Yes, according to Amazon employees interviewed and cited by 404 Media.

Workers described machines used to remove book bindings before scanning.

Why cut the spine off a book?

Removing the binding converts a book into loose pages that can pass through automated high-speed scanners much faster than a bound volume can be scanned.

Are the books usable afterward?

No.

The employee interviewed by 404 Media said pages were placed together in large bulk containers after scanning, making reconstruction impractical.

Are the books shredded?

Reporting varies in its terminology.

The directly documented process involves cutting the bindings, scanning the loose pages and discarding or recycling the resulting paper.

“Destructively scanned” is the most precise description.

Are all the books rare?

No evidence establishes that every book processed there is rare.

The employee said the operation receives many kinds of books, including new, used, foreign-language and apparently library-sourced material.

Were rare books included?

The tracked order came from a bookseller working through a used- and rare-book marketplace, and 404 Media specifically describes the tracked material as rare books.

Workers also reported unusually obscure material.

Did Amazon admit using the books for AI training?

Amazon confirmed purchasing books commercially to help develop and improve products and services.

Its quoted statement did not identify a specific AI model or explicitly describe the training program.

So how do we know AI is involved?

404 Media identifies VGT3 as an AI-training operation based on its investigation, employee accounts and the broader book-acquisition pattern.

Its follow-up employee interview specifically describes the warehouse operation it connected with AI training.

Is Amazon using the books to train Nova?

That has not been publicly established by the available evidence.

Amazon has not publicly identified which model or models receive the scanned text.

Why would AI companies want old books?

Older books provide large quantities of edited, human-generated text, including material unavailable on the public internet.

They also largely predate the modern explosion of AI-generated online content.

What is model collapse?

Model collapse describes degradation that can occur when generations of models are indiscriminately trained on data produced by previous models.

Research published in Nature demonstrated that recursive synthetic-data training can progressively distort the learned distribution.

Does AI-generated training data always destroy a model?

No.

The research concerns indiscriminate recursive training and shows that preserving genuine original data can mitigate degradation.

Synthetic data can still be useful when carefully created and managed.

Are books from before 2022 guaranteed to be human-written?

Not absolutely.

Machine-generated text existed before 2022.

But older published books overwhelmingly predate the mass availability of modern generative language models and therefore offer unusually reliable human-data provenance.

Has another AI company destroyed books this way?

Yes.

Anthropic's Project Panama involved purchasing millions of physical books, cutting off their bindings, scanning the pages and discarding the paper originals.

What was Project Panama?

Project Panama was Anthropic's large-scale physical-book acquisition and destructive-scanning program documented through court records.

The resulting digital library contributed material used in AI development.

A federal district court held that Anthropic's conversion of lawfully purchased physical books into internal digital replacements was fair use under the facts of the case.

The court separately held that its use of pirated copies to build a permanent library presented a different legal problem.

No.

It is highly relevant precedent, but fair use is fact-specific and the Amazon operation has not been adjudicated in the same case.

Can a company legally destroy a book it purchased?

Generally, the owner of a lawful physical copy has broad rights to dispose of that physical copy.

The more complicated issue is creating and using the digital reproduction.

Is destroying books after scanning a way of hiding evidence?

There is no evidence supporting that allegation.

Destruction occurs because the binding is physically removed to enable high-speed scanning.

The available reporting does not establish an evidence-destruction motive.

Could the books be scanned without destruction?

Yes.

Libraries routinely use non-destructive scanning systems.

Those techniques are generally slower and more expensive than cutting bindings and feeding loose sheets into high-speed equipment.

Why not preserve rare books before scanning them?

That is precisely one of the ethical concerns raised by the story.

There is no public evidence explaining what safeguards Amazon uses to identify historically important or uniquely valuable copies before destructive scanning.

Are authors paid when Amazon buys a used book?

Normally, authors do not receive a new royalty when an already-sold physical book is resold through the secondhand market.

AI training raises separate questions about whether additional licensing should apply to reproductions and training uses.

The Copyright Office has emphasized that fair-use analysis depends on the particular use and circumstances.

It has also recognized the growing importance of licensing markets for AI training.

No universal rule resolves every AI-training scenario.

Courts have reached important fair-use decisions, but outcomes depend on acquisition methods, purpose, market effects, outputs and other facts.

Is Amazon destroying unique human knowledge?

The physical books are destroyed during destructive scanning, but their textual contents are digitized.

It is therefore more accurate to say the operation destroys physical copies while preserving their text in digital form.

Why are bibliophiles angry?

Because a physical book can possess cultural, historical and scholarly value beyond its OCR text.

Destroying scarce books can permanently reduce the number of surviving physical copies available to readers, collectors and researchers.

How many books has Amazon scanned?

Amazon has not publicly disclosed a total.

404 Media says the facility receives very large shipments, but the overall scale remains unknown.

Does Amazon have other VGT3-style facilities?

The investigation does not establish the full number of Amazon destructive-scanning sites.

What happened to the actual AirTag book?

Its tracking signal led investigators to the Amazon facility associated with VGT3.

The reporting connects the shipment with the destructive-scanning operation, though the final physical fate of that exact individual volume was not independently observed page by page.

What is the biggest takeaway from the investigation?

The most important revelation is not that an AI company can cut a book apart.

Libraries, scanning companies and corporations have used destructive digitization for years.

What matters is scale and purpose.

Frontier AI companies need enormous amounts of high-quality human-generated data.

The open web is increasingly saturated with duplicated, low-quality and synthetic material.

Physical books therefore represent one of the largest remaining reservoirs of text that has not already been absorbed into digital training corpora.

And 404 Media's AirTag turned that abstract demand into something physical.

A bookseller accepted an anonymous bulk order.

One book carried a tracker.

The shipment crossed the United States.

The signal stopped at Amazon.

Inside, workers described an industrial pipeline that turns bound books into loose pages, loose pages into scans and scans into data.

The books are not being burned because Amazon hates literature.

They are valuable for exactly the opposite reason.

AI companies want what is inside them.

That may be the strangest part of the entire story.

For centuries, books were valuable because people could hold them, preserve them, lend them and pass them between generations.

In the AI era, some books have acquired a different kind of value.

They are uncontaminated human data.

A forgotten history sitting on a dusty shelf can contain sentences no web crawler has ever seen.

An obscure technical manual can hold terminology absent from modern websites.

A discontinued academic volume can offer hundreds of pages of coherent expert writing created before generative AI flooded the internet.

Suddenly, the old book nobody wanted becomes computationally precious.

The VGT3 story therefore is not really about one AirTag.

It is about a transition in what a book means.

To the bookseller, it is inventory.

To the reader, it is a work.

To the collector, it is an artifact.

To the historian, it may be evidence.

To the AI company, it can become something else entirely:

training data waiting to be extracted.

Once the machine sees the book that way, the binding is no longer sacred.

It is simply in the way.

Revlox Magazine Newsletter

Get the latest Revlox stories, cultural essays, and strange discoveries, handpicked for your inbox.

A cleaner edit of the week’s standout reporting, visual culture, historical mysteries, and deeper reads from across the magazine.

By signing up, you agree to the Terms & Conditions and acknowledge the Privacy Policy.

Advertisement

More stories from Revlox Magazine

Read more

Advertisement

Advertisement

Advertisement