AI Firms Are Quietly Buying and Destroying Books to Train AI

 

AI firms buying and destroying books to train AI — ZTS Infotech

It sounds like the plot of a dystopian novel, except the paper trail is real. Federal court filings unsealed this year, combined with reporting published July 28 and 29, 2026 by 404 Media, The Washington Post, Tom’s Hardware and Decrypt, confirm that major AI developers have been quietly purchasing millions of physical books — including rare, out-of-print titles that exist nowhere else — running them through hydraulic cutting machines to remove the spines, feeding the loose pages into industrial scanners, and shredding what remains. For leaders who have spent the past two years integrating generative AI into their businesses, the story is a reminder that these tools were built on foundations most vendors would rather not discuss in a sales call.

Inside Project Panama: What the Court Filings Reveal

The clearest window into this practice comes from Anthropic, the company behind the Claude family of models. An internal program known as Project Panama was described in an unsealed court document as, in the company’s own words, “our effort to destructively scan all the books in the world.” The filing noted the project used a deliberately vague code name because the company did not want the effort known publicly.

That detail matters. Companies rarely obscure a sourcing decision unless they suspect the public reaction would be unfavorable. Project Panama was not a side experiment — it was, by Anthropic’s own description, an ambition to scan the world’s books at scale, with the physical originals discarded once digitization was complete.

Anthropic Project Panama quote:

From the unsealed court filing — Anthropic’s internal description of Project Panama.

Why Books, and Why Now

The internet has run out of clean data

AI labs are running into a data problem that seemed theoretical a few years ago. So much of today’s web is now populated by AI-generated content that models trained on it tend to degrade in quality. Books published before the generative AI boom, roughly pre-2022, represent a large, dense, professionally edited body of human-written text untouched by AI-generated filler — one of the last large untapped reservoirs of trustworthy training material.

A court ruling opened the door

In June 2025, a federal judge ruled that scanning a legally purchased physical book and then destroying the original constitutes fair use under U.S. copyright law. That ruling — distinct from Anthropic’s separate $1.5 billion settlement over claims involving pirated digital copies — gave AI developers a legal pathway to build training libraries from print books without the copyright exposure tied to scraping text from the open web. Since then, per the reporting cited above, book buying has accelerated sharply, particularly from April 2026 onward.

The Marketplace Behind the Buying Spree

Much of this activity reportedly moves through a platform called ISBNdb, which now facilitates anonymous bulk orders of up to one million books per transaction, backed by strict non-disclosure agreements that keep buyer identities confidential. One independent bookseller told 404 Media that weekly sales volume jumped from around 20 books to several hundred almost overnight. Some titles being purchased and destroyed reportedly exist in fewer than ten copies worldwide — meaning a handful of buying decisions, made under NDA with no public accountability, could permanently eliminate access to editions libraries and archives never digitized.

It is worth being precise about what is and is not confirmed. Every buyer’s identity moving volume through platforms like ISBNdb remains undisclosed, and the NDAs are, by design, meant to keep it that way. What is documented is Anthropic’s own Project Panama program, the ruling that enables the practice industry-wide, and the market response booksellers are reporting.

Key figures on AI book-buying and destructive scanning

By the numbers — key figures from the court filings and July 2026 reporting.

What This Means for Business Leaders

For executives evaluating AI vendors, licensing AI-powered content tools, or building products on top of large language models, this story is less about book publishing and more about data provenance — a due-diligence question quickly becoming as important as security or uptime.

  • Vendor due diligence now includes training data. Procurement teams that once asked only about privacy and security certifications should ask AI vendors how their models were trained and what the sourcing chain looked like.
  • Reputational exposure is transferable. A company that adopts an AI tool inherits some of the reputational risk tied to how that tool’s underlying model was built.
  • Transparency is becoming a differentiator. As scrutiny of training practices grows, vendors and the businesses that integrate their tools will be judged on how openly they discuss their supply chain, not only on performance benchmarks.

Expert Perspective

The legal reasoning behind the June 2025 ruling is sound on its own narrow terms: buying a book and scanning your own copy is a different act than downloading a pirated file. But narrow legal correctness and broad public acceptability are not the same thing, and the gap between them is where this story lives.

What should concern business leaders is not the scanning itself but the pattern around it: a legally sanctioned practice conducted through anonymous, NDA-bound bulk transactions structured so buyers cannot be identified. That is not how companies behave when confident the public will approve. Expect this to accelerate calls — from authors’ groups, librarians, and increasingly enterprise customers — for AI companies to publish sourcing disclosures alongside their model cards. Regulators in the EU and UK, already pushing for training-data transparency requirements, are likely to treat this reporting as further justification for mandatory disclosure rules.

For AI vendors, the more durable advantage may belong to those willing to document sourcing practices proactively, rather than waiting for the next unsealed court filing to do it for them.

Key Takeaways

  • Anthropic’s internal “Project Panama” program bought millions of physical books and destroyed the originals after digitizing them.
  • A June 2025 federal ruling found destructive scanning of a legally purchased book qualifies as fair use, giving AI developers legal cover.
  • Pre-2022 books are prized because they are uncontaminated by AI-generated content now flooding the open web.
  • ISBNdb reportedly enables anonymous bulk book purchases of up to one million titles per transaction, backed by NDAs.
  • Some destroyed titles reportedly exist in fewer than ten copies worldwide, raising preservation concerns.
  • Business leaders should treat training-data provenance as a standard vendor due-diligence item.
  • Transparency about data sourcing is becoming a competitive differentiator as scrutiny increases.

Conclusion

This story will not stay contained to publishing and copyright circles. As AI tools embed themselves in everyday business operations, how those tools were trained is moving from a legal footnote to a boardroom question. Companies that get ahead of that shift — asking vendors harder questions now and favoring providers willing to answer them — will be better positioned than those caught flat-footed by the next unsealed court filing. Expect further disclosures like this one in the months ahead.

  • bm
    Writen by Anirban Das