Dossier mode
Project Panama
The same investigation, restaged one beat at a time. Drive it with the arrow keys, space, or autoplay. Nothing is cut from the piece — long runs are split across frames. Read the full investigation or open the Who Controls What You Get to Know hub.
They Cut Up the Books to Feed the Machine.
Anthropic gave it a codename: Project Panama — an effort, in its own words, to 'destructively scan all the books in the world.' It bought millions of print books, sheared off their spines, scanned the loose pages, and recycled the remains.
§0 · Disclosure — we have a conflict of interest, stated up front.
- Anthropic, the company at the center of this piece, makes Claude — the AI assistant this site uses in its own research and drafting workflow. We are reporting on our own tool vendor.
- We have not softened a single fact because of it. A site that won't name the company behind its own assistant has no business grading anyone else.
- Every claim below is sourced to the court record (Bartz v. Anthropic PBC, N.D. Cal.) and mainstream reporting.
The scandal isn't that it was a crime — cutting up books you bought is legal, and a court ruled training an AI on them is fair use. The scandal is that the frictionless, lawful path to the machine was to physically destroy the commons and enclose it.
The paper record went in one end; a private model you must pay to query came out the other.
“Project Panama is our effort to destructively scan all the books in the world.”
Anthropic's settlement over 7M+ books it downloaded from pirate shadow libraries — roughly $3,000 per work for some 500,000 titles, the largest copyright settlement in US history (final approval 2026).
NPR; Authors Guild
Anthropic ran 'Project Panama' to destructively scan millions of print books.
Internal documents surfaced in litigation describe it as Anthropic's 'effort to destructively scan all the books in the world.' In early 2024 the company engaged a vendor to convert an estimated 500,000 to two million books: bulk-purchased print copies had their bindings cut off by a hydraulic cutter, the pages were run through high-speed scanners, and the destroyed volumes were sent to a recycler. Reported by the Washington Post and Ars Technica from the court record.
Judge Alsup ruled that training LLMs on purchased books is fair use — the shredding wasn't the illegal part.
In June 2025 Judge William Alsup ruled on summary judgment that using books to train an LLM is 'quintessentially transformative' and therefore fair use, and that destructive scanning of books Anthropic had lawfully bought was permissible under the first-sale doctrine. The lawful, frictionless route to a training corpus was to buy books and destroy them.
Separately, Anthropic pirated 7M+ books to build a 'central library' — that piracy was infringing, and it cost ~$1.5 billion.
Alongside the books it bought, Anthropic had downloaded more than seven million books from pirate shadow libraries such as LibGen and PiLiMi. Judge Alsup ruled that piracy was NOT fair use — it was infringing — and set a damages trial. Anthropic settled instead: about $1.5 billion, roughly $3,000 per work for some 500,000 titles, the largest copyright settlement in US history.
Anthropic isn't alone: Meta trained on pirated books its own staff flagged as 'a dataset we know to be pirated.'
In Kadrey v. Meta, filings showed Meta trained its LLaMA models on the Books3 / LibGen shadow-library datasets, and that Mark Zuckerberg approved use of LibGen despite internal warnings. (On that record the court found Meta's use fair use — a separate outcome from Anthropic's piracy ruling.) The industry pattern is the same: the world's books ingested wholesale, permission treated as an afterthought.
That AI firms are vanishing 'rare book editions.'
Corrected: the destroyed books were overwhelmingly common, bulk-purchased used copies, not rarities — the same titles exist in millions of other copies and in libraries. The real loss isn't bibliographic scarcity; it's the enclosure: public, human-made knowledge converted at scale into a private model you must pay to access, while the physical copies that fed it are pulped. That's the defensible version of the alarm.
Why it matters, and where it connects.
This is the newest frontier of who controls what you get to know: not a mogul buying a newspaper, but an AI company buying and pulping the printed commons and reselling access to what it held. The enclosure is the through-line — public knowledge locked inside a private, metered model, the same fence Pergamon built around research and Media Ownership tracks around the news. We report it straight, including on our own tool vendor, because a site that softens the facts about the company behind its assistant forfeits the right to grade anyone else.