THEBLACKBOOK AUDIT
Investigation · Who Controls What You Get to Know

They cut up the books to feed the machine.

Anthropic gave it a codename: Project Panama — an internal effort, in its own words, to “destructively scan all the books in the world.” The company bought millions of print books, sheared off their spines with a hydraulic cutter, ran the loose pages through high-speed scanners, and sent the remains to a recycler. The paper record went in one end; a private model came out the other.

The most unsettling part isn't that it was a crime. Cutting up books you've bought is legal, and a court ruled that training an AI on them is fair use. The scandal is that the frictionless, lawful path to building a machine was to physically destroy the commons and lock what it held inside something you have to pay to query.

COURT EXHIBIT
Document
ReceiptFACT

Anthropic's own definition of Project Panama

Anthropic internal document (Bartz v. Anthropic PBC) · Unsealed internal memo · N.D. Cal. No. 3:24-cv-05417 · Unsealed 2026

In its own internal words — surfaced in the copyright litigation and reported by the Washington Post — Anthropic described Project Panama as a plan to destructively scan the world's books.

Project Panama is our effort to destructively scan all the books in the world.
View source →
§0 · Disclosure

We have a conflict of interest, and we are stating it up front. Anthropic — the company at the center of this piece — makes Claude, the AI assistant The Black Book Audit uses in its own research and drafting workflow. We are reporting on our own tool vendor. We have not softened a single fact because of it; if anything, a site that won't name the company behind its own assistant has no business grading anyone else. Every claim below is sourced to the court record and mainstream reporting. If we have any of it wrong, the Standing Invitation below is open to Anthropic like anyone else.

§1 · Summary Brief

What this page is about

To train large language models, AI companies need enormous amounts of high-quality human text — and printed books are the gold standard: dense, edited, authoritative, and (for older titles) free of the AI-generated “slop” now polluting the open web. In early 2024 Anthropic ramped up what internal documents called Project Panama: an effort to destructively scan books at industrial scale, hiring a vendor to convert somewhere between 500,000 and two million volumes in roughly six months.

Destructive scanning means what it sounds like. The books were bulk-purchased, their bindings sliced off with a hydraulic cutter, the pages fed through production scanners, and the physical remains recycled. Separately, Anthropic had also downloaded more than seven million books from pirate “shadow libraries.” A federal judge drew a sharp line between the two, and the piracy half of the case ended in the largest copyright settlement in US history — about $1.5 billion. This page grades exactly what happened, and what it means for who controls the record of human knowledge.

What we are NOT saying
We are not saying AI training is theft as a matter of law — a court ruled the opposite (training on lawfully acquired books is fair use). We are not claiming the books destroyed were mostly rare or irreplaceable editions; the record points to common used copies bought in bulk, and we grade the “vanishing rare books” framing down accordingly. The documented story is narrower and, we think, more important: the lawful path to building these systems runs through destroying physical books and enclosing what they contain inside a proprietary product.
§2 · Graded Claims

The record, claim by claim

Anthropic ran 'Project Panama' to destructively scan millions of print books.

FACT

Internal documents surfaced in litigation describe Project Panama as Anthropic's 'effort to destructively scan all the books in the world.' In early 2024 the company engaged a scanning vendor to convert an estimated 500,000 to two million books over about six months: bulk-purchased print copies had their bindings cut off by a hydraulic cutter, the pages were run through high-speed scanners, and the destroyed volumes were sent to a recycler. Reported by the Washington Post and Ars Technica from the court record.

Destroying purchased books, and training on them, was legal — that's the point.

FACT

In June 2025 Judge William Alsup ruled on summary judgment that using books to train an LLM is 'quintessentially transformative' and therefore fair use, and that Anthropic's destructive scanning of books it had lawfully bought was permissible under the first-sale doctrine (you may do as you like with a copy you own). In other words, the shredding wasn't the illegal part. The lawful, frictionless route to a training corpus was to buy books and destroy them.

Anthropic also pirated 7M+ books — and that cost it ~$1.5 billion.

FACT

Alongside the books it bought, Anthropic had downloaded more than seven million books from pirate shadow libraries such as LibGen and PiLiMi to build a 'central library.' Judge Alsup ruled that piracy was NOT fair use — it was infringing — and set a damages trial. Anthropic settled instead: about $1.5 billion, roughly $3,000 per work for some 500,000 titles, the largest copyright settlement in US history, granted final approval in 2026.

Anthropic isn't alone: Meta trained on pirated books its own staff flagged.

FACT

In Kadrey v. Meta, court filings showed Meta trained its LLaMA models on the Books3 / LibGen shadow-library datasets, and that Mark Zuckerberg approved use of LibGen despite internal warnings that it was 'a dataset we know to be pirated.' (On the specific record there, the court found Meta's use fair use — a separate outcome from Anthropic's piracy ruling.) The pattern across the industry is the same: the world's books, ingested wholesale, with permission treated as an afterthought.

The 'destroying rare, irreplaceable editions' framing overstates it.

SOME SMOKE

Some coverage frames this as AI firms vanishing rare book editions. Graded SOME SMOKE and corrected: the destroyed books were overwhelmingly common, bulk-purchased used copies, not rarities — the same title exists in millions of other copies and in libraries. The loss that's real isn't bibliographic scarcity; it's the enclosure. Public, human-made knowledge is converted, at scale, into a private model you must pay to access — and the physical copies that fed it are pulped. That's the defensible version of the alarm.

§3 · Record vs Narrative

The shredder is real; the “lost rarities” are not the story

  • Legal is the scandal, not the defense. “It was all lawful” is usually where a story ends. Here it's the point: the cheapest compliant way to build a frontier model was to destroy the commons. When the law rewards enclosure, “we broke no rules” is an indictment of the rules.
  • Enclosure, not scarcity. We don't claim priceless editions were lost — they mostly weren't. What's lost is public access to a public inheritance: books written by humans, digested into a product owned by a company, with the paper pulped behind it.
  • Piracy vs. purchase — keep them straight. Training on bought books = fair use. Destroying bought books = legal. Downloading 7 million pirated books = infringing, and it cost $1.5 billion. Conflating these is how the story gets attacked; we keep the lines the court drew.
§4 · Why It Matters

Who owns the record of what we know

This is the newest chapter of an old story: control of knowledge. Robert Maxwell did it with school textbooks and with the scientific journals he locked behind Pergamon's paywalls; six companies did it with the news you're allowed to see. The AI era does it at a scale and a speed those men could only dream of — ingesting the written inheritance of humanity into a handful of proprietary models, and pulping the originals on the way in. It belongs alongside Media Ownership and the coming Who Controls What You Get to Know hub because it answers the same question those pieces do — just with a hydraulic cutter instead of a printing press.

§5 · FAQ

Questions worth taking seriously

Aren't you biased, since Anthropic makes the AI you use?

We disclose exactly that at the top of the page. The mitigation for a conflict is transparency plus sourcing — and every claim here is tied to the court record and mainstream reporting (Washington Post, Ars Technica, NPR). We would rather name our own vendor plainly than pretend the conflict doesn't exist.

If it was legal, what's the problem?

That is the problem. The lawful, cheapest route to building a frontier AI was to buy and physically destroy books and enclose their contents in a private model. The piracy half cost $1.5 billion; the destruction half cost nothing, because the law allowed it. When the rules reward pulping the commons, “it was legal” is a critique of the rules, not a defense of the act.

§6 · Standing Invitation

If you are named on this page

If you represent Anthropic, Meta, or anyone named here and believe we have a fact wrong or a framing unfair, we want to hear from you and we carry responses in full. Reach us through the contact channels on our mission page.

§7 · Sources

The record

▦ Ledger gaps

Help us fill these lines.

This entry is graded on what’s on the public record. These are the blanks we know about. If you can source one, you’re rebuilding the ledger with us.

  • OpenHow many books were actually destructively scanned under Project Panama, given the piece's own estimated range of 500,000 to two million volumes?Help fill this →

Notify me when a gap is filled

We'll email you when we fill one of the gaps above.

By signing up you agree to receive emails from The Black Book Audit. Unsubscribe anytime.