Posted inPLUG IN

Anthropic bought mns of physical books, destroyed them, and a federal court said that’s fine

Project Panama gave Anthropic’s models mns of books that no one else will ever read — and the ruling that blessed it handed every AI company a playbook

📚🗑️ An AI company bought mns of physical books, cut off their spines, fed the pages through scanners, and got rid of the rest. The books were gone. The data? In Anthropic’s servers. The story isn’t new. The Washington Post reported it back in January, based on more than 4k pages of unsealed court documents. But a little over two weeks ago, on 20 July, a federal court ruling turned what Anthropic has dubbed Project Panama into an AI controversy that is dominating social media.

What’s Project Panama?

In early 2024, executives at Anthropic — Claude’s parent company — began work on an initiative they sought to keep under wraps. An internal planning document, later unsealed in a federal copyright lawsuit, described the project’s intentions as follows: “Project Panama is our effort to destructively scan all the books in the world. We don’t want it to be known that we are working on this,” The Post reported. Tens of mns of USD were then spent in the following year to accomplish just that.

Anthropic turned to Tom Turvey — former head of partnerships for Google Books — to lead the charge. The company purchased used and out-of-print books in bulk from a slew of vendors. These books were then re-routed to contractors, where hydraulic machines sliced off their spines, and scanners processed the pages of up to 11k books per day, according to Gadget.

…No one would have access to these books but Anthropic and its models.

💡FUN FACT- Anthropic co-founder and president Daniela Amodei is a literature major.

Project Panama wasn’t Anthropic’s first rodeo

Court documents revealed that Anthropic co-founder Ben Mann personally downloaded large hauls of pirated books from shadow libraries including LibGen. Another court record saw Anthropic CEO Dario Amodei call negotiating with publishers a “slog,” choosing instead to pirate the works. The presiding judge ruled that Anthropic chose piracy when legitimate methods were on the table.

The legitimate alternative? Project Panama. Physically purchasing and destroying books sidesteps the licensing problem: once you own a copy, what you do with it is no one’s business.

But why old, rare books? For Anthropic, pre-2022 books were particularly valuable. Books published before AI-generated content became widespread are considered higher-quality training material — feeding AI-generated text back into models risks what researchers call model collapse.

BACKGROUND- Training LLMs requires enormous quantities of high-quality human-produced text. For most of the past decade, it was in generous supply. We’re now running out, according to the 2026 AI Index Report by Stanford University. The remaining internet text? No longer suited for the task as AI-generated content floods the web. Simply put, internet data is now contaminated.

The case — brought forth by a slew of authors — produced two outcomes: On the piracy matter, the presiding judge ruled against Anthropic: downloading mns of books from shadow libraries was not fair use and exposed the company to liability. Anthropic settled that portion of the case for USD 1.5 bn — the largest copyright settlement in US history — covering almost 500k titles. On Project Panama? Well, the judge decided that was fair play for the reasons mentioned earlier. The problem is this sets a dangerous precedent, one that allows for an industry-wide practice: buy, scan, and destroy.

And that practice is already spreading

What gave a story initially reported in January ongoing virality were reports by 404 Media and Fortune revealing that it’s not just Anthropic. Booksellers in the Netherlands, Germany, and Switzerland have, of late, reported unusual bulk buying patterns, with some buyers requesting quotes for thousands of titles to be shipped to countries around the world, including China. For several sellers, the requests were dismissed as scams, as they simply weren’t used to this scale of requests.

404 Media also reported that ISBNdb, a firm that maintains a large book database, advertised a service brokering bulk book acquisitions for AI clients — orders ranging from 1k to 1 mn books, with non-disclosure agreements included as part of the service. ISBNdb later backtracked after 404 Media went public, deleting their claims from their website, the tech zine reported.

The Library of Alexandria, take two

This is only the beginning. Several AI giants have similarly been involved in book piracy scandals, and the recent ruling now means the buy-scan-destroy model has a legislative precedent for the tech giants looking for a “lawful” way around training their models that isn’t piracy — even if it means death to physical media.

What the ruling doesn’t address is what happens when the books being acquired are exceptionally rare or exist in only a handful of surviving copies. Unlike Google Books — which created a publicly searchable index while preserving physical copies in partner libraries — the corpus produced by Project Panama is entirely private, and the originals are gone.

The cultural and preservation concern driving much of last week’s viral reaction isn’t about paperback bestsellers with mn-copy print runs, but rather about the possibility that a machine in a warehouse somewhere is processing books that no library holds, no digital archive has, and no reader will ever be able to find again.

Arabic-language texts are safe… for now

It is, then, perhaps a good thing that Arabic has thus far been underrepresented in the AI sphere. The buy-scan-destroy pipeline is almost entirely built on English-language books, and Arabic works have largely been spared. Much of the Arabic-language data available for AI training today consists of translated English content, often missing cultural nuances and failing to reflect real-world language use accurately.

A 2025 Welo Data report found that large language models perform significantly worse in Arabic than in English on key tasks, which means that while attention may not be on Arabic, it eventually will be, and we run the risk of having the same done to our native language works.

Some institutions in the region are presenting ways around this without the need for destruction. The Hindawi Foundation, an Egyptian nonprofit publisher, has made a corpus of 1,745 Arabic books — covering fiction, nonfiction, poetry, and children’s literature — available at no cost, and the collection has been used in training Arabic language models.

How we can protect our texts: Egypt’s National AI Strategy includes an initiative for an AI patent granting system — but contains no provision addressing how training data should be acquired, attributed, or compensated. If history is not to be repeated, a governance framework needs to be developed to address any potential — present or future — destruction of physical media for the sole benefit of AI.