Millions of books were cut and destroyed to train the Claude AI model
Internal documents revealed in a legal battle with authors expose how the AI giant acquired millions of books, dismantled and scanned them - and then discarded the physical copies. The goal: to build a massive database of high-quality human text to train Claude.

New details about one of Anthropic's unusual projects are being revealed in The Guardian, following internal company documents submitted as part of a legal battle with authors in the US. According to the documents, the company developing the artificial intelligence model Claude purchased millions of printed books, dismantled them, scanned the pages, and then destroyed the physical copies. The operation was codenamed "Project Panama," and in an internal memo from April 2024, its goal was defined in words that are hard to miss: "our effort to destructively scan all the books in the world." The secrecy was also intentional. The document stated that the company uses a codename because it does not want it known that it is working on the project, and employees were asked not to speak about it publicly.
Behind the project was a central problem for artificial intelligence companies: to train a model like Claude, you need huge amounts of high-quality text. Anthropic specifically sought books from the period before the spread of AI models, because they contain human text that was not influenced by content already created using artificial intelligence. From the company's perspective, books were particularly high-quality training material thanks to the long and complex texts they contain. In court documents, it was written that Anthropic particularly valued the "well-curated facts, organized analysis, and compelling fictional narratives" in the books — content that could teach Claude to write more accurately and persuasively.
To obtain this material, Anthropic spent millions of dollars on purchasing books and turned to distributors and retailers to buy them in large quantities, including used copies. The books were then transferred to third-party vendors, who removed the covers, cut the book spines, and separated the pages to quickly pass them through scanners. From each book, a PDF file was created that included the scanned pages alongside text that a computer can read and process. At the end of the process, the paper copies were discarded or sent for recycling. This is exactly the meaning of "destructive scanning": you don't scan a book and keep it on the shelf, but physically dismantle it to turn it into a digital file — and the original copy does not survive the process.
However, the books purchased legally were only part of Anthropic's database. According to the documents, the company also used huge amounts of books that came from pirate repositories: about five million copies from LibGen, about two million from Pirate Library Mirror, and about 183 thousand from Books3. Anthropic also considered obtaining orderly licenses from copyright holders, but within the company, the legal, business, and practical process involved was described as cumbersome. The use of pirated books later became a central part of the legal battle with authors, and the company reached an out-of-court settlement of $1.5 billion regarding the pirated materials.
The court, however, made a clear distinction between books that Anthropic bought and copies that came from pirate sources. Judge William Alsup ruled that using books purchased legally for the purpose of training an AI model can be considered "fair use" — that is, use that is permitted in certain cases even without additional permission from the copyright holder. According to the ruling, Claude did not redistribute the books but learned from thousands of works about grammar, structure, and ways of writing to produce new text. Turning printed books purchased legally into digital files for Anthropic's research library was also not considered an infringement. In contrast, regarding pirated books, the judge was unequivocal: downloading a book from a pirate site is a copyright infringement in itself.
The details now being revealed illustrate primarily how much artificial intelligence companies are still dependent on human creation to build systems that are supposed to compete with it. According to court documents, Anthropic sought to create a central research library of "all the books in the world," which would be kept "forever" and allow it to select from it various book databases for training future models. The company stated that its data acquisition plans do not buy and destroy "rare" or "ancient" books. However, the project leaves behind a broader question: when millions of printed books become, for the AI industry, raw material that can be bought, cut, scanned, and then discarded — "Project Panama" may be not only a story about the way Claude was trained, but a glimpse into how the race for AI data is also changing the attitude towards books themselves.





