A hydraulic cutter slices through the binding. The pages, freed, pass through a high-speed scanner. The paper goes off for recycling. Repeat the operation between 500,000 and two million times.

This is project Panama, the internal code name of a programme launched in early 2024 by Anthropic, the company behind the assistant Claude — whose latest generation of models we analysed. Its ambition, as set out in an internal planning document made public during the court proceedings: destructively digitise every book in the world.

One useful point about the chronology, because it is almost always skipped: the information is not new. The document was unsealed in January 2026 as part of a class action filed in 2024, and the Washington Post published its investigation on 27 January 2026, after examining more than 4,000 pages of court exhibits. It was in late July and early August 2026 that the affair went viral, carried by social networks — six months after the facts were revealed.

It is presented almost everywhere in the same way: a company destroys books, it is legal, it is scandalous. All three statements deserve to be taken one by one, because two are accurate and one is false — and because what is really at stake is not what people think.

What project Panama actually did

Let us start with the documented facts.

In February 2024, Anthropic hires Tom Turvey, former head of partnerships for the Google Books project — the mass digitisation operation Google ran in the 2000s, which survived a long court battle. His mission, as reported in the proceedings’ documents: to obtain “all the books in the world”.

The chain put in place is industrial. The volumes are bought on the second-hand market, notably from Better World Books. They are then cut — the documents speak of books “cleanly cut” — to separate the pages from the binding. The sheets pass through high-speed scanners supplied by Datamation Imaging Services. Then the paper is handed to recyclers.

The scale reported runs from 500,000 to two million volumes, for a budget of several tens of millions of dollars.

It is worth measuring what that figure represents. Two million books is the order of magnitude of the holdings of a very large French university library. Except that here, no copy survived the operation.

Why paper, when the internet is free?

This is the most interesting question, and the answer says a great deal about the state of the industry.

Language models have exhausted the usable web. What remains available online is overwhelmingly recent, redundant content of uneven editorial quality. And a model learns to write by imitating what it reads.

The internal documents are explicit about the objective: to reach works older than what circulates online, in order to teach the model to write well rather than to reproduce what the teams describe as a low-quality register of writing specific to the web.

In other words: paper books were not chosen out of nostalgia, but because they contain something that no longer exists anywhere else. Worked, edited, proofread prose — and a considerable share of knowledge that has never been digitised.

It is the same finding of scarcity that explains, at the other end of the chain, why real-world data has become the sector’s strategic prize, which we documented when analysing the Chinese open model strategy: compute can be bought, quality data cannot.

Let us come to the point that makes most of the noise: yes, this part of the operation was indeed found legal in the United States.

In an order of 23 June 2025, federal judge William Alsup, of the San Francisco court, rules in Bartz v. Anthropic — a class action brought by the authors Andrea Bartz, Kirk Wallace Johnson and Charles Graeber. He finds that digitising lawfully purchased books to train a model falls under fair use.

The heart of the reasoning fits in two words: format shifting. Since the paper copy is destroyed after digitisation, the court considers that no additional copy is created, but that one medium is substituted for another. The judge’s wording is plain:

“Each purchased print copy was copied in order to save storage space and to enable searchability in digital form. The print original was destroyed. One replaced the other.”

The logic holds. It also has a counter-intuitive consequence that has to be underlined: it is precisely the destruction that makes the operation legal. Had Anthropic kept the books after scanning them, it would have ended up with two copies — the paper and the digital one — and the format-shifting argument would have collapsed.

Destroying them was not collateral damage. It was a legal condition.

The precedent everybody had in mind: Google Books

Hiring Tom Turvey was no accident of the curriculum vitae. He brought with him the only favourable American precedent on mass book digitisation — and Anthropic clearly intended to play the same tune.

A reminder of the facts. From 2004, Google digitises millions of books to build a search engine inside their content. The Authors Guild sues. The litigation lasts ten years. On 16 October 2015, the Second Circuit Court of Appeals rules in Google’s favour: digitising around four million copyrighted works, for commercial purposes, falls under fair use. The Supreme Court declines to take the case, leaving the decision final.

The court’s reasoning rested on four pillars: the use was highly transformative (it creates a search tool, not a substitute library), public display was limited to short snippets, the service was not a market substitute for the original works, and Google’s commercial character was not enough to rule out the exception.

It is precisely on those four points that the Anthropic case differs, and that needs saying, because the comparison gets used rather quickly.

Google indexed to make books findable: the reader saw a snippet, then went and bought the work. Anthropic ingests the full text to produce a machine that writes in place of the author, in the same register, on the same subjects. The market-substitute question does not arise in anything like the same terms. And the technical point mentioned below — the demonstrated ability of some models to reproduce whole passages from works overrepresented in training — weakens the analogy further.

Judge Alsup, for that matter, did not validate project Panama by relying on Google Books. He validated a far narrower point: the format shift of a purchased copy that is then destroyed. That is a librarian’s reasoning, not a search engine’s. And it says nothing about what the model does with the text afterwards.

What the viral story leaves out: the seven million other books

Here is the part that systematically disappears from the summaries, and without which the case makes no sense.

The same judge, in the same decision, refused the benefit of fair use for another part of the file. Before and during project Panama, Anthropic had downloaded more than seven million pirated books, from well-known clandestine libraries: Books3, Library Genesis (LibGen) and Pirate Library Mirror (PiLiMi). The company had built a centralised internal library from them.

There, no acrobatics about format could hold: the works had never been bought.

What followed is a record. Anthropic agrees to settle for $1.5 billion. Judge Alsup grants preliminary approval in September 2025, then retires. It is Judge Araceli Martinez-Olguin who delivers final approval on 20 July 2026.

The settlement figures deserve to be set out:

Item Value
Total amount $1.5 billion
Works on the list 482,460
Claim rate around 93%
Compensation per work around $3,000
Legal fees cut by the judge to around $101.6 million
Authors who opted out 350

According to the plaintiffs’ lawyers, it is the largest known copyright recovery in the United States.

Saying that “destroying books to train an AI is allowed” is therefore accurate — and seriously incomplete. The right formulation is: buying then destroying is allowed, downloading illegally costs $1.5 billion. These are two parts of one file, and it is the second that has set the precedent in practice.

A point almost nobody makes: this settlement creates no precedent

It has to be said clearly, because it is counter-intuitive.

A settlement ends a dispute without the court deciding the substance. The $1.5 billion settlement therefore establishes no binding precedent for the other cases under way. It creates a financial benchmark — from now on, everyone knows what a pirate library can cost — but it fixes no rule of law.

Only the June 2025 order on fair use keeps value as analysis, and it comes from a court of first instance. Which is to say that the central question — is training a model on copyrighted works lawful? — remains open in the United States.

No, these were not rare books — but it is not that simple

This is the most useful correction to make to the story going around, and it cuts both ways.

Alarmist headlines speak of the destruction of rare editions and antique works. The available documents do not support that reading. The suppliers primarily targeted non-fiction published from the 1970s onwards and carrying an ISBN — in other words ordinary second-hand copies, often unsold, which the market is awash with and which frequently end up pulped anyway.

On that precise point, the charge of cultural vandalism does not stand. A management textbook from 1994 bought in a job lot and scanned is not a loss to heritage.

But the opposite nuance exists too, and it is documented: specialist booksellers interviewed by the American press report that some bulk lots nonetheless contained older editions or copies that have become impossible to find. When you buy pallets of books by weight, you do not choose what is in them.

The honest formulation is therefore this one: it was not a targeted destruction of heritage, but a mass industrial operation whose collateral damage to irreplaceable copies is real and unquantified. Nobody, at Anthropic or at its suppliers, appears to have kept a record of what went under the blade.

In Europe, Judge Alsup’s reasoning would not hold

This is the question every European reader asks, and the answer is clear: no, this arrangement would not transpose as it stands.

European law has no American fair use, which is an open exception assessed case by case. It works through an exhaustive list of exceptions. On training, the reference is directive 2019/790 on copyright in the digital single market, which created two text and data mining exceptions, transposed in France into article L122-5-3 of the intellectual property code.

These exceptions do not work as they do in the United States:

  • The first benefits research organisations and bodies without a direct commercial purpose.
  • The second, broader one, comes with an opt-out for rights holders which must be exercisable by machine-readable means: robots.txt file, HTML tags, metadata.
  • For commercial exploitation by a large player, prior authorisation from rights holders remains the rule.

Two processes are under way, and they need careful distinguishing.

Before the Court of Justice of the European Union, a case puts head-on the question that has never been decided in Europe: does training a language model constitute reproduction within the meaning of copyright, and can that reproduction fall under the text and data mining exception? The hearing took place on 10 March, and the Advocate General’s opinion is expected on 3 September 2026, with the judgment to follow. In other words: the European answer does not yet exist. Any claim to the contrary is premature, and that is precisely what makes the current period unstable for rights holders and model vendors alike.

In France, Parliament has not waited. The bill brought by Senator Laure Darcos, tabled on 12 December 2025 with Agnes Evren, Pierre Ouzoulias, Laurent Lafon, Catherine Morin-Desailly and Karine Daniel, creates a presumption that artificial intelligence providers exploit cultural content. The Conseil d’Etat gave a favourable opinion on 19 March 2026: the national legislator is competent to create such a presumption, and the text runs contrary neither to the Constitution nor to European law, subject to drafting adjustments. The text was then examined in committee on 1 April, then passed unanimously by the Senate on 8 April 2026.

What that text changes, if it completes its passage, is considerable: it reverses the burden of proof. Today, an author has to show that their work served to train a model — proof that is nearly impossible to produce without access to the datasets. Tomorrow, it is for the provider to show it did not use the work.

Let us add a piece that is often misquoted. The AI Act does require, in its article 53, the publication of a sufficiently detailed summary of the content used for training, following a template published by the European Commission on 24 July 2025. But the timetable carries a decisive nuance: the obligation is immediate for models placed on the market from 2 August 2025, while models already in circulation have until 2 August 2027 to comply. Transparency about the training data of existing models is therefore not for right now — we set out that whole timetable in our analysis of what changes with the European AI regulation.

In concrete terms: a Panama-style operation run on the European market would expose its author to mass litigation, not to judicial validation.

The other route exists, and it has a price: pay

One element missing from the debate deserves stating, because it makes the case much less fatal than it looks: buying licences is possible, several companies do it, and the prices are known.

  • OpenAI has signed around two dozen agreements with publishers and data platforms. The largest known one covers up to $250 million over five years with News Corp, signed in May 2024.
  • Meta, long reluctant about that model, committed from March 2026 to up to $50 million a year for three years, $150 million in total.
  • Google struck its first AI content licensing agreement with the Associated Press.
  • Reddit alone disclosed $203 million of data licensing contracts at its stock market listing.

Set those amounts against the settlement. Anthropic pays $1.5 billion for seven million pirated books. The sector’s largest known licensing deal is worth $250 million over five years. In other words: the penalty is around six times the price of the biggest licensing contract ever signed.

That calculation changes how the case reads. Piracy was not only illegal: it turned out, in hindsight, to be more expensive than the lawful route. That is an argument that did not exist two years ago, when no decision had put a number on the risk. It exists now, and every general counsel in the sector has noted it.

One asymmetry remains that these agreements do not erase: they are negotiated with press groups and platforms, that is to say entities able to field lawyers. An isolated author negotiates nothing — they wait for a class action to succeed and collect $3,000.

What it changes for a European author or publisher

Three practical consequences, without extrapolation.

The opt-out has to be exercised, and technically. The opt-out provided for by the European directive is not a declaratory principle: it has to be expressed by machine-readable means. In practice that means a site’s robots.txt file, tags in the page code, or metadata in the files. A sentence in a site’s terms and conditions is not enough. A catalogue that is not technically protected is an exposed catalogue.

Collective action is, in practice, the only economically viable route. The Bartz case proves it: 482,460 works grouped together, a record settlement, and around $3,000 per title. No author would have obtained that alone, and 350 of them refused the agreement precisely because they considered that amount insufficient. In Europe, professional organisations and collecting societies play that aggregating role.

Transparency becomes a procedural weapon. The obligation on general-purpose model providers to publish a sufficiently detailed summary of their training content is not administrative paperwork: it is what will make it possible, tomorrow, to establish that a work was indeed used. Without it, the burden of proof remains nearly impossible for a rights holder to carry — and it is that very finding that motivates the French bill on the presumption of exploitation.

For the writing professions, this battle sits on top of another one, which we have documented: the value of editorial work itself, at a time when studies on exposed occupations put translation and writing at the top of the list.

The proceedings under way, and what is coming

The Anthropic file is not isolated. It is the first to reach its conclusion, which makes it a benchmark.

In France and Europe, several fronts are open. The French publishing unions have brought an action against Meta over the use of copyrighted works in training its models. Leading publishers — among them Hachette, Cengage and Elsevier — have sued Google, accusing it of exploiting millions of copyrighted works to train Gemini.

One technical point feeds these disputes and is worth knowing: researchers have shown that a model can reproduce whole passages verbatim from copyrighted works when those works are heavily represented in the training data — a form of overfitting. The argument that a model “learns” without “copying” loses force when the original text can be extracted by a single well-built prompt. The boundary between learning and reproduction is in fact at the heart of another dispute we have followed, the American accusation of copying levelled at a Chinese model — proof that the question goes well beyond books.

This legal battle comes with a lobbying battle, of which we documented another episode in the United States, when OpenAI backed a text that could exempt it. The stake is the same everywhere: get the rules fixed before the courts fix them.

What to take away

Three things, and they do not all point the same way.

The fact is real and it is documented. Anthropic did buy, cut up and destroy between 500,000 and two million books, as part of an internal programme named project Panama, with the declared ambition of digitising every book in the world. This is not a rumour: it appears in court exhibits.

The legality is real, but partial and local. An American federal judge validated purchase followed by destruction in the name of format shifting — destruction being, paradoxically, what makes the operation lawful. But the same decision ruled out fair use for seven million pirated books, which cost the company $1.5 billion and around $3,000 per work. And that settlement, being a transaction, creates no binding precedent.

The viral story exaggerates on one point and understates on another. No, these were not mostly rare books: the target was the ordinary second-hand stock, non-fiction, post-1970, with an ISBN. But yes, irreplaceable copies may have been in bulk lots, with no inventory to tell us today which ones.

That leaves the question this affair really poses, and it is not about paper. Models need a raw material the web no longer supplies: quality writing, edited, often old. That material belongs to somebody. American law answered with an acrobatic move on format; European law answers with prior authorisation. Those two answers are incompatible, and the companies concerned operate on both sides of the Atlantic.

That is where the next act plays out — not in the noise of the guillotine cutters. And it is a battle to follow just as closely as the one over the models themselves, whose performance our AI ranking measures without ever being able to say, for want of sufficient transparency, what each one was trained on.

The figure to keep

$3,000. That is what an author receives when their work is among the 482,460 covered by the Anthropic settlement. For comparison, the company was valued in the tens of billions of dollars at the time of the agreement, and the total sum — $1.5 billion — represents a fraction of what the sector raises in a quarter. The price of a book, when the courts finally set it, is still the price of a month’s rent.