“This case sets a new bar for AI training data,” said a legal expert tracking the settlement closely. Anthropic, the company behind the Claude AI model, has agreed to pay $1.5 billion following allegations it trained its AI using half a million pirated books. The ruling affects nearly 500,000 copyrighted titles, with US-registered rights holders expected to receive about $3,000 per book. It marks a significant moment in how legal frameworks are catching up with AI’s rapid growth.

The dispute originated when Anthropic was found to have sourced training data from piracy sites like Library Genesis and Pirate Library Mirror. The court made a sharp distinction: using legally obtained material might fall under fair use, but pirated content crosses clear infringement boundaries. As part of the settlement, Anthropic must destroy all pirated datasets involved in training Claude. Final court approval is anticipated by mid-2026, with payments scheduled to be released in phases after deductions.

The involvement of the Irish Writers’ Union highlights the case’s international implications. The union has negotiated its own deal with Anthropic and is advising European authors on how to file claims. Given Europe's stricter copyright and data laws, along with the rollout of the EU’s AI Act since 2024 requiring transparency on training data, this sets a precedent likely to influence AI companies operating across borders.

Anthropic isn’t alone in facing such legal battles. OpenAI faces lawsuits from authors and media giants, while Stability AI is entangled in disputes with Getty Images. Although $1.5 billion sounds huge, Anthropic’s valuation runs into tens of billions, suggesting that licensing training data will become inevitable. This shift could fundamentally alter the economics of developing large language models amidst mounting copyright enforcement pressure.