Loading...

AI Training on Copyrighted Works: What Courts Are Actually Ruling

Key takeaways

  • Judge Alsup ruled that AI training on copyrighted text is lawful, but Anthropic paid $1.5 billion for obtaining books through pirated shadow libraries.
  • Courts apply fair use based on competitive intent: Thomson Reuters' lawsuit established that building competing products from copyrighted training data doesn't qualify as transformative use.
  • Copyright law hasn't been updated since 1976, leaving judges to interpret half-century-old guidelines for technology that didn't exist when the law was written.
  • Thaler v. Perlmutter ruled that works generated entirely by AI receive no copyright protection, raising questions about the threshold for AI assistance.

AI models powering ChatGPT, Gemini, Claude, and comparable systems learned from datasets containing hundreds of millions of books, articles, and academic papers. Most authors whose works appear in these training datasets never consented to their inclusion. The fundamental legal question — whether this constitutes copyright infringement — remains deeply unsettled, with courts and legal experts divided on how to apply 50-year-old copyright law to technology that did not exist when the law was written.

The Anthropic Judgment

In one of the first major rulings on AI copyright questions, Judge William Alsup ordered Anthropic to pay $1.5 billion to a group of writers whose works were used in training the company’s AI models. The decision appears, on its surface, to be a decisive victory for authors concerned about unauthorized use of their intellectual property.

Judge Alsup’s actual ruling tells a more complicated story. He determined that Anthropic’s core practice — training AI systems on copyrighted text — was lawful. The $1.5 billion penalty addressed a distinct violation: Anthropic had obtained many books by pirating them from illegal online shadow libraries rather than acquiring legitimate copies through proper channels.

The judge compared AI training to how writers learn by studying existing literature. “Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different,” Alsup wrote. This framing treated the absorption of patterns across vast amounts of text as analogous to human reading and study, not as copyright violation.

The settlement’s significance diminishes when placed in financial context. Anthropic projects approximately $200 billion in annual revenue by 2028. Cathy Gellis, an attorney specializing in intellectual property and technology law, highlighted this disparity: “What’s a $1.5 billion fine to a company projecting about $200 billion in annual revenue by 2028?” For a company operating at that scale, the penalty may represent a manageable cost of doing business rather than a deterrent.

Fair Use and Copyright Training

Copyright disputes involving AI systems almost universally center on fair use doctrine, a principle that allows limited use of copyrighted material without explicit permission when the use is sufficiently transformative.

The Fair Use Doctrine

Courts weigh multiple factors when assessing fair use claims: the purpose and character of the work in question, how much material was used, and the effect on the market for the original work. Gellis emphasized a foundational distinction in copyright law. “Copyright law hinges on copying, but it doesn’t hinge on using the work or experiencing the work, consuming the work, reading the work,” she explained. By this interpretation, the act of AI systems ingesting text and learning patterns from it might constitute reading rather than copying, a distinction that reshapes the legal analysis.

Competitive Intent as a Deciding Factor

Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International, identified a consistent pattern in how courts have approached these cases. The deciding factor often turns on whether the AI system aims to build a directly competing product. “If what you’re doing is you’re training on somebody’s property because your purpose is to directly compete, then the courts will frown on it,” Henderson said. “If what you’re doing is not going to compete, then the courts are tending to find ways that it will be okay.”

A woman immerses in virtual reality with neon-lit goggles and gloves in a tech-savvy environment.

Thomson Reuters and Competitive Reuse

The Thomson Reuters case exemplifies how courts apply the competitive intent test. The technology and media company sued Ross Intelligence, a research firm that had trained an AI system on Thomson Reuters’ content to build a competing legal research platform. Judge Stephanos Bibas ruled against Ross, finding no fair use protection. “Ross’s use is not transformative because it does not have a ‘further purpose or different character’ than Thomson Reuters’s,” Bibas wrote.

Directly Competing Products

The decision rested on a straightforward principle: using copyrighted material to build a product that directly replaces the original copyright holder’s own offering does not qualify as transformative. Authors have made similar arguments about AI chatbots, contending that systems trained on their work compete by generating synthetic text that could substitute for human-written content. These arguments have not yet prevailed in any court ruling, though multiple pending cases may eventually shift this calculus.

Copyright law has not been updated since 1976, meaning judges must interpret guidelines written for photocopiers and VCRs when deciding cases involving systems that learn from billions of documents. Henderson noted that legal uncertainty pervades the field because most cases remain in active litigation. “They know that the AI model has been trained on so much stuff, and the law has not really caught up to that question,” he said.

AI-Generated Works and Copyright Protection

A separate legal question concerns copyright protection for works created entirely by AI systems. In Thaler v. Perlmutter, a court ruled that works generated entirely by AI with no human creative input cannot receive copyright protection. This raises thorny definitional problems that copyright law has never before had to confront: at what point does AI assistance disqualify a work from copyright protection, and how can anyone prove the extent of AI involvement in creation?

Gellis drew a comparison to existing technology that proved illuminating. “If you write your novel in Word and run spell check, we kind of feel comfortable with the idea of saying that Word does not own your novel,” she noted. AI systems force reconsideration of questions that were previously settled or ignored. The threshold question — whether AI-assisted or AI-generated works merit copyright protection at all — remains completely open.

Ongoing Legal Uncertainty

Most major AI companies currently face pending litigation over copyright questions, meaning final legal resolution remains years away. Court decisions carry weight as precedent even before reaching final judgment, influencing how AI companies, publishers, and authors make business decisions today.

Gellis cautioned that current uncertainty will likely persist. “What you are seeing is that the initial opening volleys are being influential, and that influence itself could be undone if other courts decide different things, and it’ll take later states of litigation to figure out which one will prevail,” she said. “It would be kind of foolish for the AI companies to ignore them.” Companies must navigate an environment where yesterday’s favorable ruling could be overturned by tomorrow’s appeal, and where legislative reform remains uncertain.

Frequently Asked Questions

Did the $1.5 billion settlement prove that AI training on copyrighted books is illegal?

No. Judge Alsup determined that training AI on copyrighted text is lawful. The $1.5 billion penalty was specifically for Anthropic pirating books from illegal shadow libraries rather than acquiring legitimate copies.

How do courts decide if AI training on copyrighted material violates copyright?

Courts apply fair use doctrine, weighing whether the AI system creates a competing product. Thomson Reuters' case against Ross Intelligence established that training on copyrighted content to build a directly competing product is not fair use.

Can a work that's entirely AI-generated receive copyright protection?

No. In Thaler v. Perlmutter, the court ruled that works generated entirely by AI with no human creative input cannot be copyrighted.

Written by
Marcus Feldman

Marcus Feldman analyzes cryptocurrency and blockchain markets — price movements, protocol upgrades, and the regulatory shifts reshaping crypto exchanges worldwide.