Introduction
Generative AI is now one of the major issues in both U.S. and European copyright litigation. The debate is no longer limited to whether copyrighted works may be used to train models; it now also concerns issues such as the provenance of datasets, the lawfulness of model outputs, licensing opportunities, copyright management information (CMI) removal, technological protection measures, and the legal effect of opt-outs. In practice, this means that rightsholders are testing whether existing copyright rules can be used to control both the inputs and outputs of generative AI systems. Meanwhile, AI developers are seeking to rely on fair use in the United States and text and data mining (TDM) exceptions in Europe.
United States: Training, outputs, and fair use
The training of AI models is a frequent subject of litigation. Rightsholders are increasingly looking into the nature of the input materials allegedly used to train models. They are bringing claims against AI developers, asserting that copyrighted content was unlawfully used to train LLMs, and that the resulting models therefore reproduce protected content in their outputs. Where the input material is difficult to establish, some plaintiffs seek to rely on allegations of output-based direct infringement by pointing out the alleged substantial similarity between at least some LLM outputs and the alleged source material on which they rely.
We are also observing an increase of AI-related disputes in the music industry, with similar issues being raised. Rightsholders are pursuing certain AI models, alleging that they are carrying out mass copying of copyrighted sound recordings for the purpose of AI music generation. While some disputes seem to be progressing toward trial, many other such disputes (often involving the same defendants) are being settled through licensing models. These negotiated arrangements are emerging as an increasingly popular mechanism to avoid complex and costly litigation.
The issue of fair use remains unresolved in many of the ongoing U.S. cases, which causes difficulties for plaintiffs and defendants alike. One aspect of fair use that is often considered by courts is whether the defendant is alleged to have circumvented technological protection measures to access content for training AI music generation models. If so, that defendant is unlikely to be able to rely on fair use defenses.
Recent cases illustrate these difficulties. In one decision that highlights the challenges of applying fair use principles to AI models, the court held that even though the training of the LLM itself was transformative enough and constituted fair use, the developer’s retention of allegedly pirated books in a central library did not. That case settled before the court issued a final decision.
In other industries, such as the visual arts, landmark litigation against image generation LLMs continues, with courts permitting claims to proceed to trial for allegations of induced infringement or under the Lanham Act. We are also seeing a rise in character-based output claims, whereby rightsholders (such as film studios) are targeting AI-generated images or videos of protected characters.
Europe and the United Kingdom: TDM, outputs, and model weights
In Europe, the central battleground is the relationship between AI training and the DSM Directive’s TDM exceptions. Like Company v. Google Ireland is the first CJEU case on generative AI and copyright. This case revolves around whether the reproduction of press content in chatbot outputs and Google’s alleged use of that content for LLM training infringe EU copyright, or whether such use is covered by the TDM exception under Article 4 of the DSM Directive. This case is still pending, and any developments are highly anticipated.
It is also becoming clear from a series of recent decisions that the implementation of machine-readable opt-outs is becoming a decisive factor in European domestic AI litigation. German courts accepted that TDM exceptions applied to dataset creation and AI training where no valid machine-readable opt-out was declared. Danish and Dutch courts similarly found that human-readable reservations or selective bot exclusions may be insufficient unless the opt-out is machine-readable and actionable.
Finally, although the case is still ongoing and likely to progress to the Court of Appeal of England and Wales, the High Court handed down its judgment in Getty Images v. Stability AI in 2025. During the trial, Getty abandoned its primary claim (which relates to the training and development of the LLM) because the relevant training and development did not occur in the United Kingdom. As a result, the judgment does not consider whether training on Getty images was itself infringing, and this issue remains unresolved. Getty’s secondary infringement claim, that an “infringing copy” of its images was being made, was rejected. The court found that the models did not contain or store reproductions of the relevant works on which they were trained and were not therefore “infringing copies”.
Emerging trends
AI copyright disputes continue to raise fundamental questions about the use of copyrighted works for AI training. However, despite the growing number of claims brought in the United States, relatively few cases have resulted in substantive judgments on these issues. Many disputes have instead been resolved through settlements or commercial licensing agreements before courts issued a ruling on the merits, meaning that there is still limited judicial guidance on several key issues. Recent cases also suggest that courts in both the United States and Europe are reluctant to adopt bright-line rules, instead favoring highly fact-specific assessments. Nevertheless, the prevalence of settlements and licensing arrangements, and the current absence of general legal rules, do not mean that litigation is declining. The number of pending proceedings remains significant. Further judicial guidance is therefore expected as these cases progress through the courts.