/ 5 min read / Entertainment & Media Guide to AI: Three years on

Challenges of data licensing in the U.S.: What is ‘fair use’?

Introduction

The use, ownership, and exploitation of data is extremely valuable. The era of AI has ushered in a veritable gold rush of companies and individuals seeking to mine this man-made resource, which, unlike gold, is available in great abundance. However, the alchemy involved in turning seemingly infinite data into something valuable requires tremendous computational power and investment. Since this guide was first published, the legal landscape for data licensing and fair use in AI has evolved. U.S. courts have issued the first fair use rulings on AI training data; the U.S. Copyright Office has released a report concluding that AI training can constitute copyright infringement; the EU AI Act has imposed obligations on AI providers; and a growing licensing market has emerged as a mechanism for securing rights in training data. 

Understanding text and data mining

Text and data mining (TDM) generally involves the identification of patterns or relationships in data sets that were previously unknown. TDM can be used to build predictive models of behavior in the retail context, so that when a customer navigates Amazon’s website or opens their Facebook page, they are presented with advertising keyed to their individual tastes and preferences.

In the media and entertainment context, one form of TDM – machine learning – is being used to train AI programs to create content, whether in text, audio, visual, or audiovisual form. Machine learning, like traditional TDM, is intended to discover novel and useful knowledge in data. However, a fundamental difference between machine learning and traditional TDM is that TDM can, in itself, extract data for human comprehension, whereas machine learning extracts data to improve an AI program’s ability to produce output. In addition, TDM does not necessarily involve rule or pattern discovery, while machine learning almost always does.

TDM in the U.S.: What is 'fair use' anyway?

The legality of making copies of the text or data through TDM has become a serious issue. As AI search engines crawl the World Wide Web, continually seeking, digesting, and aggregating content, they inevitably process copyrighted works such as music videos, songs, novels, and news stories. Because this processing – which generally requires making a copy – is frequently performed without the express consent of the copyright holder, its legality often depends on whether it is permitted under an exception to, or outside the framework of, copyright law. Under U.S. copyright law, the exception most frequently relied upon is fair use.

Under section 107 of the Copyright Act, fair use factors include: (1) the purpose and character of the use; (2) the nature of the copyrighted work; (3) the amount and substantiality of the portion used in relation to the whole; and (4) the effect of the use on the potential market for, or value of, the copyrighted work. Fair use of a copyrighted work for purposes such as teaching, scholarship, and research is specifically permitted by section 107. A key consideration that courts use in deciding whether fair use exists is whether the use is “transformative.”

Whether copying copyrighted material for machine learning constitutes fair use is a hotly debated topic that will affect the future of AI in the United States. For example, Thomson Reuters and West Publishing Corp. sued Ross Intelligence, Inc. over, among other things, its alleged use of machine learning to create a legal research platform based on the Westlaw database. The U.S. District Court for the District of Delaware granted partial summary judgment to Thomson Reuters, ruling that Ross Intelligence’s use of Westlaw headnotes to train its AI legal research tool was not fair use. This landmark decision found that Ross’s use was not transformative because it was intended to create a competing legal research product serving the same purpose as the original works, and that the market-effect factor favored Thomson Reuters because Ross was a direct competitor. Although the court confined its holding to non-generative AI, the ruling underscores that AI training on copyrighted material is not categorically fair use and that the provenance of training data carries significant litigation risk. The fair use issue is currently on interlocutory appeal before the U.S. Court of Appeals for the Third Circuit.

Will fair use protect machine learning?

In a seminal 2015 case, the Second Circuit found that Google Books’ scanning of more than 20 million books, many of which were subject to copyright, constituted a non-expressive and transformative fair use of the texts because it enabled users to search for information about copyrighted books, as opposed to the expressive content of the books themselves. A key takeaway from the case is the distinction between “expressive” and “non-expressive” use of copyrighted materials, the latter being deemed fair use by the court. Applied to AI, does this suggest that, so long as the original text does not “express” in the final work product, the act of machine reading is fair use?

Since this guide was first published, U.S. courts have issued the first fair use rulings addressing AI training. In June 2025, Northern District of California judges ruled in favor of AI developers in Kadrey v. Meta Platforms and Bartz v. Anthropic. The courts found that using copyrighted books to train large language models was “highly transformative” because the purpose of copying was to train an AI system, not to reproduce or distribute the original works. However, the court in Kadrey cautioned that the ruling was limited and that a market dilution theory under the fourth factor could have changed the outcome had the plaintiffs developed an evidentiary record on that point. These cases and Thomson Reuters v. Ross confirm that whether use of copyrighted works for AI training constitutes fair use is neither categorical nor foreclosed. Instead, the analysis will depend on the specific facts, including the purpose of the training, the nature of the copyrighted works, and the evidence of market impact.

In practice, the licensing market for AI training data has evolved rapidly since this guide was first published. Major content owners, including news publishers, academic publishers, music rights holders, and stock image companies, have entered into licensing agreements with AI developers, and the AI training data licensing market has grown. The U.S. Copyright Office endorsed voluntary licensing as the preferred market-based solution. These developments confirm that licensing, rather than unilateral reliance on fair use or TDM exceptions, is becoming the primary mechanism for securing rights in AI training data.

Related Insights