Authors
Introduction
The predominant way that rights to collect, use, and share data are allocated in advance, or to create business certainty, is typically through licensing. A license is a right or a permission for a person or company to use another party’s intellectual property, often in exchange for a fee. The benefit of the licensing model is that it offers tremendous flexibility to slice, dice, allocate, monetize, expand, and limit collection, use, and disclosure in an area where traditional intellectual property rights – such as patent, copyright, trademark, and trade secret law – may be less clear, or where there may be differing opinions or points of view. Licensing can help address these issues among and between businesses and even consumers. In particular, licensing as a tool has broadly enabled many of the data-focused innovations of the Internet age. Licensing also helps address privacy and data protection issues in many legal systems; for example, in the United States, not only do privacy policies often address these issues, but terms of use or terms of service frequently include license grants that extend to things that may or may not be subject to traditional intellectual property grants.
Since this guide was first published, new regulatory frameworks and guidance have reinforced the centrality of licensing while adding compliance obligations. In the European Union, the AI Act requires providers of general-purpose AI models to implement copyright compliance policies and to disclose summaries of their training data. In the United States, the Copyright Office has described voluntary licensing as a workable market-based mechanism for allocating rights in copyrighted training data.
Can data be owned?
Data is free-flowing information, and there is no standard definition of the term “data.” The Joint Technical Committee of the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC) proposes the following definition:
“Reinterpretable representation of information in a formalized manner suitable for communication, interpretation, or processing.”
At the most basic level, data is just information. For example, the fact that a texture belongs to a genre called “brick” or “steel” is information that is not capable of appropriation in itself. This lack of ownership of data stems from the fundamental principle that information, ideas, methods, and techniques are free and flow freely.
Data is information. Whether any person has a proprietary or “ownership” interest in data rests on the question of whether the law has created a specific property regime for that type of data, also called intellectual property. In most countries, there exist only four types of data eligible for intellectual property protection: (1) works of art and other subject matters from the creative industries (i.e., copyrighted works); (2) trade secrets; (3) trademarks; and (4) patentable inventions. Simply put, data that does not fall within one of the aforementioned categories may not be “owned” as intellectual property. And those categories do not provide complete ownership of information; rather, they provide restrictions on what others can do. For example, others can copy the idea of a creative work if they do not copy the expression, independently create data that may be protected by another as a trade secret, and copy and build on data disclosed in a patent application as long as they do not infringe the patent. Of course, this does not mean that there is complete freedom to use and reuse data that is not intellectual property, since other types of restrictions might apply to the data, as discussed below.
Ownership of property regimes applicable to data
Ownership of or right to control input data
Parties providing data intentionally for use in machine learning or AI development frequently seek to assert ownership or control of the data or otherwise assert a right to exclusively share and use it. However, ownership, in the sense of property, is often not available, with personal information being a good example. Personal information is not a proprietary right; it is an access right that is almost exclusively controlled by the individual to whom the information relates or identifies, and it is difficult for a party other than that individual to assert rights over that specific data. For other types of data, such as confidential business information, a customer may want to ensure that it maintains explicit confidentiality rights in its data and that no rights are transferred to the vendor by virtue of the performance of services for the customer. For a vendor, there may be significant value in controlling the input data so that it may continue to use such input data in its AI tool without breaching another party’s rights. Many vendors provide services freely or at low cost in order to generate input data that can be used to train and improve their models.
Since this guide was first published, the legal landscape for input data has been shifting. In Thomson Reuters v. Ross Intelligence, the first U.S. court to rule on fair use in AI training held that using copyrighted headnotes to train a legal research AI was not fair use where the purpose was to create a competing product. Though confined to non-generative AI, the ruling underscored the risk of using copyrighted material without a license. Two subsequent decisions in the Northern District of California, Kadrey v. Meta Platforms and Bartz v. Anthropic, found that using copyrighted books to train large language models was fair use because it was “highly transformative,” though the decisions were fact- and theory-specific and could differ where market harm is demonstrated. As the law continues to develop, parties negotiating input data provisions should account for risk by addressing training data provenance, audit rights, indemnification for intellectual property claims, and compliance with applicable obligations.
Ownership or control of output data
The ownership status of output data can be a highly contested and difficult provision to negotiate in data-related contracts. The output data of an AI model may include direct end-user output data created for use by the AI customer and indirect output data that is input by the customer and used by the model to improve functionality and efficiency. Output data varies depending on the type of model used and its purpose. Many customers desire to “own” the output data since it was created using input data provided by the customers. A customer could argue that output data was a derivative work (as that term is used under the U.S. Copyright Act) and, therefore, that copyright automatically flows through to the customer. Unfortunately, data used in AI development or model training may often have uncertain copyright provenance or be unequivocally outside copyright protection. Even trade secret status is frequently unclear. However, vendors will sometimes argue that it is important that they “own” the output data they created, since they used their own proprietary model to create it. The AI vendor may also want to ensure it retains at least partial “ownership” of the output data so that it can continue to use that data to train the AI model.
The legal landscape for output data has also been evolving. The U.S. Copyright Office maintains that works entirely generated by AI are not copyrightable, and that the mere selection of prompts does not yield a copyrightable work. Only where a human has contributed sufficient expressive elements may copyright attach to those contributions. Trade secrets have emerged as an increasingly important alternative. Unlike copyright, trade secret law does not require human authorship, making it a viable mechanism for protecting commercially valuable AI outputs, provided that the party asserting trade secret status implements reasonable measures to maintain secrecy. Parties should carefully review contractual language to ensure that each of their interests is protected, with particular attention to intellectual property ownership, trade secret protections, confidentiality obligations, and representations regarding the copyrightability of output data.
Use of derived data
Derived data is new data and insights derived from output data and may not have been available from the existing data. Since such derived data is valuable, both customers and vendors may have potential uses for such data outside the contractual agreement for the AI model. One of the issues with derived data is who can control it. A customer could argue that it should be afforded the right to control the derived data since it was the original inputter of the data used to create it, but since derived data is created by combining and transforming data into a new type of data, a vendor could argue against this. The contractual framework for derived data must now also account for new regulatory requirements. At the same time, the growing licensing market for AI training data offers new monetization opportunities for both customers and vendors seeking to commercialize derived data.