The TDM exception in the age of generative AI: opt-out, training and the transformation of the exclusive right of reproduction

Reading time: 14 minutes

Abstract

The training of generative artificial intelligence systems is based on the use of enormous quantities of digital content, often protected by copyright. In Europe — and beyond — this phenomenon has reopened the debate on the limits within which works, articles, images or databases may be used to train models without the consent of rightsholders.

This article analyses the relationship between the text and data mining exceptions provided for by the DSM Directive and the rules governing general-purpose artificial intelligence models introduced by the AI Act, asking whether the training of foundation models on copyright-protected works is lawful. Starting from the traditional distinction between analytical use and substitutive use of works, the article reconstructs the genealogy of functional copies in European copyright law to show how generative AI profoundly alters this balance. Finally, the article examines the tensions that this transformation raises in relation to the international three-step test, the proprietary structure of European copyright and the protection of intellectual property as a fundamental right.

Introduction: From text and data mining to the training of foundation models

Generative artificial intelligence has forced European copyright law to confront a question that only appears to concern technology: to what extent may a protected work be used to train an AI model without the consent of the rightsholder?

In recent years, European law seems to have begun providing an implicit answer to this question through the progressive connection between the text and data mining (TDM) exceptions introduced by Directive (EU) 2019/790 (“DSM Directive”) and the rules governing general-purpose AI models laid down by Regulation (EU) 2024/1689 (“AI Act”).

TDM refers to the set of automated techniques that enable software and computer systems to analyse large quantities of digital content in order to extract information, correlations, recurring patterns and models. The “DSM Directive”, adopted to adapt European copyright law to the digital economy, introduced two specific exceptions for these activities:

– one intended for scientific research carried out by research organisations and cultural heritage institutions (Art. 3 DSM);

– and one applicable also to commercial uses, subject to the rightsholder’s opt-out (Art. 4 DSM).

The AI Act operates on a different plane. The Regulation was not adopted to regulate copyright, but to govern the development, placing on the market and use of artificial intelligence systems, with particular attention to general-purpose AI models (GPAI). However, in regulating such models, the AI Act expressly brings their training within the framework of the text and data mining exceptions provided for by the DSM Directive, consolidating the idea that the training of foundation models may take place on the basis of Art. 4 DSM, unless the rightsholder has reserved their rights.

It is precisely in this connection that the problematic issue emerges. The TDM exceptions introduced by the DSM Directive were conceived to allow automated data analysis activities. The AI Act, however, ends up projecting them into the context of generative AI, extending rules designed for analytical uses to systems whose purpose is no longer merely to extract information from works, but to generate new content from them.

The thesis of this contribution is that the TDM exception, interpreted in light of the AI Act and applied to the training of foundation models, risks operating no longer as a simple exception to copyright, but as a form of “implicit compulsory licence” in favour of the artificial intelligence industry. This transformation raises profound tensions with the proprietary structure of European copyright (see Art. 2 of Directive 2001/29/EC), with the international three-step test (see Art. 9(2) of the Berne Convention for the Protection of Literary and Artistic Works; Art. 13 TRIPS; Art. 10 WCT) and with the protection of intellectual property as a fundamental right (see Art. 17(2) of the Charter of Fundamental Rights of the EU).

The genealogy of analytical use in copyright law: Interoperability, search engines, Google Books and functional copies

The idea that protected works may be used without authorisation for purely analytical purposes did not originate with text and data mining. For some time, European and comparative copyright law has tolerated certain forms of “functional” reproduction, where copying the work is not the purpose of the use but the necessary means to obtain information, ensure interoperability or enable automated analysis processes.

This logic first emerged in the rules on reverse engineering and software interoperability. Directive 91/250/EEC, later incorporated into Directive 2009/24/EC on the legal protection of computer programs, allowed, in specific cases, decompilation activities necessary to obtain the information indispensable for enabling different computer systems to communicate. In this context, reproducing the software was not regarded as an economic exploitation of the work as such, but as a necessary technical step to gain access to functional information.

A similar logic gradually became established in the digital economy with search engines and automated web-indexing systems (see Art. 5(1) of Directive 2001/29/EC; CJEU, C-360/13, Public Relations Consultants Association Ltd v Newspaper Licensing Agency Ltd). Crawling, caching and indexing inevitably involve making temporary or functional copies of the content analysed. However, such copies have generally been regarded as compatible with copyright insofar as they served to locate, organise or make online information discoverable, without directly replacing the enjoyment of the original works.

The Google Books litigation also reflects this approach (see Authors Guild, Inc. v. Google Inc., 804 F.3d 202, 2d Cir. 2015). The mass digitisation of millions of published works was held permissible, in the US context, because it was primarily intended for text search and content indexing, while the display of the works remained limited and, as a result, did not substitute for the original publishing market.

Text and data mining governed by the “DSM Directive” fits into this framework. The exceptions provided for in Arts. 3 and 4 are based on the same underlying idea: allowing automated analysis of works in order to extract data, correlations and patterns, without giving central importance to the expressive enjoyment of the copied content.

The common element of these experiences is the distinction between analytical use and substitutive use. Copying is tolerated when it serves to understand, classify or analyse the work, not when it enables economic competition with it. It is on this distinction that the main critical issues arising today from the application of Art. 4 of the DSM Directive to generative AI are built.

The AI Act and the transformation of European copyright: From authorisation to compliance: opt-out, GPAI and implicit compulsory licence

The decisive shift occurs with the AI Act. Formally, the Regulation does not amend the DSM Directive or introduce new copyright exceptions. However, in regulating general-purpose AI models, it consolidates an interpretation of Art. 4 DSM as applicable to the training of foundation models.

Recital 105 is the most significant connecting point. The European legislature expressly recognises that training generative AI models requires access to large quantities of data and that, in this context, text and data mining techniques are widely used. The recital further adds that, “where the rights holder has expressly reserved the right to opt out in an appropriate manner, providers of general-purpose AI models need to obtain an authorisation from rights holders if they wish to carry out text and data mining over such works.”

The a contrario reasoning is clear. If authorisation becomes necessary only where a valid opt-out has been exercised, the system implicitly ends up assuming that, in the absence of such a reservation, training may take place within the exception provided for by Art. 4 DSM.

The AI Act does not merely invoke this connection at a theoretical level. Art. 53 imposes specific copyright-compliance obligations on GPAI providers, including the adoption of policies suitable for identifying and complying with reservations of rights expressed by rightsholders pursuant to Art. 4 DSM (see Art. 53(1)(c) AI Act), as well as transparency obligations concerning the content used for training (Art. 53(1)(d) AI Act).

A structural shift in the system therefore occurs.

– In the classic copyright model, use of a work presupposes the rightsholder’s prior authorisation.

– In the model emerging from the interaction between Art. 4 DSM and the AI Act, by contrast, use tends to become presumptively lawful unless the rightsholder opts out (see Recital 105 AI Act; Art. 53 AI Act; Art. 4 DSM), while the main issue moves from substantive authorisation to procedural compliance.

It is in this shift that the TDM exception risks taking on a function different from that originally envisaged by the European legislature. Rather than operating as a narrowly circumscribed limitation on the exclusive right, Art. 4 DSM progressively tends to function as a form of “implicit compulsory licence” in favour of the generative-AI industry: a system in which the use of works appears to be permitted by default unless the rightsholder objects in a manner that can be technically recognised by model providers.

Generative AI and the collapse of the distinction between analytical and substitutive use: From data mining to the industrial extraction of creative value

The rise of generative AI has profoundly altered the balance on which traditional forms of analytical use of protected works were based. Foundation models are trained on datasets of enormous size, often obtained through scraping and text and data mining techniques applied to content found online. From a technical standpoint, these activities bear many similarities to practices already tolerated in the contexts of reverse engineering, search engines or automated data analysis. However, the economic function of the use changes radically.

In traditional forms of analytical use, copying a work was a means of obtaining information about the works themselves: ensuring interoperability, organising content, identifying correlations or enabling automated searches. The analytical activity remained external to the expressive market of the work and did not produce content intended to substitute for its economic exploitation.

Generative AI breaks this distinction. The training of foundation models uses works not only to extract data or informational patterns, but to build systems capable of generating text, images, code, music or other content that may potentially compete with the content used in the training dataset.

A chatbot trained on journalistic articles can provide summaries and answers that reduce the need to consult the original sources. An image generator trained on photographs, illustrations or works of art can produce content that serves as an alternative to that offered by photographers, illustrators or stock platforms. A model trained on code repositories can generate portions of software that are economically substitutive for the work of human developers.

It is at this point that the systemic rupture occurs. Analysis is no longer confined to knowledge of the work, but becomes productive capacity. The functional copy no longer serves only to understand or classify content, but fuels systems that operate directly in creative, information and professional markets.

It is precisely this shift that makes the application of Art. 4 DSM to the training of foundation models problematic. The TDM exceptions were conceived in a context in which automated use of works appeared substantially non-substitutive. Generative AI, by contrast, introduces a form of exploitation that tends to move ever closer to the economic core of the exclusive right.

The systemic conflict with international law and fundamental rights: Three-step test, intellectual property and transformation of the exclusive right

The transformation produced by the interaction between Art. 4 DSM and the AI Act does not raise only issues of legislative policy or economic balance between the AI industry and rightsholders. It also raises deeper questions as to whether the current European framework is compatible with the limits imposed by international copyright law and with the protection of intellectual property as a fundamental right.

The first area of tension concerns the three-step test, enshrined in Art. 9(2) of the Berne Convention, Art. 13 of the TRIPS Agreement and Art. 10 of the WIPO Copyright Treaty. According to this criterion, exceptions and limitations to copyright are permissible only:

– in certain special cases;

– insofar as they do not conflict with the normal exploitation of the work;

– and do not unreasonably prejudice the legitimate interests of the rightsholder.

Applying Art. 4 DSM to the training of foundation models raises doubts in relation to each of these requirements. The broad scope of the exception, potentially applicable to enormous quantities of content and to models intended to operate on an industrial scale, makes it difficult to characterise generative training as a “special case” within the meaning required by the three-step test.

The second requirement also appears problematic. European case law has traditionally interpreted copyright exceptions restrictively precisely because of their impact on the rightsholder’s exclusive right (see CJEU, Infopaq International, C-5/08; Pelham, C-476/17). Traditional forms of analytical use had been considered compatible with the three-step test insofar as they did not directly interfere with the market for the original works. Generative AI, by contrast, increasingly tends to operate in the same creative and information markets from which it extracts the content used for training.

The issue is particularly evident in sectors in which licensing markets specifically intended for AI training are emerging. The growing spread of agreements between generative-AI platforms and publishers, photographic agencies, collecting societies and holders of large digital archives demonstrates that the use of works for training is acquiring an autonomous economic value. In this context, it is increasingly difficult to maintain that such use does not interfere with the “normal exploitation of the work” within the meaning of the three-step test.

There is also a constitutional dimension. Art. 17(2) of the Charter of Fundamental Rights of the European Union expressly provides that “intellectual property shall be protected”. The Court of Justice of the European Union has repeatedly recognised that copyright constitutes a form of property protected by EU law, although it must be balanced against other fundamental rights and general interests (see CJEU, Promusicae, C-275/06; Scarlet Extended, C-70/10; UPC Telekabel Wien, C-314/12).

The European Court of Human Rights has also progressively brought intellectual property within the scope of Art. 1 of Protocol No. 1 to the ECHR, recognising it as a “possession” capable of Convention protection (see ECtHR, Anheuser-Busch Inc. v. Portugal, GC, 2007).

From this perspective, the issue concerns not only the abstract legitimacy of the TDM exceptions, but their functional transformation in the context of generative AI. The more Art. 4 DSM tends to operate as a general mechanism for access to content for industrial training purposes, the more the system risks approaching a form of regulatory reallocation of the economic value of works, without a corresponding compensation mechanism for rightsholders.

It is in this shift that the central thesis of this contribution emerges. Interpreted in light of the AI Act and applied to the training of foundation models, the TDM exception in fact risks operating no longer as a simple, narrowly circumscribed limitation on the exclusive right, but as a form of “implicit compulsory licence” in favour of the artificial intelligence industry. The use of works tends to be regarded as generally lawful unless the rightsholder objects by opting out, while effective control over the economic exploitation of content progressively shifts from prior consent to the technical management of the reservation of rights.

This transformation is likely to fuel growing litigation. The actions already brought in the United States and the United Kingdom against OpenAI, Stability AI, Meta or Anthropic (NYT v. OpenAI; Getty v. Stability AI; Kadrey v. Meta; Bartz v. Anthropic) show that the conflict no longer concerns only the technical lawfulness of training, but the redistribution of the economic value produced by generative-AI systems. In the European legal order too, litigation can be expected to shift progressively from the issue of mere technical reproduction of works to scrutiny of whether the current model is systemically compatible with the proprietary structure of European copyright, the limits imposed by the international three-step test and the protection of intellectual property as a fundamental right.

From this perspective, the current regulatory framework may prove to be only a transitional phase in a broader process of transformation of European copyright. The more the use of works for AI training tends to be regarded as generally lawful unless the rightsholder opts out, the more the exclusive right risks moving, functionally, towards a model based no longer on the rightsholder’s prior consent, but on forms of generalised access potentially accompanied by compensation mechanisms.

In this scenario, the real issue no longer seems to be whether AI training should be permitted, but under what economic and legal conditions such use may take place without structurally altering the proprietary function of European copyright.

If AI training is permitted unless the rightsholder opts out, the exclusive right does not formally disappear. But its nature changes. From a right to authorise, it tends to become a right to object. From a rule based on prior consent, it becomes a reservation mechanism. From fully negotiable property, it risks turning into a legal position conditioned on a technical burden of protection.

It is in this transformation that one of the most significant battles in contemporary copyright law is being fought. The TDM exception, born to enable automated data analysis, risks becoming the Trojan horse of a new industrial infrastructure for access to content: a form of “implicit compulsory licence”, not expressly declared by the legislature and constructed not through an organic reform of copyright, but through the regulation of artificial intelligence.

Reviewed by: Arlo Canella
Publication date: 27 May 2026
© Canella Camaiora S.t.A. S.r.l. - All rights reserved.

Textual reproduction of the article is permitted, even for commercial purposes, within the limit of 15% of its entirety, provided that the source is clearly indicated. In the case of online reproduction, a link to the original article must be included. Unauthorised reproduction or paraphrasing without indication of source will be prosecuted.

Celeste Martinez Di Leo

Praticante avvocato, laureata in Giurisprudenza presso l’Università degli Studi di Pavia e in “Abogacía” presso l’Universidad de Belgrano (Argentina) a pieni voti.

Leggi la bio
error: Content is protected !!