Squaring the Copyright and AI Circle

The Commission's AI and copyright dilemma, in 432 responses
Analysis
July 27, 2026

On 25 June, the European Commission’s Call for Evidence for a targeted legislative initiative on copyright closed with 432 responses. The Call for Evidence is the formal kick-off to the next round of EU copyright legislation, and its most consequential strand deals with AI and copyright. This analysis takes a closer look at the submissions in response to the consultation to understand where the various stakeholders in the debate about AI and copyright stand. The Call for Evidence bundled four distinct issues — AI and copyright, online piracy of live events, protection of sound recordings of third-country nationals, and research-related uses of protected content. The following analysis concerns itself only with the AI and copyright strand, with a brief look at the research exception where it illuminates the wider dynamics; the other two strands are left aside.

The analysis is based on machine-assisted coding of all 432 responses that have been published on the Commission’s website, together with the 339 attached position papers. The responses have been coded on a fixed set of positions: the stance taken on the two text and data mining (TDM) exceptions, opt-out mechanisms, transparency, remuneration, rights infrastructure, and the research exception.[1] This post summarizes the trends that have emerged from the full-corpus analysis.

Of the 432 responses, 359 take a position on at least one of the AI and copyright issues listed above.

Dual objectives 

The main objective of this analysis is to understand how the responses deal with the core policy tension at the center of the call for evidence. In the roadmap document, the Commission states that it seeks to “enhance the licensing and enforcement of copyright and related rights in the AI context, improve the conditions for creators’ remuneration, while making it easier for providers of generative AI to access copyright-protected content.” This dual objective — more control and remuneration for rightholders, and easier access to training data for AI developers — is the core policy dilemma that any legislative intervention in this space needs to address. The Commission is caught between the overall strategic objective of building an “AI continent” and its deep-rooted commitment to protect the economic position of Europe’s cultural and creative sectors. At this stage it is far from obvious that those objectives can be honoured at once.

Read against this ambition, the overwhelming majority of the responses are unsurprising: they map relatively cleanly onto the economic and ideological interests of the entities who filed them, and most responses ask the Commission to deliver one half of the objective at the expense of the other. Nevertheless, what makes the overall corpus of responses interesting is a small set of submissions — ours among them — that take the dual objective seriously and propose ways to meet both halves at once.

Who responded

Coded by the type of organization and the positions actually taken,[2] the field breaks down as follows: rightsholders account for 241 of the 432 responses, by far the largest bloc, split between large corporate rightsholders (93), collective management organizations (86), individual-creator organizations (46), and mixed coalitions (16).

The library, research and open knowledge sector comes in second with 65 responses. Cross-sector business federations submitted 34 responses. By contrast, the AI and technology industry submitted a mere 33 responses while individual citizens responded 23 times. Finally, the sports and anti-piracy coalition submitted 21 responses, almost all of which address only the live-piracy strand.

When compared to previous consultations on copyright policy issues, two absences stand out. The digital rights and consumer organizations that were central to the discussions over the DSM Directive are almost entirely missing: neither EDRi nor BEUC filed a response, and from the broader digital rights community, only Wikimedia chapters and Open Future responded. Euroconsumers is the lone consumer organization in the entire corpus.

Similarly, the responses skew towards professional interest representatives. Responses from individual citizens are marginal— and mostly unrelated to the topic of the call for evidence. Compared with previous copyright-related consultations, the absence of users’ concerns about regulatory overreach (save the internet!) and individual creators’ concerns about the erosion of protections they enjoy under copyright (save copyright!) is notable. While this is likely due to the somewhat obscure and technical nature of such calls for evidence, it is still somewhat surprising given the pervasive nature of anti-AI sentiment among many types of creators.

Whatever the reasons, the effect is that the user-and-public-interest perspective is carried almost exclusively by the library, research, and open knowledge sector.

TDM exceptions: harden, exclude, or keep?

Unsurprisingly, the central battleground that emerges from the responses is Article 4 of the DSM Directive, which allows text and data mining, including the use of copyrighted works for AI training, unless rightsholders have explicitly reserved their rights. The responses largely mirror the positions taken by various stakeholders since the connections between the TDM exceptions in the DSM Directive and the use of copyrighted works to train generative AI models became clear.

The rightsholder field overwhelmingly wants to limit or abolish this arrangement, but it is split on how. The largest group (105 submissions, 95 of them rightsholders) wants to keep the overall architecture but harden the opt-out by making rights reservations enforceable and backed by stronger transparency obligations. A second group, 72 submissions strong, goes further and wants generative AI training excluded from the scope of the TDM exception altogether, so that model training requires a licence by default. This exclusionist position is led by the audiovisual sector and collecting societies, and includes major organized players. On the other hand, the big international umbrella organizations  (CISAC, IFPI, MPA) notably stay within the existing architecture.

Of these two general positions, the demand to exclude model training from the scope of the TDM exceptions is the more consequential one. It would reverse the existing default for the use of publicly available content from “permitted unless reserved” to “prohibited unless licensed,” which would render large swaths of publicly available information effectively unusable for model training. This does not merely fail to address the Commission’s objective of “making it easier for providers of generative AI to access copyright-protected content”; it would also make developing competitive AI models in the EU  more difficult. This would be exaggerated by the fact that any changes to the scope of text and data mining would not only affect (commercial) uses under Article 4 of the CDSM Directive but also research relying on Article 3. Such an outcome could hardly be reconciled with the AI content ambitions that drive much of the Commission’s economic and industrial strategy these days.

On the other side, the common position of the technology industry is to preserve the status quo: of the 33 tech submissions, 19 want Article 4 kept as is and 9 want it clarified. Only a handful of submissions push to weaken or remove the rights reservation, with Meta taking the hardest-line position in arguing that the possibility for rightholders to opt-out should simply be abolished.

Though the library and research sector largely aligns with the tech industry in defending the TDM framework, its primary concern is access to knowledge rather than competitiveness, and its core objective is to maintain and strengthen the scientific-research TDM exception in Article 3 of the CDSM Directive. This alignment is coincidental rather than structural — both groups arrive at the position independently — which shows in the consultation’s research questions, where the libraries face a unified rightsholder bloc while the tech sector largely falls silent.

Positions on the Article 4 TDM exception by stakeholder group. 303 responses take a position on Article 4; the four groups shown account for 292 of them.

Beneath the positions on the exception itself lies a shared diagnosis: 151 submissions, from all camps, consider the current state of opt-out signalling to be broken. But agreement ends there; support for specific machine-readable standards is thin, and most standards are carried by their own, largely non-overlapping constituencies. The IETF AI-preferences work (13 endorsements) is backed almost entirely by the technology camp — Google, GitHub, CCIA, BSA, DOT Europe. On the other hand, TDMRep (14 endorsements) is supported by a mix of publishers and technology firms. Robots.txt collects the largest number of endorsements (34), yet it is the standard that rightsholders attack the most vigorously — CISAC calls it “an unacceptable format for opt-outs due to its inefficiency and opacity.”

Unsurprisingly for anyone who has followed this topic over the past years, nearly everyone mentions “machine-readable rights reservations,” but nobody agrees on how they can work in practice or, more consequentially, on who should govern them.

A similar dynamic plays out when it comes to the issue of opt-out registries that have been proposed as a solution to recording opt-outs in a more robust way. Where one might expect those who consider opt-out signalling broken to demand a central opt-out registry, the opposite is the case. Opposition is stronger among CMOs, with CISAC dismissing a centralized registry as adding “administrative and bureaucratic hurdles to an already imperfect procedure”.

In the end, the main divergence among those who agree that opt-out signalling is broken is between those who locate the fix in standardization and better infrastructure, and those who propose a different allocation of the burden of proof.

Transparency and remuneration

That different allocation of the burden of proof has a concrete shape, and it recurs across the major rightsholder submissions as their central operational demand. The first element is mandatory work-level disclosure: AI developers must reveal which specific works were used in training, so that rightsholders can see whether their reservations were honoured, substantiate claims, and identify whom to license to. The second element is the fallback for when disclosure fails: a rebuttable presumption of use — under which a developer that does not provide granular transparency, or refuses to cooperate — is presumed to have trained on the works in question and must prove otherwise. Both the European Parliament’s own initiative Report on Copyright and AI Generative AI and the French Senate bill are cited in support of this approach. The effect of this move would be to shift the compliance burden: where the opt-out required rightsholders to act in advance, these proposals make AI developers demonstrate compliance.

If the Commission is looking for the path of least resistance, it is in the area of transparency. Work-level disclosure of training data is demanded by 185 submissions and a near-unanimous demand among rightsholders (172 in favour, 3 preferring summaries only). On the other side, the technology industries’ counter-position — aggregate summaries, respect for trade secrets — musters just 13 submissions. No other concrete demand in the corpus commands support this broad or opposition this weak.

A third area is quietly consensual: 155 submissions call for shared rights infrastructure — registries, rights-data and metadata frameworks, content identifiers, provenance databases — against only 25 who oppose interventions in this area. Unlike transparency, this support is genuinely cross-camp, spanning rightsholders, technology companies, academics, and public authorities.

The research strand is a useful control case. On a mandatory, contract-proof research exception and a secondary publication right, the coalition geometry reverts to the classic rightsholders-versus-users axis: 73 in support (51 from the library and research sector) against 58 opposed (57 of them rightsholders). The libraries that stand alongside the technology industry on TDM stand alone here. This confirms that the shared defence of the TDM framework is an alliance of convenience, not of conviction — and that the public-interest case for it needs to be made on its own terms.

Finally, on remuneration, the field thinks almost exclusively in terms of licensing. Direct market licensing is the preferred route in 125 submissions — typically the large rightsholders, who point to existing deals with AI developers as proof the market works. A further 76 want licensing administered collectively through CMOs, and 49 (mostly creator and performer organizations) frame the question more fundamentally as one of prior consent, the remuneration-side counterpart of the demand to exclude AI training from the TDM exception. Forty submissions, mostly tech and cross-sector business, oppose any new remuneration mechanism.  Statutory remuneration in the form of a levy on AI providers is endorsed by just seven submissions — including our own.

The demand side of the consultation, in other words, asks the Commission to make licensing work; almost nobody asks whether licensing can work.

Bridging the gap

Measured against the Commission’s dual objective, almost the entire corpus of responses fails on one of the dual objectives:

Hardening the opt-out or excluding training from the exception delivers control and a licensing obligation, but makes access to training data harder. More importantly, it rests on the assumption that licensing markets can reach the breadth of content that model training actually requires — an assumption unlikely to hold. If it does not, European AI developers will face a structural disadvantage against anyone operating outside the EU.

Preserving the status quo would continue to deliver access but would leave the remuneration objective empty. Structurally, it benefits the incumbent AI developers. As long as there is no workable opt-out system, it also creates significant legal uncertainty for smaller EU-based model developers.

As is to be expected at this early stage of the process, each camp resolves the Commission’s tension by ignoring the half of it that would solve the issues faced by stakeholders on the other side.

But there is a small set of submissions that do something different: They relocate the point at which the obligation attaches. Rather than fighting over authorization at the training stage, these responses attach the obligation for model developers downstream from training at the deployment stage where AI systems generate revenue. There is a considerable amount of variation among this group of proposals, reflecting the largely exploratory nature of these responses. Our own submission argues for replacing the Article 4 opt-out with a non-waivable remuneration right — a statutory levy on the commercial deployment of models trained on publicly available content, paired with broad redistribution beyond the group of traditional beneficiaries of copyright.

The Max Planck Institute for Innovation and Competition, whose long-term proposal creates new exclusive rights at both the training and deployment stages, carves web-scraped content out of the training-stage right, replacing exclusivity with a statutory fair-compensation claim, and recommends a flat-rate benefit-sharing levy on developers as a transitional measure while the larger reform is built.

Mistral proposes a two-track model: a statutory levy on commercial AI revenue for pre-training, with licensing for fine-tuning and retrieval downstream— a refined version of positions that the company has advocated publicly since early 2026.  Relatedly, France Digitale, also representing the developer side, asks the Commission to explore a full TDM exception without opt-out but with mandatory compensation. The proposal from the UK-based CREATe centre reaches a similar destination via a lifecycle split: free pre-market research and development, licensing obligations at commercialization— a position previously articulated in response to the UK government’s consultation on AI and copyright. Euroconsumers, the one consumer voice, argues for compensating creators “at the monetisation phase” rather than at training.

Finally, COMMUNIA and Creative Commons both propose an arrangement that keeps the TDM exception plus opt-out in place but adds a levy on commercial AI systems to finance an information ecosystem fund — sustaining the shared infrastructure AI depends on rather than compensating individual acts of copying.

These submissions differ on the instrument — and on how far they would relieve the training stage — but they share a common architecture: the payment obligation attaches at deployment, not at ingestion. This builds on the acknowledgement that training and deployment are two fundamentally different phases of the AI economy. And the distinction between training-stage and deployment-stage uses is acknowledged well beyond the group that makes it the lever of its proposals: 109 submissions treat training and inference as distinct acts, 73 of them from rightsholders. But the rightsholder majority uses the distinction to add a second control point, licensing inference and retrieval on top of a controlled training stage. The bridging group uses it the other way around: relieving the training stage is precisely what buys the guaranteed, market-wide remuneration downstream.

This approach is, as far as the responses to the call for evidence are concerned, the only approach with the potential and the intent to meet both of the Commission’s objectives at once. It would give developers the friction-free access to training data that the competitiveness agenda demands, and give creators a remuneration claim that — unlike licensing — reaches all commercial deployments, including those of developers who would never sign a licence. On the distribution side, it can provide the basis for mechanisms that reach the long tail of creators who lack the bargaining power to strike deals. While licensing seems like the natural answer to the questions raised by the consultation, in a market that is highly concentrated on the demand side it will only work for those large enough to negotiate, and only reach the models whose developers choose to participate. Hardly any of the responses engage with this limitation — even though, as we have argued elsewhere, the inability of licensing to reach beyond the largest rightsholders and the willing developers is by now well documented.

What the Commission should take from this

Two findings from this corpus are immediately actionable. Work-level transparency and shared rights infrastructure command broader and less contested support than any position on the exceptions themselves, and both are compatible with every remuneration model on the table. They are the obvious constructive ground, and the Commission should build on them regardless of where it lands on the harder questions.

But transparency and infrastructure are enablers; they do not resolve the underlying conflict over authorization and remuneration. On that conflict, the consultation offers the Commission two well-organized camps, each demanding half of the Commission’s own objective, and a small, heterogeneous group pointing at the only visible way to deliver both halves. That group’s proposals are heterogeneous and largely exploratory — the differences between a compensated exception, a transitional levy and a full statutory remuneration right are not details. But the record of this consultation supports one narrow conclusion: of everything the Commission has received, these are the only submissions that treat its stated objective as a single problem rather than as two halves to choose between.

Paul Keller

Footnotes

  1. This coding is by definition lossy and will not capture the full nuance of every response. Methodology used: the full text of all 432 inline comments and 339 attachments has been downloaded from the Commission’s website, and each response was classified by a large language model (Claude Sonnet 4.6) against a fixed rubric of fourteen axes — four describing who the respondent is (camp, sector, actor type), ten describing the positions taken — each with a controlled vocabulary and per-value definitions. Submissions that are silent on an axis are coded as taking no position, and a subset of submissions was additionally read in full. This type of machine classification carries a margin of error: the counts are reliable in broad direction, but are best read as orders of magnitude and not as precise tallies.^
  2. The Commission’s own respondent-category labels do not align well with the composition of the field and have therefore not been used. This shortcoming is most visible when it comes to the sizable group of collective management organisations that self-identified as a mix of “other”, “NGO” or “company/business”. ^
keep up to date
and subscribe
to our newsletter
Subscribe