Over the past three decades, American output per hour increased by 85% versus 29% in Europe, reversing much of the post-World War II convergence. AI offers a window of opportunity to reverse this decline.

Yet European policymakers risk slamming this window shut by tinkering with the legal architecture that makes training possible: the Text and Data Mining exemption in the bloc’s copyright rules.

To understand the stakes, consider how AI models are built. Models do not copy or store data like a digital library. They learn from it, analyzing massive datasets of text, images, and code.

In 2019, the European Union established a forward-thinking approach via the Digital Single Market Copyright Directive. Articles 3 and 4 created a clear baseline:

  • Permission to Learn: Companies can use automated techniques to analyze publicly available data.
  • Rightsholder Control: Content owners retain the right to “opt out” via machine-readable code like robots.txt.

This stable framework gave European researchers, open-source developers, and startups the predictable legal environment needed to innovate without fear of constant lawsuits. It effectively institutionalized a machine-age equivalent of the human “freedom to learn.”

Get the Latest
Sign up to receive regular Bandwidth emails and stay informed about CEPA's work.

Despite its success, the European Parliament and several national governments are considering rolling back text and data mining protections. Media and copyright rightsholders argue for a shift toward a prior-authorization (“opt-in”) default or mandatory statutory levies on AI training data. In April 2026, the French Senate even adopted a bill introducing a restrictive legal presumption that AI models will exploit cultural content.

Reopening the text and data mining exception at this critical juncture would be a policy mistake. It threatens Europe’s competitiveness in three profound ways:

  1. Crushing Startups and Open Source: Shifting to a transaction-heavy, opt-in regime creates massive compliance and licensing costs. Large tech giants can afford armies of lawyers to negotiate these deals. European startups and open-source projects cannot. This regulatory burden directly drives market concentration.
  2. Exacerbating the Compute Deficit: Europe already faces massive energy and compute constraints compared to the US. If policymakers make Europe data-poor on top of being compute-poor, local AI development will stall entirely.
  3. Undermining Sovereignty: If European models are barred from learning from the open web, and open models cannot be freely fine-tuned by enterprise, SMEs, and individuals on open European data, choice will be restricted and dependence on offshore models, trained under more permissive legal frameworks, will grow. 

Although the desire to protect creators and rightsholders is entirely legitimate, targeting model inputs (the data used for learning) represents the wrong mechanism. A rigorous distinction must be maintained between training inputs and generated outputs. If computers are restricted in what they can read and develop, it will stifle development and create high entry barriers.

If an AI model generates an output that duplicates a copyrighted work, or arguably unduly exploits a creator’s likeness, copyright law should punish that output. But preventing a model from analyzing public text to understand grammatical structure or scientific concepts is the equivalent of banning a human art student from visiting a public museum.

If Europe wants true technological autonomy, it must possess the domestic capability to develop, adapt, and deploy frontier AI models. Policymakers should resist the temptation to unravel the 2019 text and data mining framework. Instead, the immediate task is to make the existing opt-out mechanisms workable, standardized, and predictable.

Europe’s AI developers must retain the freedom to allow their machines to learn in order to reap the vast rewards of the AI transformation in the arts, medicine, science, and industry. If the text and data mining exception is reopened, it will not save traditional creative industries. It will only ensure that Europe watches the next digital revolution from the sidelines.

Brian Williamson is a partner at the London-based Communications Chambers consultancy. He works at the intersection of technology, economics, and policy. Clients have included governments, regulators, telcos, and tech companies. This article is adapted from a paper “The Text and Data Mining Exception is Key to AI-Driven Progress.

Bandwidth is CEPA’s online journal dedicated to advancing transatlantic cooperation on tech policy. All opinions expressed on Bandwidth are those of the author alone and may not represent those of the institutions they represent or the Center for European Policy Analysis. CEPA maintains a strict intellectual independence policy across all its projects and publications.

Tech 2030

A Roadmap for Europe-US Tech Cooperation

Learn More
Read More From Bandwidth
CEPA’s online journal dedicated to advancing transatlantic cooperation on tech policy.
Read More