Copyright & AI: Why the European Commission Should Avoid Flawed Opt-Outs and Support International Standardisation
Main takeaways
- Mandating flawed, unstandardised opt-outs like TDMRep and CAWG to exclude content from TDM, including AI model training, would fragment the web, increase compliance costs, and weaken European AI innovation
- Asset-level metadata is not a scalable solution for copyright opt-outs: it can be altered, is routinely stripped, and cannot reliably reflect ownership
- Instead of picking winners, the Commission should rely on consensus-driven international standards bodies, such as the IETF, to improve established protocols
Will the European Commission make AI innovation significantly harder and more expensive by mandating ineffective copyright opt-out mechanisms for AI training? The question is technical, but the answer is critical to Europe’s innovation potential.
First, some context. The AI Act imposes copyright rules on providers of general-purpose AI (GPAI) models, including an obligation to identify and comply with reservations expressed by rightsholders under Article 4 of the EU Copyright Directive (EUCD). Article 4 creates a text-and-data-mining (TDM) exception, which enables AI training on publicly available content unless a rightsholder objects in a “machine-readable” format, known as an ‘opt-out’.
Article 4 intentionally uses broad wording to avoid locking the law into rigid technical protocols or picking winners, allowing markets and experts to develop technical solutions and standards. This AI Act obligation was recently translated into a voluntary Code of Practice for GPAI providers, under which signatories commit to read and follow instructions expressed through the robust and universal Robots.txt Protocol.
Signatories also commit to identify and comply with other protocols adopted by standardisation organisations, as well as state-of-the-art, technically implementable, widely adopted protocols that have been “generally agreed” through inclusive EU-level discussions.
1. The Commission’s unbalanced process
In this context, the European Commission asked stakeholders in early 2026 for their views on opt-out solutions that could be included in a list of generally agreed mechanisms. The Commission expected to reach an agreement after two stakeholder workshops – an extremely ambitious timeline. As anticipated, stakeholders across the board expressed disagreement at the first workshop in June, which led the Commission to delay the second.
Surprisingly, the Commission excluded trade associations representing technology companies from the workshops, leading to significant underrepresentation of AI developers. The process also raises other concerns. The Commission, which is not a standardisation body and is therefore only empowered to facilitate discussions, should not exceed its mandate by unilaterally imposing specific technical standards.
None of the seven opt-out solutions explored so far is fit for purpose, and two receiving particular attention from the Commission would be entirely unworkable: the TDM Reservation Protocol (TDMRep) and the Creator Assertions Working Group Protocol (CAWG).
2. TDMRep: A redundant and technically unworkable protocol
Developed by French publishers via a W3C community group, rather than a formal standards body, TDMRep is not a W3C standard and lacks balanced multistakeholder governance. It introduces TDM-reservation (a binary opt-out flag) and TDM-policy (a JSON file using ODRL to express terms like a “duty to compensate”).
In practice, TDMRep’s location-based controls offer no additional functionality over Robots.txt. Worse, conflicting instructions (such as ’TDM-reservation: 1’ – a binary exclusion – alongside ‘GPT-Bot: allow’) create severe legal uncertainty for web crawlers and rightsholders alike.
Furthermore, TDM-policy files attempt to create a veneer of machine-readability by embedding legal terms of service (ToS) in JSON files. But these preferences are not structured in a way that allows them to be translated into actionable instructions at web scale. That is why such files should be excluded from any mandatory implementation.
Controlled by a narrow group of publishers with vested interests, mandating TDMRep leaves developers vulnerable to arbitrary future requirements without technical oversight.
3. CAWG: Conflating provenance and copyright, compromising privacy
The CAWG protocol attempts to attach rights-reservation signals to individual digital assets – such as an image – using metadata layered onto the Coalition for Content Provenance and Authenticity (C2PA) specification. But C2PA is designed to record information about an asset’s provenance – where it came from and how it was modified – not to determine legal copyright ownership. A C2PA manifest can contain a camera or software signature, but that does not by itself verify who owns the underlying copyright (for example, a photo agency).
This conflation creates severe privacy and safety risks by attaching personal identity data directly to digital assets. C2PA deliberately excluded identity features because human rights groups warned that this would threaten journalists, whistleblowers, and dissidents. Forcing copyright enforcement into provenance tools risks fracturing the ecosystem and stalling C2PA adoption.
More generally, asset-based metadata mechanisms fail at scale. Embedded metadata lacks verification, inviting fraud as anyone can alter it. Standard web platforms routinely strip metadata during file compression, meaning persistence would require a disruptive overhaul of web architecture. Finally, static metadata snapshots cannot capture dynamic, transferable copyright – resulting in conflicting signals across duplicate files.
4. Rely on consensus-driven international technical standards
As it stands, no solution other than the Robots.txt Protocol offers the maturity, scalability, and technical robustness needed to express opt-outs. The Commission’s process should not mandate ineffective solutions, which would only raise the cost of innovation in the EU.
Rather than introducing fragmented, overlapping, or legally mandated opt-out mechanisms that risk making the TDM framework unworkable, the EU should rely on industry-led, international consensus standards. Specifically, the EU should support ongoing work within international bodies, such as the Internet Engineering Task Force’s (IETF) AI Preferences Working Group, to update widely adopted protocols like Robots.txt, enabling more effective and granular rights reservations.
Establishing a standardised, machine-readable vocabulary at global level ensures that rights reservations can be processed automatically and reliably at scale across jurisdictions, which also benefits rightsholders. This is key to avoiding conflicting signals or regional protocols that would defeat the purpose of rights reservations.
Conclusion
The success of the EU’s framework for AI and copyright hinges on robust, universally agreed, and market-led standards. Instead of mandating inappropriate technical tools, the European Commission should support the ongoing work of standardisation bodies with the relevant expertise and ability to foster market adoption.