The Disclosure Rule Nobody in Clipping Noticed
Since 2 August 2025, providers of general-purpose AI (GPAI) models selling into the EU have been legally required to publish a public summary of what they trained their models on. The obligation comes from Article 53(1)(d) of the AI Act (Regulation (EU) 2024/1689), and on 24 July 2025 the European Commission's AI Office published the actual tool for doing it: a standardized template and explanatory notice, built to give "a common minimal baseline for the information to be made publicly available in the Summary of Training Content," as the Commission's own library page describes it.
Alongside the template, the Commission's AI Office finalized the General-Purpose AI Code of Practice on 10 July 2025 — a voluntary framework that, per Article 56, gives signatories a "presumption of conformity" with Articles 53 and 55 while the AI Office develops harmonized standards. The Code has three chapters: transparency (documentation for every model), copyright (rights-respecting data practices), and safety-and-security (for the small number of models trained above the 10^25 FLOP systemic-risk threshold).
Two dates matter for anyone tracking enforcement. Models placed on the EU market from 2 August 2025 onward had to comply immediately, with no grace period. Models already on the market before that date have until 2 August 2027 to publish a compliant summary — but the Commission's power to actually verify compliance and issue corrective measures against GPAI providers starts 2 August 2026, according to the AI Act's implementation timeline and Article 111's transitional rules. That's less than two weeks from today. Fines for non-compliance under Article 101 run up to €15 million or 3% of a provider's global annual turnover, whichever is higher.
What the Summary Actually Has to Say
The template isn't a checkbox exercise. Per a summary of its contents published by WilmerHale, providers have to identify the model and its version, the modalities it was trained on (text, image, video, audio), estimated dataset size, and language and demographic coverage. On data provenance specifically, providers must name the major public datasets used, describe licensed and third-party sources, and — for anything scraped from the open web — list a meaningful share of the domains crawled, along with when the crawl ran and how it was conducted. WilmerHale's reading of the template puts that domain-disclosure threshold at roughly the top 10% of scraped domains for most providers, a lower bar (around 5%, or 1,000 domains) for smaller companies; that specific percentage could not be independently verified against the Commission's raw template document here and is presented as WilmerHale's characterization rather than a direct quote from Brussels.
What is directly attributable to the Commission's own framing is the purpose: the summary exists, in the AI Office's words, to give the public "a comprehensive overview of the data used to train a model" and to help enforce both copyright and data protection law simultaneously. That dual purpose is what makes the template relevant to platforms that have never touched an AI lab. If your platform hosts a public library of creator clips, product UGC, or campaign content that's crawlable from the open web, that domain can show up — or fail to show up — in a foreign AI company's legally mandated disclosure. For the first time, there's a standard document to check.
The Opt-Out Mechanism Buried in the Code of Practice
The transparency template only tells you what a model was trained on after the fact. The actual lever for keeping content out in the first place sits in a different piece of law: Article 4(3) of the DSM Copyright Directive (2019/790), which entitles rightsholders to reserve their content against the EU's general text-and-data-mining (TDM) exception — commonly called the TDM opt-out. Ordinarily, TDM over lawfully accessible content is copyright-exempt across the EU; Article 4(3) lets a rightsholder switch that exemption off for their own material.
The GPAI Code of Practice operationalizes that opt-out for AI training specifically. Under what Clifford Chance's analysis of the Code calls Measure 3, signatories commit to "use web crawlers that identify rights reservations via the Robot Exclusion Protocol (robots.txt), and identify and comply with other appropriate machine-readable protocols, such as llms.txt," and to give rightsholders visibility into how their crawlers and reservation-handling actually work. Crucially, the Code doesn't require robots.txt specifically — it explicitly preserves "rightsholders' ability to reserve rights by any appropriate means," so a platform isn't locked into one technical standard to make a reservation stick.
For a platform whose entire catalog is public-facing UGC — clip pages, creator profiles, campaign landing pages — this is the mechanism that determines whether that catalog is fair game for a Code-signatory AI lab's next training run, or whether it's been affirmatively reserved against it.
Parliament Wants a Registry and an Itemized List
The Code of Practice and the training-data template are Commission-level implementation tools. The European Parliament thinks they don't go far enough, and said so formally. On 10 March 2026, MEPs adopted a resolution on copyright and generative AI by a vote of 460 to 71, with 88 abstentions, according to the Parliament's own press release. The resolution asks the Commission to go beyond aggregate domain lists and require AI providers to disclose "an itemised list of all copyrighted works used to train AI," along with detailed crawling records.
On opt-outs, Parliament wants the mechanism formalized further: it floated giving the European Union Intellectual Property Office (EUIPO) a role managing a centralized opt-out registry, rather than leaving reservation entirely to scattered robots.txt files and per-site llms.txt signals. Rapporteur Axel Voss framed the goal as symmetry: "Legal certainty would let AI developers know which content can be used and how licences can be obtained," he said, while "rightsholders would be protected against unauthorised use." Parliament separately rejected a flat-rate licensing fee floated during the debate as a blanket compensation mechanism, and it added a media-specific provision: news outlets whose traffic is diverted by AI systems should retain the right to refuse their content's use entirely, with full compensation when it is used.
The resolution is non-binding — it's a request to the Commission, not new law — but it signals where the next round of legislative pressure is headed: itemized disclosure rather than domain-level summaries, and a registry rather than a patchwork of machine-readable files.