The EU's AI Training-Data Rules Reach UGC Clip Libraries
Oarized · 22 July 2026
The Disclosure Rule Nobody in Clipping Noticed
Since 2 August 2025, providers of general-purpose AI (GPAI) models selling into the EU have been legally required to publish a public summary of what they trained their models on. The obligation comes from Article 53(1)(d) of the AI Act (Regulation (EU) 2024/1689), and on 24 July 2025 the European Commission's AI Office published the actual tool for doing it: a standardized template and explanatory notice, built to give "a common minimal baseline for the information to be made publicly available in the Summary of Training Content," as the Commission's own library page describes it.
Alongside the template, the Commission's AI Office finalized the General-Purpose AI Code of Practice on 10 July 2025 — a voluntary framework that, per Article 56, gives signatories a "presumption of conformity" with Articles 53 and 55 while the AI Office develops harmonized standards. The Code has three chapters: transparency (documentation for every model), copyright (rights-respecting data practices), and safety-and-security (for the small number of models trained above the 10^25 FLOP systemic-risk threshold).
Two dates matter for anyone tracking enforcement. Models placed on the EU market from 2 August 2025 onward had to comply immediately, with no grace period. Models already on the market before that date have until 2 August 2027 to publish a compliant summary — but the Commission's power to actually verify compliance and issue corrective measures against GPAI providers starts 2 August 2026, according to the AI Act's implementation timeline and Article 111's transitional rules. That's less than two weeks from today. Fines for non-compliance under Article 101 run up to €15 million or 3% of a provider's global annual turnover, whichever is higher.
What the Summary Actually Has to Say
The template isn't a checkbox exercise. Per a summary of its contents published by WilmerHale, providers have to identify the model and its version, the modalities it was trained on (text, image, video, audio), estimated dataset size, and language and demographic coverage. On data provenance specifically, providers must name the major public datasets used, describe licensed and third-party sources, and — for anything scraped from the open web — list a meaningful share of the domains crawled, along with when the crawl ran and how it was conducted. WilmerHale's reading of the template puts that domain-disclosure threshold at roughly the top 10% of scraped domains for most providers, a lower bar (around 5%, or 1,000 domains) for smaller companies; that specific percentage could not be independently verified against the Commission's raw template document here and is presented as WilmerHale's characterization rather than a direct quote from Brussels.
What is directly attributable to the Commission's own framing is the purpose: the summary exists, in the AI Office's words, to give the public "a comprehensive overview of the data used to train a model" and to help enforce both copyright and data protection law simultaneously. That dual purpose is what makes the template relevant to platforms that have never touched an AI lab. If your platform hosts a public library of creator clips, product UGC, or campaign content that's crawlable from the open web, that domain can show up — or fail to show up — in a foreign AI company's legally mandated disclosure. For the first time, there's a standard document to check.
The Opt-Out Mechanism Buried in the Code of Practice
The transparency template only tells you what a model was trained on after the fact. The actual lever for keeping content out in the first place sits in a different piece of law: Article 4(3) of the DSM Copyright Directive (2019/790), which entitles rightsholders to reserve their content against the EU's general text-and-data-mining (TDM) exception — commonly called the TDM opt-out. Ordinarily, TDM over lawfully accessible content is copyright-exempt across the EU; Article 4(3) lets a rightsholder switch that exemption off for their own material.
The GPAI Code of Practice operationalizes that opt-out for AI training specifically. Under what Clifford Chance's analysis of the Code calls Measure 3, signatories commit to "use web crawlers that identify rights reservations via the Robot Exclusion Protocol (robots.txt), and identify and comply with other appropriate machine-readable protocols, such as llms.txt," and to give rightsholders visibility into how their crawlers and reservation-handling actually work. Crucially, the Code doesn't require robots.txt specifically — it explicitly preserves "rightsholders' ability to reserve rights by any appropriate means," so a platform isn't locked into one technical standard to make a reservation stick.
For a platform whose entire catalog is public-facing UGC — clip pages, creator profiles, campaign landing pages — this is the mechanism that determines whether that catalog is fair game for a Code-signatory AI lab's next training run, or whether it's been affirmatively reserved against it.
Parliament Wants a Registry and an Itemized List
The Code of Practice and the training-data template are Commission-level implementation tools. The European Parliament thinks they don't go far enough, and said so formally. On 10 March 2026, MEPs adopted a resolution on copyright and generative AI by a vote of 460 to 71, with 88 abstentions, according to the Parliament's own press release. The resolution asks the Commission to go beyond aggregate domain lists and require AI providers to disclose "an itemised list of all copyrighted works used to train AI," along with detailed crawling records.
On opt-outs, Parliament wants the mechanism formalized further: it floated giving the European Union Intellectual Property Office (EUIPO) a role managing a centralized opt-out registry, rather than leaving reservation entirely to scattered robots.txt files and per-site llms.txt signals. Rapporteur Axel Voss framed the goal as symmetry: "Legal certainty would let AI developers know which content can be used and how licences can be obtained," he said, while "rightsholders would be protected against unauthorised use." Parliament separately rejected a flat-rate licensing fee floated during the debate as a blanket compensation mechanism, and it added a media-specific provision: news outlets whose traffic is diverted by AI systems should retain the right to refuse their content's use entirely, with full compensation when it is used.
The resolution is non-binding — it's a request to the Commission, not new law — but it signals where the next round of legislative pressure is headed: itemized disclosure rather than domain-level summaries, and a registry rather than a patchwork of machine-readable files.
What It Means for Clipping and Payout Platforms
None of this legislation was written with UGC clipping or creator-payout platforms in mind — it targets the AI labs building foundation models. But the mechanics land squarely on any platform whose value is a public library of creator content. A few things worth acting on now rather than after the 2 August 2026 enforcement date passes:
Check whether your domain shows up in a training-data summary. GPAI providers are now required to publish which domains they crawled. A platform with a large public clip or campaign archive can search its own domain against major providers' published summaries as those roll out through 2026 and 2027.
Decide on a reservation posture before assuming one by default. The Code of Practice recognizes robots.txt and llms.txt as valid machine-readable opt-out signals under DSM Article 4(3), but silence is not automatically read as consent or refusal — a platform that wants its public UGC library excluded from AI training needs to actively publish a reservation, not rely on ambiguity.
Separate the platform's reservation from the creator's. A clip or campaign page is typically a mix of platform-owned presentation and creator-owned content. Decide, in creator agreements and technical crawler policy alike, who has the right to opt content out of AI training — the platform, the creator, or both jointly — before a dispute forces the question.
Watch the EUIPO registry proposal, not just the Commission's existing template. If Parliament's push for a centralized opt-out registry advances, it would replace ad hoc per-domain signals with a single place to register a reservation — potentially simplifying compliance for platforms that currently rely on robots.txt alone.
The compliance deadlines here are aimed at AI labs, not at UGC platforms directly. But a platform's own content library is exactly the kind of asset these rules were built to make visible — and, for the first time, there's a documented, EU-mandated way to check who's used it and a real mechanism to say no in advance.