What OpenAI's Statement Covers — and What It Doesn't
On July 31, OpenAI published Advancing Responsible AI Across Europe, a blog post timed to the EU AI Act's next enforcement phase. The document is organized around three areas that map to two of the three chapters of the GPAI Code of Practice: the Transparency chapter and the Safety & Security chapter.
On the Transparency side, OpenAI points to its Frontier Governance Framework (published May 28, 2026), which maps internal safety practices onto disclosure obligations. It references the Preparedness Framework (in place since 2023, updated April 2025), published system cards, the Red Teaming Network, and the public Model Spec. All of this is substantive — but it's about how OpenAI governs itself internally, not about what data went into the models.
On Safety & Security, the post describes the Trusted Access for Cyber (TAC) program, launched under the EU Cyber Action Plan in May 2026, under which OpenAI works with EU and national cyber agencies and critical infrastructure operators. This aligns with the European Commission's Action Plan on Cybersecurity and AI from July 7, 2026.
On provenance, the post details C2PA Content Credentials and SynthID watermarking — more on that below. All solid infrastructure work.
What's missing is the third chapter: Copyright. The GPAI Code requires two things from every signatory — including OpenAI as an endorsed signatory: a publicly available summary of training data, using the mandatory template the European Commission published on July 24, 2025, and a documented policy for complying with EU copyright law and the text-and-data-mining opt-out regime. The July 31 statement addresses neither.
The Quiet Help Center Update
While the blog post sidesteps the Copyright chapter, OpenAI's EU AI Act Help Center page was updated roughly 22 hours before this article — it now links to two training data summaries in PDF format: one for GPT-5.5 and one for GPT-5.6 Luna. The timestamps suggest this update was pushed between the Tech Times article going live and the August 2 deadline.
But the Help Center update is a stopgap, not a strategy. The blog post — the high-profile, press-facing statement — still contains zero mention of training data, copyright compliance, or Article 53 obligations. And critically, GPT-5, released on August 7, 2025 — five days after GPAI obligations first took legal effect — still has no published training data summary. Under the timeline analyzed by WilmerHale, models released after August 2, 2025 had no transitional window: they were required to comply on their launch date.
What the EU AI Act Actually Requires — and What's at Stake
Starting August 2, 2026, the European AI Office gains the power to request information, access models, and impose fines for non-compliance with GPAI obligations. The numbers are not symbolic:
- €15 million (approximately $17.2 million at late-July 2026 exchange rates) or 3% of global annual turnover — whichever is higher.
- For context, OpenAI's estimated 2025 revenue was in the $3.7 billion range. Three percent of that is roughly $111 million — well above the €15 million flat cap.
The enforcement powers apply to all three chapters of the GPAI Code: Transparency, Safety & Security, and Copyright. The Copyright chapter's training data summary is not optional, and it's not something you can satisfy by pointing to internal governance frameworks. It requires a published document, in a specific format, available for rights holders and regulators to inspect.
The template the Commission published in July 2025 doesn't require full dataset disclosure — it requires category-level information: what type of data was used (text, image, video, audio), what the main sources were, and how copyrighted materials were handled. This is a disclosure framework, not a data dump. The fact that OpenAI didn't even mention it in its flagship compliance blog post is the gap that caught attention.
C2PA and SynthID: What They Actually Do
The provenance work OpenAI describes is technically real and worth explaining, because it's easy to either overhype or dismiss this technology.
C2PA Content Credentials embed a cryptographic manifest in a file's metadata layer. The manifest records which AI system generated the content, when, and whether the file was subsequently edited. Any C2PA-compatible tool can read it. The catch: metadata can be stripped by a screenshot, a platform re-encode, or a format conversion. It's a good first layer; it's not survivable.
SynthID, developed by Google DeepMind and licensed to OpenAI, addresses that limitation by embedding the provenance signal in the pixel data itself — imperceptible modifications that persist through compression, resizing, and cropping. The watermark is detectable by a neural network, invisible to human perception. OpenAI brought this to all ChatGPT and API image outputs in May 2026, and is expanding to audio.
The combination — C2PA for rich context, SynthID for survivability — is a layered approach that the GPAI Code explicitly endorses. No single watermarking technology currently meets all four criteria Article 50 imposes: effectiveness, interoperability, robustness, and reliability. A layered architecture is the Code's suggested answer.
But provenance technology addresses Article 50 — transparency of AI-generated content. It does not address Article 53(1)(d) — the obligation to publish a training data summary. These are separate requirements in separate sections of the Act, and compliance with one doesn't substitute for the other.
Industry Context: Everyone Is Scrambling
OpenAI is not alone in racing the August 2 deadline. Google announced on July 24 that it was signing the Transparency Code of Practice and expanding SynthID watermarking partnerships to Apple, ElevenLabs, Kakao, NVIDIA, and OpenAI. Google also cautioned that layering additional rules onto evolving technical frameworks could cause "disclosure fatigue" — a diplomatic way of saying the industry thinks the compliance infrastructure is immature.
The European Commission published a list of over 180 organizations that have signed the Code of Practice on transparency of AI-generated content. But a benchmark study published in mid-2026 examining documentation quality across GPAI models found that signatories "score only marginally higher overall than non-signatories," with the advantage concentrated in downstream-facing documentation rather than "upstream disclosures on training data, copyright-relevant data use, bias mitigation, computing and energy consumption."
Separately, the industry's broader training-data problem isn't theoretical. Since 2022, regulators and courts have imposed over $3.5 billion in fines and settlements on major technology companies for AI-related data violations — led by Anthropic's $1.5 billion settlement for training AI on pirated books and Meta's $1.4 billion for biometric data collection.
The Dublin Dimension
For OpenAI, the stakes are geographically concentrated. OpenAI Ireland Limited is the EU data controller, and Ireland's AI Office handles frontline GPAI enforcement. OpenAI signed an 88,000-square-foot lease at Dublin's Tropical Fruit Warehouse on July 27, 2026, committing to 250 new roles over two years and a €105 million (approximately $120.7 million) total investment.
Non-compliance wouldn't just trigger fines — it would risk disrupting European commercial operations at a time when OpenAI, approaching a potential IPO, needs to demonstrate structural durability to public-market investors. The Grok investigation — where the European Commission opened formal DSA proceedings against X after Grok produced roughly 3 million non-consensual intimate images — provides a template for how quickly EU enforcement can escalate when a specific AI harm is documented.
What Happens Next
OpenAI's blog post is legally careful. It describes compliance practices as "supportive of the EU framework" rather than claiming verified compliance, acknowledges that provenance technology "remains imperfect," and frames its approach as ongoing — "we will keep strengthening our compliance approach and learning from regulators and the broader ecosystem."
That's prudent. But it also means the statement does not function as a compliance certification. The European AI Office can now begin verifying GPAI compliance through information requests and, where necessary, enforcement actions. The Copyright chapter's training data summary requirement cannot be met by publishing internal frameworks or watermarking technology — it requires a published document, in a specific format, for each covered model.
The Help Center update with GPT-5.5 and GPT-5.6 summaries suggests OpenAI knows the gap exists and is closing it piecemeal. But for a company that just leased 88,000 square feet in Dublin and is counting on European revenue for its growth story, the Copyright chapter deserves better than a last-minute PDF upload and a blog post that pretends the chapter doesn't exist.
What exactly does the EU AI Act require from OpenAI regarding training data?
Article 53(1)(d) of the EU AI Act requires every provider of a general-purpose AI (GPAI) model — including OpenAI — to publish a publicly available summary of the content used to train their models, using a mandatory template the European Commission published on July 24, 2025. The summary must describe data types, main sources, and how copyrighted materials were handled. A documented copyright compliance policy covering the EU's text-and-data-mining opt-out regime is also required under Article 53(1)(c).
What fines can OpenAI face for non-compliance?
From August 2, 2026, the European AI Office can impose fines of up to €15 million (approximately $17.2 million) or 3% of global annual turnover, whichever is higher. For OpenAI, with estimated 2025 revenue around $3.7 billion, 3% would be approximately $111 million — well above the €15 million flat cap. Pre-August 2025 models have a transitional deadline of August 2, 2027. Models released after August 2, 2025 — including GPT-5 — had no transitional window.
Has OpenAI published any training data summaries at all?
Yes, but selectively. In a quiet update to its EU AI Act Help Center page timed around the enforcement deadline, OpenAI added PDF summaries for GPT-5.5 and GPT-5.6 Luna. No summary exists for GPT-5, released August 7, 2025 — the model that fell within the immediate compliance window. The company's July 31 blog statement "Advancing Responsible AI Across Europe" does not mention training data, copyright compliance, or Article 53 obligations at all.