← All insights
EU AI ActTransparencyPolicyAI training

The EU is forcing AI makers to disclose their training data. You still need your own record

Slava Spitsyn · June 18, 2026 · 3 min read

For years, the honest answer to "what was this model trained on?" was a shrug. The EU AI Act tries to end the shrug. From 2025, the makers of general-purpose AI have to publish a summary of their training content. It's a real shift — and it quietly raises a question for everyone whose work is in that data.

What's new

  • Article 53(1)(d) of the EU AI Act requires providers of general-purpose AI models to draw up and make public a "sufficiently detailed summary" of the content used to train the model.
  • The summary follows a template issued by the EU's AI Office, so disclosures are comparable across providers rather than freeform PR.
  • Combined with the copyright-policy duty in Article 53(1)(c), the message to AI makers is: show your work, and show that you respected opt-outs.

How it works

The obligation is deliberately a summary, not a file list:

  • It's meant to be detailed enough that rightsholders and regulators can understand the scope and main sources of the training data.
  • It is not a line-by-line manifest of every URL ingested. The point is accountability and the ability to exercise rights — not a full audit trail.
  • Enforcement and the exact bar for "sufficiently detailed" are still being tested as the first disclosures arrive.

So you get a map of the territory, drawn by the people who crossed it.

Behind the news

This is genuine progress. But notice the direction the information flows: from the AI maker, on their schedule, in their words.

  • A summary tells you the categories of data. It rarely tells you that your specific archive, on your domain, was taken on a specific date.
  • You are reading the other side's account of what happened to your property. Useful — but it's their ledger, not yours.
  • Transparency you don't control is a courtesy. It can be vague, late, or simply silent on the one crawl you care about.

The asymmetry is the problem. They know exactly what they took. You get a paragraph.

Why it matters

To actually exercise the rights the EU just handed you, you need a second source of truth — your own.

If a provider's summary and your own access record disagree, the disagreement is the leverage. You can't negotiate, license, or challenge anything when the only record of the event belongs to the other party.

A summary describes the past in general. A record proves a specific event. Rights live on the second one.

How I see it

I read the transparency rules as a half-built bridge. The EU made AI makers account for themselves — that's the far side. The near side, the part owners stand on, is still missing: a record kept by you, of who reached your content and when.

That's the layer I'm building with WARD. Not to replace the AI Act's disclosures, but to give the other side of the table something to check them against:

  • Your declaration of how content may be used.
  • Your log of who actually accessed it.

When their summary meets your record, transparency stops being a courtesy and starts being verifiable.

getward.org

WARD is an open standard for AI access control on the web.

Declare how AI may use your content — and keep a record of who came, when, and what they took.