← All insights
AI crawlingCopyrightPolicyEU AI ActJapan

The same AI crawler is a thief in Tokyo and a guest in Brussels

Slava Spitsyn · June 20, 2026 · 3 min read

A single AI crawler can hit a Japanese site, a German site, and an American site in the same minute — and be legal in one place, restricted in another, and headed for court in the third. The bot doesn't change. The law underneath it does. And most site owners have no idea which rulebook is theirs.

What's new

Three of the largest content markets on earth have landed in three different places:

  • Japan — broadly permissive. Article 30-4 lets models train on copyrighted work without consent, treating it as non-expressive data use.
  • European Union — leaning toward consent and transparency. The DSM Directive lets you reserve your rights; the AI Act makes GPAI providers respect those reservations and disclose what they trained on.
  • United Statesunsettled. There's no training statute; it's being fought case by case, with "fair use" argued on both sides.

How it works

The split comes from how each system frames the same act:

  • Japan asks: is the model enjoying the expression? If not, it's allowed.
  • The EU asks: did the rightsholder reserve this use? If yes, respect it.
  • The US asks: is this a transformative fair use? And then lets the courts decide, slowly, one dispute at a time.

Same crawl, three completely different questions. Which one governs you can depend on where you're based, where the model maker is based, and where the data sits.

Behind the news

For a website owner, the patchwork has a brutal practical edge:

  • You probably don't know which regime applies to any given crawl — and it may be more than one.
  • The law is moving slowly. Directives, court cases, and codes of practice take years. Crawlers don't wait for them to settle.
  • The crawl already happened. While the jurisdictional question is being argued, your content has been taken, logged as one more anonymous IP, with no record of what left.

The map of rights is fragmented. The act of access is not — it happens the same way everywhere.

Why it matters

You can't pick your jurisdiction, and you can't speed up the law. But there's one thing that stays constant across all three regimes: you can't govern what you can't see.

Whether your protection comes from a reservation (EU), a lawsuit (US), or nothing much at all (Japan), every one of those paths needs the same raw material — a record of who accessed your content and when. That record is jurisdiction-agnostic. It's useful no matter which rulebook turns out to be yours.

How I see it

I stopped trying to predict which legal regime will win. They may never converge — and building on top of a moving, fragmented law is building on sand.

So WARD is built one layer below the law, on the thing that's the same everywhere:

  • A declaration of how your content may be used — which maps onto whichever regime applies.
  • A record of who actually accessed it — which is evidence under all of them.

The map of rights is a patchwork. The record of access shouldn't be.

getward.org

WARD is an open standard for AI access control on the web.

Declare how AI may use your content — and keep a record of who came, when, and what they took.