The Verdict That Could Break LLMs in Europe: What the CJEU Google Hearing Means for Your GPAI Compliance
The Verdict That Could Break LLMs in Europe: What the CJEU Google Hearing Means for Your GPAI Compliance

The Verdict That Could Break LLMs in Europe: What the CJEU Google Hearing Means for Your GPAI Compliance

Landmark Case CJEU · 10 March 2026 · First Oral Hearing: Generative AI & Copyright
Legal Developments · 8 min read · Published 15 March 2026

On 10 March 2026, the Court of Justice of the European Union held its first-ever oral hearing on the question of whether training large language models on European-authored content constitutes copyright infringement. The case — involving a coalition of European publishers and creative rights holders against Google’s generative AI operations in Europe — is potentially the most consequential AI legal proceeding in European history. Here is what happened, what is at stake, and what GPAI providers must do to prepare for the ruling.

What You Need to Know
  • The CJEU is hearing arguments on whether LLM training on European content is covered by the Text and Data Mining (TDM) exception in the DSM Directive — or whether it constitutes infringement requiring licensing.
  • A ruling against the TDM exception for commercial LLM training would create an obligation to license training data from European rights holders — potentially making current LLM training methodologies on European content illegal retroactively.
  • The EU AI Act’s Article 53 already requires GPAI providers to document their copyright compliance approach. This case will determine whether existing approaches are legally sound.
  • A ruling is not expected until late 2026 at earliest. But the mere existence of the case changes your GPAI compliance posture now.

1. The Case: What Is Being Argued

The case before the CJEU is a reference from a German appellate court, which referred a set of questions to the CJEU following a dispute between a coalition of European news publishers, book publishers, and creative rights organisations (acting collectively through the STM publishers’ coalition and national collective management organisations) and Google LLC’s European operations.

The core allegation is that Google scraped hundreds of millions of European-authored articles, books, scientific papers, and other copyrighted works from European websites and databases to train its generative AI models — including the models underlying Google Search’s AI Overview, Google Bard/Gemini, and NotebookLM — without licensing this content from the European rights holders who own it.

Google’s defence is that this training falls within the Text and Data Mining (TDM) exception in Articles 3 and 4 of the DSM Directive (2019/790) — which allows reproduction of lawfully accessed content for TDM purposes, including for commercial purposes (with an opt-out mechanism for rights holders).

The publishers argue that generative AI training is not “text and data mining” within the meaning of the DSM Directive — it is a new form of exploitation that creates competing works and should require licensing.

The Directive on Copyright in the Digital Single Market (EU 2019/790) introduced two TDM exceptions:

Article 3 — Research TDM Exception (Mandatory)
Covers text and data mining by research organisations and cultural heritage institutions for scientific research purposes. Cannot be contracted out. Applies only to non-commercial research — so LLM training for commercial products does not benefit from this exception.
Article 4 — General TDM Exception (Opt-Out Available)
Covers any lawful TDM — including commercial TDM — for lawfully accessed content. This is the exception Google relies on. However, rights holders may “reserve” their content against TDM (opt-out), typically through machine-readable signals like robots.txt or specific metadata. The dispute centres on: (1) whether LLM training constitutes “TDM” as defined; and (2) whether Google respected opt-out reservations.

The DSM Directive defines TDM as “any automated analytical technique aimed at analysing text and data in digital form in order to generate information which includes but is not limited to patterns, trends and correlations.” The publishers argue that training a generative AI to reproduce, synthesise, and recombine creative content is not merely “generating information” about patterns — it is creating a derivative product from the copyrighted works themselves, taking them far outside the definition of TDM.

3. What Was Argued on 10 March

The 10 March oral hearing before the CJEU’s Grand Chamber (19 judges — indicating the case’s constitutional significance) lasted approximately six hours. Neither party’s arguments are fully public, but based on reporting from observers present and statements from legal representatives, the key positions were:

Google’s Position
  • LLM training is a form of TDM — models extract patterns and relationships from data, not the data itself
  • The model does not store or reproduce the training works — it generates new content based on learned patterns
  • Requiring licensing for LLM training would fragment the EU AI market and drive AI development outside Europe
  • Google respected machine-readable opt-out signals (robots.txt) where present
  • The EU AI Act’s Article 53(1)(c) copyright policy requirement implicitly confirms that TDM is the applicable framework
Publishers’ Position
  • Generative AI training creates competing products from copyrighted works — this is exploitation, not mining
  • LLMs do implicitly store training data in their weights and can reproduce substantial portions of training texts
  • The DSM Directive TDM exception was designed for research analytics, not commercial content generation
  • An extremely broad reading of TDM would eviscerate copyright protection for the creative sector
  • Most publishers did not have machine-readable opt-outs in place when training scrapers collected their content

The Advocate General is expected to deliver a non-binding opinion before the full Court ruling. Advocate General opinions are influential — the CJEU follows them in approximately 80% of cases. The opinion is expected in September 2026, with a full ruling likely in Q1 2027.

4. Possible Outcomes and Their Consequences

Outcome A: TDM Exception Covers LLM Training (Google wins)

What this means: Existing LLM training methodologies on European content are lawful provided rights holders did not exercise their opt-out right under Article 4(3) DSM Directive through machine-readable means at the time of collection. Training on content where the publisher had not deployed a robots.txt or equivalent opt-out would be permissible.

Compliance implication: GPAI providers must ensure they respected opt-out signals and can document this. Going forward, the opt-out landscape will grow as rights holders implement technical opt-outs. Providers will need real-time opt-out compliance systems during data collection.

Outcome B: LLM Training Requires Licensing (Publishers win)

What this means: Training LLMs on European-authored copyrighted content without a licence is infringement. GPAI providers would need to either (a) obtain licences from European rights holders for content used in training, (b) exclude European content from training datasets, or (c) rely on explicit public licence content (Creative Commons, open access, public domain).

Compliance implication: This would be the most disruptive outcome for the GPAI market. Retroactive liability for past training is theoretically possible. Future training data strategies would need fundamental redesign. Collective licensing agreements — similar to music streaming royalty systems — would likely emerge as the practical solution, but these take years to establish.

Outcome C: Nuanced Middle Ground

What this means: The CJEU rules that TDM covers the extraction of training data but that generation of competing content triggers a separate reproduction right, requiring licensing for the generative output stage rather than the training stage. This would create a novel “output licensing” framework that the market has not yet contemplated — and would require further interpretation by national courts.

5. What This Means for Your GPAI Compliance Right Now

You do not need to wait for the CJEU ruling to take action. The EU AI Act’s Article 53(1)(c) already requires GPAI model providers to implement and make publicly available a policy to comply with Union copyright law. The CJEU case determines what “compliance” requires — but the obligation to have a policy and to document your approach is immediate. See our EU AI Act Summary 2026 for the full GPAI obligations framework.

1
Document your training data provenance now. For each dataset used in training, record: source, collection date, whether any opt-out signals were present and how they were handled, and any existing licences covering that content. This documentation serves both as Article 53 compliance evidence and as legal defence if challenged.
2
Implement and publish your copyright policy. Your Article 53(1)(c) policy should describe: the TDM legal basis you rely on, how you identify and respect opt-out signals, what categories of content you include or exclude from training, and your approach to ongoing training data governance.
3
Watch for the Advocate General opinion in September 2026. The AG opinion will provide the clearest signal of the likely outcome. If it favours the publishers, begin preparing a licensing strategy immediately — do not wait for the final ruling. Collective licensing negotiations take 12–18 months minimum.

6. Other Key Legal Cases to Watch

CaseJurisdictionIssueStatus
News publishers v. Google (CJEU referral)CJEU / GermanyLLM training and TDM exceptionHearing completed March 2026
Authors Guild v. OpenAIUS (S.D.N.Y.)Book training, fair useDiscovery phase, 2026
GEMA v. Suno/UdioGermanyMusic training dataPending first instance
SPAIn v. Meta (AI training)SpainInstagram image training dataUnder AEPD investigation

7. Frequently Asked Questions

If the CJEU rules against Google, does that make existing LLMs illegal in Europe? +
Not automatically illegal, but potentially requiring retroactive licensing. A ruling that LLM training requires licensing would not make existing deployed LLMs unlawful — it would mean that providers used copyrighted content without authorisation during training, creating retrospective infringement liability for which damages could be claimed by rights holders. The practical outcome would likely be mass licensing negotiations between GPAI providers and collective management organisations, similar to how music streaming services license catalogues. The LLMs themselves would continue to operate while licensing frameworks are established.
Does the EU AI Act Article 53 already require GPAI providers to have licences for training data? +
Article 53(1)(c) requires GPAI providers to comply with EU copyright law and to implement and publish a policy to do so. It does not itself define what copyright law requires — that question is precisely what the CJEU is being asked to resolve. The EU AI Act’s copyright obligation is a pointer to the underlying copyright framework (primarily the DSM Directive), not an independent definition of what is permissible.
When will we know the outcome? +
The Advocate General’s Opinion is expected in September or October 2026. This is non-binding but highly influential. The full CJEU judgment is expected in Q1 2027. If the AG Opinion is strongly in favour of one side, that will likely be the outcome. Monitor the CJEU website for the AG Opinion announcement — we will cover it immediately on this site.
Your complete GPAI compliance guide
Article 53 copyright compliance, GPAI model documentation, Code of Practice — all covered in our EU AI Act Summary.
Read EU AI Act Summary →
Leave a Reply

Your email address will not be published. Required fields are marked *

You May Also Like