The biggest legal question in AI right now does not have one answer. In 2026, courts on four continents are simultaneously writing the rules for who owns AI-generated content, whether scraping copyrighted material to train a model is theft, and what it means to be the “author” of a work a machine produced. The U.S. Copyright Office just published its third and most consequential report. The U.S. Department of Justice weighed in for the first time. The EU AI Act’s Article 53 training-data transparency rules went from law to enforceable. The UK High Court dismissed the central copyright theory in Getty v. Stability AI. Beijing handed down the world’s first ruling that a Stable Diffusion image qualifies as a copyrightable work — and awarded it to the prompter, not the model.
What follows is a jurisdiction-by-jurisdiction read of where AI copyright law actually stands in September 2026, what it means for creators shipping work into AI training pipelines, for businesses deploying models that ingest third-party data, and for developers who want to know whether the code, copy, or image they generated yesterday is theirs today.
The two questions courts are answering
Every AI copyright case in 2026 reduces to one of two questions, and a complete answer requires both.
Question 1 — Inputs. Is it legal to train a generative AI model on copyrighted material without the copyright holder’s permission? This is the question at the center of The New York Times v. OpenAI (S.D.N.Y., filed December 2023, now in summary judgment posture per the AI Lawsuit Tracker), Bartz v. Anthropic (N.D. Cal., settled August 2025 for $1.5 billion — the largest U.S. copyright settlement in history), Kadrey v. Meta, Authors Guild v. OpenAI, Getty Images v. Stability AI (UK High Court, November 2025), and dozens more. U.S. courts apply the four-factor fair-use test from 17 U.S.C. § 107. EU regulators are using Article 53 of the AI Act. Japan leans on Article 30-4 of its Copyright Act. China is building doctrine case-by-case from the Beijing Internet Court.
Question 2 — Outputs. Can AI-generated output be copyrighted, and if so, who owns it? The U.S. answered this in March 2025 when the D.C. Circuit Court of Appeals affirmed Thaler v. Perlmutter — AI cannot be an author under the Copyright Act. The U.S. Supreme Court denied certiorari in March 2026, leaving that answer as final settled law. China went the opposite direction in Li Yunkai v. Liu Yuanchun (Beijing Internet Court, November 2023), holding that a Stable Diffusion output IS a copyrightable work — but the prompter is the author, not the model, and the AI tool must be prominently labeled. The EU has not issued a definitive ruling on output copyrightability; the AI Act is silent on it. See our related coverage of AI in creative industries for the artist-side perspective.
The remainder of this post walks each jurisdiction’s current rule, the cases that produced it, and the practical playbook for creators, businesses, and developers operating across borders.
United States — Inputs: fair use for training is unsettled, piracy is not
On May 9, 2025, the U.S. Copyright Office released the pre-publication version of Copyright and Artificial Intelligence, Part 3: Generative AI Training — a 108-page document the Office calls “the most consequential” in its AI series. The headline finding: training on copyrighted works “likely qualifies as fair use in some circumstances, but not in others.” The Office refused to endorse a categorical rule and recommended against compulsory licensing, instead endorsing voluntary collective licensing.
What the report says is concrete. Transformative use and market effects are the two fair-use factors judges will weight most heavily. Training is “highly transformative” when the model is doing something research-driven and non-expressive. Training is less transformative — and far more likely to weigh against fair use — when the resulting model competes directly with the market for the works it learned from. Market dilution, where AI’s speed and scale produces “unprecedented competition for sales of an author’s works,” is the new harm that did not exist at the time of Campbell v. Acuff-Rose. And training on pirated datasets — LibGen, PiLiMi, Books3, or any other shadow library — “almost certainly exceeds the boundaries of fair use.” That last line is the one that broke Anthropic’s case.
The first major test came in Bartz v. Anthropic. In June 2025, Judge William Alsup (N.D. Cal.) ruled that Anthropic’s use of legally acquired books for training was “quintessentially transformative” — fair use. But Anthropic had also downloaded over 7 million books from LibGen and PiLiMi, and Alsup held that downloading and keeping pirated copies was NOT fair use. The case was certified as a class action in August 2025. Anthropic settled for $1.5 billion — the largest copyright settlement in U.S. history — and agreed to destroy the unlawfully obtained files. The training-fair-use question remains alive in other circuits; Anthropic won the legal-acquisition argument but did not get a circuit ruling on it.
Meanwhile, The New York Times v. OpenAI (1:23-cv-11195, S.D.N.Y., Judge Sidney H. Stein) has become the bellwether. In March 2025, Judge Stein partially denied OpenAI’s motion to dismiss — direct and contributory infringement claims survive, DMCA §1202 claims related to stripped copyright management information survive, common-law unfair competition was dismissed. Discovery disputes dominated 2025 — including a Magistrate Judge order requiring OpenAI to produce 20 million ChatGPT logs (affirmed by Judge Stein in January 2026). In September 2026, the case entered summary judgment posture.
Then, on September 2, 2026, the U.S. Department of Justice filed a Statement of Interest — its first intervention in any AI copyright case — backing OpenAI and Microsoft. The DOJ argues that training large language models on copyrighted works “does not violate copyright law” because there is a legal distinction between copying for training and what the model actually outputs, and warns that restricting training would “severely hamper” American technological innovation. A Statement of Interest is not binding, but it is the first time the U.S. government has formally told a court that AI training is fair use.
United States — Outputs: AI is not an author, but humans using AI can be
On the output side, the U.S. rule is settled at the Supreme Court level. In Thaler v. Perlmutter, the D.C. Circuit Court of Appeals affirmed the Copyright Office’s denial of registration for an image Stephen Thaler claimed was autonomously generated by his “Creativity Machine” (DABUS). Judge Bradley wrote that the human authorship requirement is a “bedrock principle” of copyright law. The U.S. Supreme Court denied certiorari in March 2026, leaving the D.C. Circuit ruling as final law. AI cannot be an author.
But the Copyright Office’s 2023 Registration Guidance carved out a meaningful exception: a work that incorporates AI-generated material CAN be registered if a human author makes sufficient creative contributions. The “Zarya of the Dawn” precedent (the AI-assisted comic book) is the canonical example — the text was human-authored and the selection, coordination, and arrangement of AI-generated images were creative enough to support a partial registration. The line is human creativity, not human tooling.
For creators, the practical U.S. position is: if you prompt, iterate, select, and substantially edit AI output, you may own a copyright in the resulting work. If you accept raw first-pass output from a model, you own nothing — the work is in the public domain in the U.S. Either way, downstream users can copy, redistribute, and modify your output without asking, because there is no exclusive right to enforce.
European Union — Article 53 forces transparency, but training itself is not yet ruled on
On May 7, 2026, the European Parliament, Council, and Commission reached a provisional political agreement on the “Digital Omnibus” package — confirming that Article 53 of the EU AI Act’s transparency obligations for general-purpose AI (GPAI) providers stay anchored to the August 2, 2026 enforcement deadline. After that date, the AI Office can demand documents, order corrections, and impose fines of up to €15 million or 3% of global turnover.
Article 53 imposes four obligations on every GPAI provider placing a model on the EU market. (1) Maintain technical documentation of the model, including training and testing processes. (2) Provide information to downstream providers integrating the model. (3) Implement a copyright policy that complies with EU copyright law — particularly the text-and-data-mining (TDM) exception in Article 4 of Directive 2019/790, the DSM Copyright Directive. (4) Publish a “sufficiently detailed summary” of the content used for training the model, following the European AI Office’s Training Data Summary Template. The summary is public, meaning competitors can read it. (For a practical deep-dive on inheritance via Article 25, see the Disclos analysis.)
What does compliance look like in practice? The published summaries as of August 2026 show a gap. Google, Meta, and Microsoft filled in the template boxes. OpenAI filed one too — and was forced to disclose more than competitors initially expected. Anthropic, Mistral, and xAI replaced the structured template with narrative phrasing like “a proprietary mix of publicly available information and licensed data.” An academic paper from Trinity College Dublin and the Mozilla Foundation, presented at the 2026 ACM FAccT conference, evaluated the published summaries and concluded that even the filled-in ones do not permit real comparison. One industry commentator likened the state of things to “a treasure hunt — better than before, but stopping short of real comparability.”
The substantive question — whether training on copyrighted works without a license is itself a copyright violation under EU law — is currently before the Court of Justice of the European Union via a referral from the Hungarian courts. Judgment is not expected until 2027. Until then, Article 53 forces disclosure without resolving the underlying infringement question. The rightsholder opt-out regime from Article 4(3) of the DSM Directive requires AI developers to check whether a data source has a copyright reservation (commonly expressed as a robots.txt directive or an AI-specific opt-out metadata standard like the TDMRep protocol), exclude or license that content, and keep records of compliance.
United Kingdom, Japan, and China — three different models
On November 4, 2025, the UK High Court handed down its judgment in Getty Images v. Stability AI [2025] EWHC 2863 (Ch) — billed as the UK’s first major test of whether training an AI model on copyrighted images is infringement. By the end of trial, Getty had abandoned its primary copyright infringement and database rights claims because there was no evidence that training of Stable Diffusion took place in the UK. The court was left with two narrow questions: does Stability’s release of Stable Diffusion in the UK amount to secondary copyright infringement by importation, and does the model’s output of watermarked images constitute trademark infringement?
Justice Joanna Smith answered “no” to the first question and “yes, narrowly” to the second. On secondary copyright infringement, Stable Diffusion’s weights do not “store or reproduce” the training images — reproduction is critical for infringement, and absent reproduction the model is not an infringing copy upon importation. Getty’s secondary copyright claim was dismissed. On trademark, the court found that Stability had infringed Getty’s marks in narrow circumstances where its model generated outputs bearing Getty watermarks, but characterized the win as “extremely limited in scope.” The deepest signal was jurisdictional: primary infringement of UK copyright works could be evaded by training models elsewhere. The parallel US case, refiled in California in August 2025, is now the more determinative proceeding.
Japan is the jurisdiction where the “training AI on copyrighted works is legal” reputation lives. The provision behind that reputation is Article 30-4 of the Copyright Act, in force since January 1, 2019. It permits exploitation of a copyrighted work without the rights holder’s authorization where the use is “not for the purpose of enjoying or causing another person to enjoy the thoughts or sentiments expressed in it” — the “non-enjoyment” test. The exception does not apply where doing so “would unreasonably prejudice the copyright owner’s interests.” The Japanese approach is broader than the EU TDM exception in one dimension (commercial training is permitted) and narrower in another (the prejudice proviso is the operative constraint). What Article 30-4 covers is the training stage. When a model produces something similar to an existing work, ordinary infringement analysis applies: similarity and reliance.
On November 27, 2023, the Beijing Internet Court issued its judgment in Li Yunkai v. Liu Yuanchun — the first Chinese judicial decision on AI-generated content copyrightability. The court held that the AI-generated image QUALIFIES as a work of fine art under Chinese copyright law: the plaintiff’s iterative refinement through prompt selection, parameter adjustment, and aesthetic judgment reflected “personalized expression and originality.” The plaintiff (the prompter) is the author and copyright holder, but the plaintiff should prominently label the AI technology or model employed. China has no binding judicial decisions on AI training copyright infringement as of mid-2026, per a peer-reviewed analysis in MDPI Laws. The 2025 Interim Measures for the Management of Generative AI Services require pre-training, optimization training, and all data-related activities to “adhere to existing laws and regulations” without prescribing specific training-data rules.
What this means for creators, businesses, and developers right now
If you create original content — writing, art, music, photography, code — and you have not opted out of AI training, your work is most likely in the training corpora of multiple commercial models right now. The U.S. fair-use answer is unsettled. The EU requires opt-out compliance. Japan allows broad training under Article 30-4. The UK has a jurisdictional loophole. China’s training-stage rule is still emerging. There is no global rule yet. For a wider view of how this shapes the AI deployment landscape, see our coverage of how to use AI for your business in 2026.
The practical playbook for creators who want to control AI training of their work: (1) Express opt-out via TDMRep metadata and standard opt-out mechanisms — enforceable under EU law, a strong fair-use signal in the U.S., a trigger for Japan’s “unreasonable prejudice” proviso. (2) Register your copyrights with the U.S. Copyright Office — registration unlocks statutory damages and attorney’s fees. (3) Document your creative process — keep prompt logs, iteration histories, selection notes, and editing trails. (4) Watch the licensing market — Getty, the AP, the Financial Times, and News Corp have all signed training-data deals with major labs in the past 18 months.
The practical playbook for businesses deploying AI models: (1) If you fine-tune a foundation model substantially, you may inherit Provider obligations under EU Article 53. (2) Know the jurisdictional posture of every market you serve — selling into the EU requires Article 53 compliance; the UK loophole is unlikely to survive a follow-up case where training CAN be shown to have happened in the UK. (3) Implement guardrails to prevent infringing outputs — the Bartz ruling leans in favor of fair use when a model has effective technical guardrails. (4) Don’t train on LibGen, PiLiMi, or any other shadow library — this is now per se evidence of piracy that defeats fair use regardless of how transformative the use is. The technical side of on-device AI and small language models may eventually offer some relief here — smaller models trained on curated data shrink training-stage copyright exposure.
The practical playbook for developers shipping AI output: (1) If your output is purely AI-generated with no human creative contribution, you do not own it in the U.S. — the work is in the public domain. (2) If you meaningfully prompt, iterate, select, and edit, you likely own copyright in the U.S., China, and most civil-law jurisdictions. (3) Label AI involvement — China requires it explicitly under Li v. Liu; the EU AI Act’s Article 50 transparency rules require AI-generated content to be labeled for users; California’s AB 2013 requires disclosure of AI-generated content in political and commercial contexts. Voluntary labeling is now best practice globally.
What to watch over the next 12 months
Five concrete inflection points will move the needle on AI copyright between now and September 2027:
- The NYT v. OpenAI summary judgment ruling. Judge Stein’s decision is expected by mid-2027. A defense verdict would settle U.S. fair use for training in the short term; a plaintiff verdict would force the entire industry toward licensed data.
- The EU CJEU ruling on the Hungarian TDM referral. Expected in 2027. The ruling will determine whether AI training under the TDM exception is permissible without rightsholder opt-out — and may force a renegotiation of Article 53’s interaction with the DSM Directive.
- The parallel Getty Images v. Stability AI case in California. Getty voluntarily dismissed and refiled in California in August 2025. Unlike the UK case, training DID happen in the U.S. The ruling will be determinative on U.S. secondary infringement of AI models.
- Chinese training-stage doctrine. The Beijing Internet Court or another specialized IP court will eventually rule on whether AI training itself is infringement.
- The voluntary licensing market. If the U.S. Copyright Office’s preferred outcome — voluntary collective licensing — succeeds, the market will formalize around CMOs for AI training rights. If it fails, legislative pressure for a compulsory regime will grow.
The single most important thing to understand about AI copyright in 2026 is that there is no single answer. The legal landscape is a patchwork of fair-use tests, TDM exceptions, statutory carve-outs, and emerging doctrines — and the patch is being stitched in real time by courts, regulators, and legislatures on four continents. The right move for anyone operating in this space is to track the rule in every jurisdiction you serve, document your data provenance and creative process, and assume that today’s permissive posture is tomorrow’s litigation target. The licensing pressure is already pushing the industry toward synthetic-data substitutes and curated datasets as a way out of the legal exposure — but that is its own can of worms.
Frequently asked questions
Who owns content generated by AI in the United States?
Under the U.S. Copyright Act, purely AI-generated works without sufficient human creative contribution cannot be copyrighted at all — the work is in the public domain. The D.C. Circuit affirmed this in Thaler v. Perlmutter (March 2025), and the U.S. Supreme Court denied certiorari in March 2026, leaving the rule as final settled law. Works that combine AI-generated elements with meaningful human authorship — for example, human-authored text with AI-generated illustrations, or human-directed selection and arrangement of AI outputs — CAN be registered and owned, following the Copyright Office’s 2023 Registration Guidance and the “Zarya of the Dawn” precedent.
Is it legal to train an AI model on copyrighted material without permission?
The answer depends on the jurisdiction. In the U.S., the question is unsettled and turns on fair-use analysis — transformative purpose and market effects are the two factors judges weight most heavily, per the U.S. Copyright Office’s May 2025 Part 3 report. Training on legally acquired works is currently defensible; training on pirated datasets is “almost certainly” not fair use. In the EU, Article 53 of the AI Act (in force since August 2025, enforceable from August 2026) requires GPAI providers to publish a training-data summary, maintain a copyright policy, and respect rightsholder opt-outs. In Japan, Article 30-4 of the Copyright Act permits training for “non-enjoyment” purposes unless it would “unreasonably prejudice” the copyright owner’s interests. In the UK, the High Court’s November 2025 ruling in Getty v. Stability AI left the question unresolved because training happened outside the UK. China has no binding ruling on training-stage infringement as of mid-2026.
What is EU AI Act Article 53?
Article 53 of the EU AI Act imposes four obligations on providers of general-purpose AI (GPAI) models — that is, foundation models placed on the EU market. Providers must: (1) maintain technical documentation of the model, including training and testing processes; (2) provide information to downstream providers integrating the model; (3) implement a copyright policy that complies with EU law, particularly the Article 4 text-and-data-mining exception of the DSM Copyright Directive; (4) publish a “sufficiently detailed summary” of the content used for training the model. The summary is public and must follow the European AI Office’s Training Data Summary Template. Non-compliance can trigger fines up to €15 million or 3% of global turnover. Article 53 took effect on August 2, 2025, and has been enforceable since August 2, 2026.
What was the Anthropic $1.5 billion settlement about?
The August 2025 settlement in Bartz v. Anthropic resolved a class-action lawsuit alleging that Anthropic downloaded over 7 million books from shadow libraries LibGen and Pirate Library Mirror (PiLiMi) to train its Claude models. Judge William Alsup (N.D. Cal.) had ruled in June 2025 that Anthropic’s use of LEGALLY acquired books for training was “quintessentially transformative” fair use, but its downloading of pirated copies was NOT fair use. The class was certified to include every beneficial or legal copyright owner of books in those datasets — approximately 482,460 books after filtering. The $1.5 billion settlement is the largest copyright settlement in U.S. history, and Anthropic agreed to destroy the unlawfully obtained files.
How can creators opt out of having their work used for AI training?
Three mechanisms are in active use as of 2026. First, AI-specific robots.txt directives and the TDMRep metadata standard, both designed to express a rightsholder reservation under the EU DSM Directive’s Article 4(3) opt-out regime. Second, direct registration with major training-data licensing platforms — services like Spawning AI’s “Have I Been Trained?” let creators discover and request removal from common datasets. Third, copyright registration with the U.S. Copyright Office, which establishes the legal record needed to bring an infringement claim and unlocks statutory damages. Under EU law, opt-outs are enforceable against AI providers operating in the EU; under U.S. law, opt-outs are evidence in a fair-use analysis but do not categorically defeat it; under Japanese law, an opt-out can trigger the “unreasonable prejudice” proviso that removes Article 30-4’s protection.
This article cites 13 primary sources — court rulings, regulator reports, and peer-reviewed legal analysis — verified accessible as of September 7, 2026. For practitioners operating across borders, the recommended baseline is documentation of training-data provenance, output creative process, and jurisdiction-specific opt-out compliance.