On September 1, 2026, the United States filed a Statement of Interest in the consolidated OpenAI copyright litigation in the Southern District of New York, urging the court to reject arguments that training LLMs on copyrighted texts violates copyright law.
That is a formal executive-branch litigation position, not a ruling, and it leaves acquisition and output questions aside.
It is also more aggressive than a casual policy remark and narrower than the bluntest headlines make it sound.
The government's filing does not say every use of copyrighted material by an AI company is lawful. It separates the pipeline into acquisition, training, and output, then says the United States is focused on whether the use of copyrighted works at the training stage constitutes fair use.
That distinction is doing real work.
What The Government Actually Asked The Court To Do
The filing is a Statement of Interest under 28 U.S.C. § 517, not a merits ruling and not a government complaint against anyone. It is the executive branch telling Judge Sidney Stein how it thinks federal copyright law should be applied in this litigation.
Its central ask is hard to miss. The brief says the United States has a strong interest in the court rejecting any argument that training LLMs on copyrighted texts violates copyright law.
The government then frames LLM development as a staged process:
- acquisition of data;
- model training; and
- model outputs.
The brief says each stage may present distinct copyright questions. But the United States limits its own position to the training stage, defined as copying works in order to feed data into the model as learning material.
That means the filing is best understood as an effort to establish a strong presumption for training-stage fair use while leaving acquisition and output practices for separate analysis.
Why This Is More Than A Repackaged Talking Point
The administration had already stated in its March 2026 National Policy Framework for Artificial Intelligence that training an LLM on copyrighted material, "in and of itself," does not violate copyright law. A July 2026 DOJ journal article also addressed the doctrine, but it expressly disclaimed that the authors' views necessarily reflected DOJ policy.
This filing is different.
It is a formal Department of Justice submission in pending federal litigation. It was submitted under a signature block listing Associate Attorney General Stanley Woodward Jr. and Civil Division Assistant Attorney General Brett Shumate, and signed by Senior Counsel Michael Weisbuch. That does not make it controlling law, but it does make it a real statement of the executive branch's litigation position in this case.
So the legal significance is not that the issue is now settled. It is that a federal court now has before it an express government argument that OpenAI's training use is fair use and that courts should reject a rule generally making LLM training impermissible without licensing.
The Fair-Use Theory The Government Is Pushing
The brief leans heavily on the idea that copyright law must distinguish among uses rather than flatten every copy into the same category.
That is why it emphasizes a use-by-use analysis and relies on cases like Authors Guild v. Google and Google v. Oracle America. The basic argument is familiar but now stated at full executive-branch volume: copying during training is transformative because it uses works as learning material to develop a model that recognizes patterns and relationships rather than to distribute the original works as a substitute library.
The filing's language gets especially strong in its fourth-factor discussion. It says OpenAI's model training using New York Times articles is fair use and argues that the training copy does not serve as a substitute for the original or shrink protected market opportunities in the relevant copyright sense.
That is a more aggressive proposition than merely saying the law is unsettled.
The brief also frames a contrary rule as harmful to innovation, scientific progress, and U.S. competitiveness. It argues that broad copyright liability for training would hamper AI development under a misunderstanding of fair use doctrine.
The Filing's Sharpest Move Is Its Attack On Kadrey
The most consequential doctrinal section may be the government's critique of Kadrey v. Meta Platforms.
The filing says the Kadrey court, "without the benefit of briefing," adopted an "indirect substitution" theory of "market dilution" based on the possibility that LLMs might create books competing with human-authored books. The government argues that this approach is untethered from the core copyright inquiry because it conflates training copies, which are not publicly accessible, with outputs, which may be accessible but will often lack substantial similarity.
That matters because the fourth fair-use factor has become one of the main battlegrounds in AI copyright cases.
If courts accept a broad market-dilution theory, many AI developers could face a much harder path on fair use even when their models are not outputting close substitutes for specific source works. If courts instead require a tighter connection between the challenged use and legally cognizable substitution, training-stage defendants get much more room.
So the filing is not just defending OpenAI's posture in one case. It is trying to narrow one of the strongest emerging theories plaintiffs have used against AI training.
The Government's Policy Case Is Broader Than Doctrine
The filing also makes an overt policy argument about structure and competition.
It says an erroneous ruling against fair use would hamper competition in the LLM market because only the largest technology companies might have the capital to pay universal licensing fees. It also says such fees would disproportionately benefit legacy media outlets because of the sheer volume of their written publications.
Then the filing makes the point in blunter terms: it is not in the public interest for the largest technology companies to have an oligopoly on LLM training due to licensing entry barriers that function primarily as subsidies for old mainstream media companies.
That argument will be attractive to many AI companies and many policymakers who worry about incumbency barriers.
It is also not the whole story.
In a footnote, the government says it takes no position on whether a licensing regime would actually be financially or logistically feasible. So the filing is clearly hostile to broad mandatory licensing as a legal consequence of the case, but it stops short of claiming to have solved the practical licensing debate.
What The Filing Does Not Resolve
This is where readers should slow down.
The government did not ask the court to bless pirated acquisition or infringing outputs. The filing expressly separates those questions from the training-stage issue it wants resolved.
That means several high-stakes questions remain open even if the court finds the filing persuasive:
- how the copyrighted works were obtained;
- whether retained libraries or other collection-stage conduct raises separate infringement problems;
- whether RAG or output behavior can create substitution or substantial-similarity issues;
- whether training on some categories of works presents different market-harm facts from others; and
- whether courts will accept or reject broader theories of lost licensing markets.
Those limits matter even more because the U.S. Copyright Office's May 2025 prepublication Part 3 report takes a more qualified approach.
The Office says various uses of copyrighted works in AI training are likely to be transformative. But it immediately adds that the result depends on what works were used, from what source, for what purpose, and with what controls on the outputs. It also warns that making commercial use of vast troves of copyrighted works to produce expressive content that competes with them in existing markets, especially through illegal access, can go beyond established fair-use boundaries.
That report predates this filing. DOJ addresses it directly in footnote 17, arguing that similar market-dilution reasoning deserves no deference and overlooks the required use-by-use analysis.
That is not a direct rejection of DOJ's position. But it is a reminder that the government has advanced one strong reading of fair use, not the only plausible one, and the filing itself says the fair-use inquiry still turns on the specific facts and uses at issue in each case.
Why This Matters Beyond OpenAI
This filing lands in the middle of a larger litigation map that includes authors, publishers, answer-engine disputes, training-data fights, and increasingly explicit market-substitution theories.
That broader setting matters because the filing is trying to influence not just one factual record but the legal frame that future courts may use.
If Judge Stein embraces the government's approach, the training-stage fair-use defense gets a much firmer doctrinal platform in one of the most important AI copyright proceedings in the country. If he does not, the executive branch will still have shown the argument it wants courts to take seriously: separate training from acquisition and outputs, reject generalized market-dilution theories, and treat transformative training use as consistent with copyright's constitutional purpose.
Either way, the filing gives courts and litigants a cleaner version of the pro-training position than they had before.
Bottom Line
DOJ did not tell a court that everything about AI companies' use of copyrighted material is lawful.
It did something narrower and more important.
It told the Southern District of New York that OpenAI's training use is fair use and that training LLMs on copyrighted texts should not, in general, be treated as copyright infringement. At the same time, it left acquisition and output questions for separate analysis and did not claim every training use in every case will necessarily come out the same way.
The fight now is over whether courts will accept that carveout, and how sharply they will separate training from the acquisition and output questions DOJ left unresolved.
Sources and Related Clearon Coverage
- Statement of Interest of the United States, In re OpenAI, Inc. Copyright Infringement Litigation, 25-md-3143 (S.D.N.Y. filed Sept. 1, 2026)
- U.S. Copyright Office, Copyright and Artificial Intelligence, Part 3: Generative AI Training (pre-publication version)
- DOJ Journal of Federal Law and Practice, "Artificial Intelligence and the Fractured Intellectual Property Landscape"
- Agenccy summary of the filing (non-authoritative reporting)
- clearon-ai.com
- clearon-ai.com

Leave a Reply