The U.S. Justice Department has formally told a federal judge that copying copyrighted written works to train large language models should be treated as fair use. But DOJ has not changed copyright law by itself, has not declared pirated datasets legal, and has not said that outputs reproducing copyrighted material are immune from infringement claims.
The September 1, 2026 filing draws a much more important legal boundary. DOJ treats acquiring source material, copying it during model training, and generating outputs for users as distinct uses that can raise different copyright questions. Its brief focuses on the middle step and calls the use of copies for LLM training “extraordinarily transformative.”
There is also a deeper disagreement hidden beneath the simpler “government sides with OpenAI” headline. DOJ argues that authors generally cannot establish copyright market harm merely because AI generates new works in the same genre or category that compete for customers without reproducing protected expression. The U.S. Copyright Office and Judge Vince Chhabria in Kadrey v. Meta have left considerably more room for such a market-dilution theory.
That dispute over what kind of market harm copyright law recognizes may ultimately be more consequential than the fight over whether training is transformative.
| The shorthand claim | What the evidence actually supports |
|---|---|
| “The government legalized AI training.” | No. DOJ submitted a legal position. Judge Sidney Stein decides the case. |
| “AI companies can use any copyrighted material they want.” | Too broad. The brief concerns LLM training on written works and separates training from acquisition. |
| “Pirated training data is now fair use.” | No. DOJ does not resolve unlawful acquisition. |
| “If training is fair use, the outputs cannot infringe.” | False. DOJ expressly treats outputs as a separate copyright question. |
| “AI companies would never need licenses.” | False. DOJ itself says licensing can provide specialized access to real-time, paywalled and proprietary content. |
| “Courts have already settled generative-AI training.” | No. Important district courts have reached fact-specific results, while the first major AI-training appellate fight is still pending. |
What exactly did the Justice Department file?
On September 1, the United States filed a Statement of Interest in In re OpenAI, Inc. Copyright Infringement Litigation, the multidistrict litigation pending in the Southern District of New York.
The court-stamped document is Document 316 on a member docket and is captioned as relating to “All Matters” in MDL No. 25-md-3143. DOJ says it is appearing under 28 U.S.C. §517, which authorizes Justice Department officers to attend to the interests of the United States in pending litigation.
That procedural posture matters.
The United States did not join the litigation as a plaintiff or defendant. It has given Judge Stein the federal government’s interpretation of the law; it has not issued a binding regulation or judicial ruling.
So headlines saying the government has “ruled” AI training legal would be incorrect.
The administration had already taken this position before September
September 1 was also not the first time the Trump administration publicly said copyrighted material can be used to train AI without violating copyright law.
The White House’s March 2026 National Policy Framework for Artificial Intelligence had already stated that the administration believes training AI models on copyrighted material does not violate copyright law. Importantly, the same document acknowledged that contrary arguments exist and said courts should be allowed to resolve whether such training qualifies as fair use. It also said Congress could facilitate collective licensing systems without legislating when licensing is legally required. (The White House)
The September development is therefore more precise:
The administration has now carried that policy position into one of the country’s central AI copyright cases and supplied a detailed fair-use argument for a federal judge to adopt.
That is different from merely issuing a policy document.
DOJ’s position is sweeping—but not absolute
Fair use is governed by 17 U.S.C. §107, which directs courts to consider the purpose and character of a use, the nature of the copyrighted work, the amount used, and the effect on the potential market for or value of the copyrighted work. It is a contextual test rather than an automatic exemption for any particular technology.
DOJ nevertheless takes a strong position on LLM training.
The brief argues that training copyrighted text is highly transformative because the training copies serve a different purpose from the original books and articles: they are processed to help develop a model capable of learning linguistic relationships and generating responses, rather than supplied to readers as substitute copies of the originals. DOJ concludes that using copies for LLM training is “extraordinarily transformative.”
It also argues that requiring licenses as a general condition of LLM training would be legally wrong.
But DOJ’s conclusion contains a caveat that should not disappear from the headline: the brief expressly says the fair-use inquiry depends on the specific facts and uses involved in each case.
The fairest summary is therefore:
DOJ is advocating a broad pro-fair-use rule for the training-stage copying of written works. It is not proposing blanket immunity for every act involved in building or operating an AI system.
The key distinction: acquisition, training and output are different uses
This is the most useful way to understand the filing.
DOJ adopts a three-stage description of LLM development already used by Judge Stein:
| Stage | What happens | What DOJ argues | What can remain disputed |
|---|---|---|---|
| Acquisition / collection | A developer obtains and stores books, articles or other material | Not the question DOJ asks the court to decide here | Piracy, unauthorized acquisition, access restrictions, contracts and other claims |
| Training | Copies of works are processed as learning material for the model | DOJ argues this use strongly favors fair use | Fair use still must be applied to the particular case |
| Output / deployment | The model responds to users | A distinct use requiring separate analysis | Reproduction, substantial similarity, market substitution and other output claims |
The DOJ brief expressly says that each stage can present distinct copyright questions and that the United States is focusing on copies made at the training stage.
That distinction also prevents a common analytical mistake: treating every form of “AI use” as one continuous act.
A company could potentially have a strong fair-use defense for a particular training copy but face a different claim over how it obtained its source copy. And a model trained through fair use could still produce a particular output that raises an infringement claim.
Retrieval-augmented generation and other systems that fetch source material after training can introduce still other uses. The Copyright Office likewise treats initial pre-training, subsequent training and retrieval uses as potentially distinct. (U.S. Copyright Office)
No, the DOJ filing does not legalize piracy
The strongest example comes from Bartz v. Anthropic.
In June 2025, Judge William Alsup ruled that reproductions of the plaintiff authors’ books made specifically in the process of training Anthropic’s LLMs qualified as fair use. But he separately analyzed millions of digital books Anthropic had downloaded from pirate libraries and retained in a central research library. He denied Anthropic fair-use protection for that separate pirate-library use on the summary-judgment record. (Justia Dockets & Filings)
There is an important procedural qualification here.
That ruling did not amount to a final post-trial judgment that every pirate download constituted infringement. Anthropic was the party seeking summary judgment, so factual disputes on that portion of the case were viewed in the authors’ favor. A later court opinion summarizing the ruling specifically noted that the authors had a strong piracy case but that success at trial was not guaranteed. (Justia Law)
What Bartz establishes for present purposes is narrower and extremely useful:
A court can treat the training use as fair while treating the acquisition and retention of the source copies as a separate legal problem.
The Copyright Office’s May 2025 pre-publication report takes a related position. In the Office’s view, knowingly using a dataset composed of pirated or illegally accessed works should weigh against fair use, though it should not automatically determine the result. (U.S. Copyright Office)
“Scraping” should therefore not be used as a synonym for piracy. Downloading a freely accessible webpage, obtaining material through a commercial license, circumventing access controls, copying a subscription database and downloading books from a pirate repository present materially different facts.
Fair-use training does not immunize infringing outputs
DOJ also explicitly separates model training from what the resulting system gives users.
The brief recognizes that an output use may not be transformative if an LLM reconstructs and disseminates a copyrighted work. DOJ’s position is that such outputs should be evaluated separately rather than used automatically to turn the underlying training process into infringement.
The OpenAI MDL already contains a useful example of the distinction.
In October 2025, Judge Stein denied OpenAI’s motion to dismiss an output-based direct-infringement claim brought by the consolidated author plaintiffs. He ruled that the allegations, taken as required at that procedural stage, stated a prima facie infringement claim as to at least some alleged ChatGPT outputs. (Justia Dockets & Filings)
That was not a judicial finding that those outputs actually infringe. It simply allowed that claim to proceed.
The important answer is much simpler:
Yes, an AI training use can potentially be fair while a particular output from the resulting model still infringes.
The biggest unresolved fight may be over market harm
Analysis: Much of the AI copyright debate focuses on the first fair-use factor and the question of whether LLM training is transformative. The DOJ filing may be even more consequential on the fourth factor: what kinds of economic harm copyright law actually recognizes.
Consider a hypothetical.
An AI system trains on thousands of human-written romance novels. It does not reproduce passages from any particular book or generate a substantially similar copy. But users can now produce huge numbers of new romance novels cheaply, and those books compete with human authors for readers.
Does that competition count against fair use?
DOJ: ordinary competition is not enough
DOJ says generally no.
Its argument is that copyright protects an author’s particular expression, not the author’s freedom from competition in a genre or category. If an AI output does not reproduce or substantially resemble protected aspects of the original, DOJ says competition from that output is not the kind of substitutive market harm that should count against the earlier training use.
That is why DOJ directly attacks the “market dilution” reasoning in Judge Vince Chhabria’s Kadrey v. Meta decision.
The Copyright Office sees a broader potential harm
The Copyright Office’s May 2025 report takes a materially different view.
It acknowledges that this is “uncharted territory” but says the speed and scale of generative AI could dilute markets for works of the same kind as those used in training. Its example is thousands of AI-generated romance novels entering a market and making it harder for human-authored novels to sell or even be found. (U.S. Copyright Office)
The Office does not say every such competitive effect defeats fair use. Its broader conclusion is expressly case-specific: some generative-AI training uses will be fair and some will not. At one end of its spectrum, noncommercial research that does not reproduce source works is likely fair; at the other, it says copying expressive works from pirate sources to generate unrestricted competing content where licensing is reasonably available is unlikely to be fair. (U.S. Copyright Office)
As of September 4, 2026, the Copyright Office still lists Part 3 as a pre-publication version and says a final version will be published later without substantive changes expected to its analysis or conclusions. (U.S. Copyright Office)
What Kadrey v. Meta actually held
Kadrey is frequently summarized too broadly.
Judge Chhabria granted Meta summary judgment on the thirteen plaintiff authors’ claim that using their books to train Llama infringed copyright. But his opinion explicitly warned readers not to turn that result into a general rule that Meta’s use of copyrighted books for LLM training was lawful.
The reason Meta won was evidentiary.
Chhabria said the authors’ theories based on regurgitation and a lost licensing market failed on the record before him. He described market dilution from a flood of competing AI-generated works as a potentially stronger theory, but the plaintiffs had barely developed it and had not presented enough evidence to create a genuine dispute of material fact. (Justia Law)
That produces a considerably different takeaway from “Meta won, therefore AI training is fair use”:
Meta won this fair-use dispute on this record. Judge Chhabria simultaneously suggested that a better-supported market-dilution case could produce a different result.
DOJ now argues that the latter part of Kadrey is legally wrong.
DOJ and the Copyright Office disagree—but not quite the way DOJ says
There is a particularly important correction buried in the primary documents.
DOJ faults the Copyright Office’s market-dilution analysis for ignoring case law requiring a use-by-use fair-use analysis.
But the Copyright Office’s report expressly says:
different uses during AI development and deployment require “separate consideration.”
So DOJ’s criticism overstates the disagreement if it is read to mean that the Copyright Office simply treats training and deployment as one use.
It does not.
The actual disagreement is more precise:
| Issue | DOJ, September 2026 | Copyright Office, May 2025 |
|---|---|---|
| Different AI uses require separate analysis | Yes | Yes |
| Is LLM training transformative? | DOJ says extraordinarily so | Various training uses are likely transformative; fair use remains fact-specific |
| Does unlawful source acquisition matter? | Acquisition is bracketed from DOJ’s training argument | Knowing use of pirated/illegally accessed datasets can weigh against fair use |
| Can downstream outputs affect the training analysis? | DOJ argues strongly that output questions should not bear on the training use | Output controls, intended uses and market consequences can affect the analysis |
| Can non-infringing AI competition dilute the market? | DOJ generally rejects genre/category-level competitive harm | The Office says large-scale market dilution can potentially matter |
| Can a licensing market matter? | DOJ rejects broad liability that would generally require training licenses | Existing or reasonably developing licensing markets can weigh in the fair-use balance |
| Is either position binding on the court? | No | No |
The real dispute is therefore how insulated the training-stage analysis should be from the economic consequences of the resulting generative system.
That is a substantially harder question than simply asking whether training and outputs are “different.”
DOJ’s training-output firewall has a tension of its own
Analysis: DOJ’s use-by-use approach is powerful, but its strongest formulation also raises an unresolved logical question.
The government says whatever copyright issues particular outputs may raise should not bear on the transformative nature or any other aspect of the training use.
But factor four requires courts to assess market effects.
If evidence showing that a trained model does not expose meaningful portions of source works supports the conclusion that the training use is non-substitutive, it is not self-evident that evidence pointing strongly in the opposite direction must always be irrelevant.
Training and outputs can be legally distinct uses without being factually unrelated.
The Supreme Court’s Andy Warhol Foundation v. Goldsmith decision reinforces the need to identify the specific challenged use rather than treating every downstream exploitation as identical. But it does not create an express evidentiary rule forbidding courts from considering downstream consequences when evaluating a particular use’s purpose or market effects. (Supreme Court)
That does not make DOJ’s position wrong. It identifies a question Judge Stein may eventually have to resolve: how separate is separate?
What courts have actually ruled about AI training and analogous copying
The law is developing, but it is not a blank slate.
| Case | Court | What it actually establishes |
|---|---|---|
| Authors Guild v. Google | Second Circuit, 2015 | Google’s copying of complete books to enable search and analytical functions was highly transformative fair use. It is binding Second Circuit precedent and a major analogy for DOJ, but Google Books was not a generative LLM. (GovInfo) |
| Andy Warhol Foundation v. Goldsmith | U.S. Supreme Court, 2023 | Fair use must focus on the particular challenged use, and a claimed new meaning alone does not automatically make a commercial use fair. (Supreme Court) |
| Bartz v. Anthropic | N.D. California, 2025 | Copies made specifically for LLM training were held fair use; Anthropic’s separate pirate-library use was not granted fair-use protection on the summary-judgment record. (Justia Dockets & Filings) |
| Kadrey v. Meta | N.D. California, 2025 | Meta won fair use for these authors on the record presented, while the judge left open a potentially stronger market-dilution theory. (Justia Law) |
| Thomson Reuters v. ROSS Intelligence | D. Delaware, 2025 | Fair use was rejected where copyrighted Westlaw headnotes were used to help build a directly competing legal-research product. The court specifically emphasized that ROSS was not generative AI. (Justia Law) |
That leaves an important limit on claims that “the courts have decided AI training is fair use.”
No federal appellate court has yet issued a merits decision squarely deciding fair use for general-purpose generative-LLM training on copyrighted expressive works.
The ROSS case is now before the Third Circuit, which heard oral argument on June 11, 2026. The appellate case itself concerns a non-generative legal-research system, so even that forthcoming ruling may not cleanly resolve the OpenAI question. (Third Circuit Court)
Why Thomson Reuters v. ROSS matters
ROSS is the clearest recent warning against treating the words “AI training” as a magic fair-use category.
ROSS used material derived from Westlaw’s copyrighted headnotes in developing a commercial legal-research system that competed with Westlaw. Judge Stephanos Bibas found the use commercial and non-transformative and held that the important first and fourth factors favored Thomson Reuters. (Justia Law)
The differences from OpenAI are substantial.
ROSS was not a general-purpose generative system. The court said it returned already-written judicial opinions rather than generating new content, and Thomson Reuters and ROSS competed in the same legal-research market.
So ROSS does not establish that LLM training is infringement.
What it establishes is narrower:
A court will not necessarily treat intermediate copying as fair merely because copyrighted material was used during development of an AI product. Purpose, market relationship and the particular use still matter.
If training is fair use, why are AI companies still licensing content?
Because fair use and licensing are not mutually exclusive.
A publisher can sell an AI developer more than permission to make a training copy. A deal can provide:
- structured access to archives;
- real-time or continuously updated material;
- paywalled content;
- APIs or proprietary feeds;
- retrieval and display rights;
- attribution arrangements;
- contractual certainty;
- rights covering uses beyond initial model training.
DOJ makes this point itself. Footnote 13 says that regardless of whether LLM training qualifies as fair use, mainstream and independent publishers can enter—and have entered—agreements giving developers specialized access to real-time, paywalled, proprietary and other content. DOJ simultaneously says it takes no position on exactly how financially or logistically feasible a broad licensing regime would be.
The March White House framework similarly asks Congress to consider mechanisms for collective licensing while expressly declining to legislate when such licensing is actually required. (The White House)
So neither of these claims follows:
“OpenAI licenses content, therefore OpenAI admits training requires a license.”
or
“Training is fair use, therefore content licenses have no value.”
Both are too simplistic.
DOJ’s competition and national-security claims are arguments, not findings
The government also argues that broadly requiring licenses for LLM training could make advanced model development harder for smaller companies, advantage the largest firms capable of paying, slow innovation and put U.S. developers at a disadvantage against foreign competitors.
DOJ connects those concerns to economic competitiveness and national security.
Those claims explain the government’s interest in the litigation, but they should not be presented as findings by the court.
Judge Stein has not determined that a licensing regime would create an AI oligopoly or harm national security. Indeed, DOJ’s own footnote acknowledges that the United States is not deciding precisely whether a licensing system would be financially or logistically workable.
The distinction is straightforward:
Verified fact: DOJ has made those arguments.
Not yet established: that those predicted consequences would actually occur at the scale DOJ suggests.
Does DOJ’s brief decide copyright law for images, music or every other AI model?
No.
The filing repeatedly frames its interest around training LLMs on written works and copyrighted text in the OpenAI litigation. Its reasoning about transformative intermediate copying and market substitution could certainly be cited in cases involving other media, but this brief does not itself resolve image-generation, music-generation or every other AI-training context.
That is why “DOJ says all AI training is fair use” is broader than the actual litigation position developed here.
What happens next in the OpenAI copyright litigation?
Timing explains why DOJ filed when it did.
Judge Stein’s March 23 scheduling order set September 4, 2026 as the deadline for summary-judgment motions, October 9 for oppositions and November 6 for replies. (Justia Dockets & Filings)
The DOJ brief therefore arrived three days before the summary-judgment deadline.
That stage of the litigation is important because the fair-use dispute will increasingly turn on a developed evidentiary record rather than allegations alone: how source works were acquired, which copies were used for which purposes, what outputs the systems can actually generate, how often reconstructive outputs occur, what safeguards exist, and what evidence the parties can produce about licensing and market substitution.
Because this article is being finalized on September 4 while merits filings are arriving and public versions may still be affected by sealing or redaction, new statistics asserted by either side should be treated as party evidence—not as established facts—until the underlying filings are public and the court evaluates them.
What happens if DOJ’s legal theory wins?
Reasonable inference: AI copyright litigation would probably not disappear. The center of gravity would move.
If courts conclude that making training copies of lawfully obtained written works is generally fair use, plaintiffs would have stronger incentives to focus on questions surrounding the training step rather than simply proving that their work appeared in a corpus.
Future litigation could concentrate more heavily on:
- how source copies were obtained;
- pirate libraries and unlawful access;
- contracts and access restrictions;
- model outputs that reproduce protected expression;
- retrieval or grounding systems that deliver source material after training;
- licensing rights covering uses other than bare training;
- distribution or retention of source libraries.
The practical question could shift from:
“Was my work in the training data?”
toward:
“How did you obtain it, what did you do with the copy, and what can the finished system reproduce or substitute for?”
That is an inference about the likely direction of litigation, not something DOJ has already established as law.
The bottom line
The simplest accurate version of the September 1 filing is this:
DOJ wants Judge Stein to treat copying copyrighted written works for LLM training as a strongly transformative fair use. It does not ask the court to declare every method of obtaining those works lawful, and it does not give infringing model outputs a free pass.
The filing is significant because the federal government has now moved beyond a White House policy statement and put a detailed copyright theory directly before the court hearing the consolidated OpenAI cases.
But the most consequential dispute may not be whether training is transformative.
It may be whether authors can invoke the fourth fair-use factor when AI systems trained on their work generate enormous amounts of different, non-infringing expression that nevertheless competes with them economically.
DOJ says copyright generally does not protect creators against that kind of generalized competition.
The Copyright Office says generative AI’s unprecedented speed and scale may make such market dilution legally relevant.
Judge Chhabria has said the theory could matter but found the plaintiffs before him had not proved it.
Judge Stein now has to decide how those principles apply to the much larger evidentiary record in the OpenAI litigation.
That question—not the slogan “AI training is fair use”—is where a large part of the next phase of U.S. AI copyright law is likely to be fought.
References and Further Reading
Primary DOJ filing and federal law
Statement of Interest of the United States — In re OpenAI Copyright Infringement Litigation, Document 316, Sept. 1, 2026 Court-stamped copy of the DOJ filing. This is the primary source for the government’s fair-use argument, its acquisition/training/output distinction, its criticism of market dilution and its competition and national-security arguments.
28 U.S.C. §517 — Interests of United States in Pending Suits The statutory authority DOJ invoked to present the interests of the United States in the pending litigation without joining the case as a party.
17 U.S.C. §107 — Fair Use The federal statute setting out the four fair-use factors.
White House and Copyright Office
White House National Policy Framework for Artificial Intelligence — Legislative Recommendations, March 2026 Shows that the administration publicly adopted its pro-training copyright position before DOJ’s September filing while still saying courts should resolve fair use and Congress should not dictate when licensing is required.
U.S. Copyright Office — Copyright and Artificial Intelligence, Part 3: Generative AI Training, Pre-Publication Version The Office’s detailed May 2025 analysis of AI training, fair use, source legality, licensing and market dilution.
U.S. Copyright Office AI Study — Current Report Status The Office’s current landing page, which as of September 4, 2026 still identifies Part 3 as a pre-publication report.
Controlling and influential cases
Authors Guild v. Google — U.S. Court of Appeals for the Second Circuit Binding Second Circuit precedent finding Google’s complete copying of books for transformative search and analytical purposes to be fair use.
Andy Warhol Foundation for the Visual Arts v. Goldsmith — U.S. Supreme Court The Supreme Court’s modern treatment of the first fair-use factor and the requirement to focus on the particular challenged use.
Bartz v. Anthropic — Order on Fair Use, June 23, 2025 Separately analyzes LLM training copies and Anthropic’s central library, providing the clearest judicial example of why training and acquisition should not automatically be treated as the same copyright use.
Kadrey v. Meta Platforms — Order on Partial Summary Judgment, June 25, 2025 Granted Meta summary judgment on the record presented while developing the competing theory that large-scale generative output could cause market dilution.
Thomson Reuters Enterprise Centre v. ROSS Intelligence — Memorandum Opinion, Feb. 11, 2025 Rejected fair use for a non-generative AI legal-research product built using copyrighted Westlaw headnotes, illustrating why “AI training” is not itself a categorical fair-use exemption.
OpenAI litigation
October 2025 Opinion on ChatGPT Output-Based Copyright Claims Judge Stein’s opinion denying OpenAI’s motion to dismiss an output-based infringement claim because the author plaintiffs had adequately pleaded a claim as to at least some alleged outputs. It is a pleading ruling, not a final infringement finding.
OpenAI Copyright MDL Docket and March 23, 2026 Scheduling Order The docket record establishing the September 4 summary-judgment deadline and the scheduled October 9 opposition and November 6 reply deadlines.
Third Circuit Oral Argument Records — Thomson Reuters v. ROSS Intelligence, No. 25-2153 Confirms that the Third Circuit heard the ROSS appeal on June 11, 2026; no appellate merits ruling resolving that AI-training fair-use dispute had issued as of this article’s publication.
Editorial currency note: This article reflects the public record available on September 4, 2026. Summary-judgment briefing in the OpenAI multidistrict litigation is actively developing. Assertions contained in newly filed briefs and expert reports remain party positions unless and until adopted as factual or legal findings by the court. The Copyright Office’s Part 3 report also remains labeled pre-publication as of this date.



