OpenAI Found 14 Near-Verbatim Book Matches in 95 Million Outputs. Does That Prove AI Training Is Fair Use?

OpenAI says an expert found only 14 instances of verbatim or near-verbatim book regurgitation in a sample containing 95 million ChatGPT outputs. That is meaningful evidence—but it does not answer every copyright question in the case, especially how books were acquired, what copies were made during training, and what market harm legally counts.
A courtroom-style scene with law books, stacks of case files, a judge’s gavel, and a floating AI interface over a large digital text cube connected to bookshelves.
Contents

No. OpenAI’s extremely low observed rate of verbatim book regurgitation could be important evidence that ordinary ChatGPT use does not substitute for reading the plaintiffs’ books. But it does not, by itself, establish that OpenAI’s acquisition of copyrighted books, the copies made during model training, or every resulting use of those books qualifies as fair use.

That distinction has become unusually important because the sprawling OpenAI copyright litigation in federal court in New York has finally reached summary judgment. OpenAI, Microsoft, author plaintiffs and news publishers are now asking U.S. District Judge Sidney Stein to resolve major copyright questions on a developed evidentiary record rather than merely deciding whether allegations are sufficient to proceed.[1][2]

The most striking number in OpenAI’s September 4 filing is easy to repeat and just as easy to misunderstand.

OpenAI says its expert searched 95 million model outputs contained in 20 million ChatGPT conversation logs and identified only 14 instances of verbatim or near-verbatim regurgitation from the books at issue. OpenAI describes that as an alleged regurgitation rate of 0.00007%.[1:1]

But those numbers use two different denominators.

Fourteen divided by 20 million conversations is 0.00007%. Fourteen divided by 95 million outputs is about 0.000015%.

That does not necessarily mean OpenAI made a mathematical error. Elsewhere its brief explicitly describes 0.00007% as the percentage of a sample of 20 million chats containing such text.[1:2] It does mean, however, that “14 out of 95 million equals 0.00007%” is not an accurate description of the calculation.

More importantly, even a vanishingly small regurgitation rate answers only part of the legal question.

What OpenAI’s 14-in-95-million evidence actually says

According to OpenAI’s filing, its expert examined a sample containing 95 million outputs from 20 million produced ChatGPT conversation logs. He identified 14 instances containing verbatim or near-verbatim material from the books at issue.[1:3]

OpenAI says:

  • 12 of the 14 instances contained 32 words or fewer;
  • the other two contained 67 and 122 words;
  • none amounted to more than 0.017% of an asserted book; and
  • the author plaintiffs’ own review of a larger, partially overlapping dataset allegedly identified 31 instances in 108 million conversation logs, or about 0.00003%.[1:4]

Those are potentially powerful numbers for OpenAI.

But they are still litigant evidence, supported by expert work submitted in an adversarial case. Judge Stein has not adopted the methodology or made those numbers judicial findings.

The distinction matters.

Question Does the 14-instance evidence help answer it?
Does ordinary ChatGPT use frequently reproduce these books verbatim? Yes, substantially—if the methodology holds up.
Is ChatGPT ordinarily a practical substitute for obtaining the full books? Relevant evidence, and favorable to OpenAI.
Can targeted prompting sometimes extract copyrighted passages? No. That requires separate testing.
Did OpenAI copy books while training its models? No. Different question.
Did OpenAI use LibGen-derived books? No—but OpenAI separately acknowledges that it did.
Was obtaining books from LibGen legally protected fair use? No.
Were intermediate training copies fair use? No.
Can specific model outputs still infringe copyright? Not categorically.
Can AI-generated works economically compete with authors without copying passages verbatim? No.
Does an AI-training licensing market count as legally cognizable market harm? No.

The statistic is therefore important without being dispositive.

The LibGen issue is no longer merely an allegation

One of the most consequential facts in the new briefing comes from OpenAI itself.

OpenAI states that compilations of books from Library Genesis, or LibGen, were used in training GPT-3 in 2020 and GPT-3.5 in 2021. OpenAI says none of the other large language models at issue in the book case were trained on material downloaded from LibGen.[1:5]

That is different from saying plaintiffs merely accuse OpenAI of secretly training on LibGen.

The parties still sharply dispute what legal consequences follow from that acquisition, but the basic fact that LibGen-derived compilations were used for GPT-3 and GPT-3.5 training is acknowledged in OpenAI’s own summary-judgment brief.

The plaintiffs go considerably further. Their September 5 public brief alleges that OpenAI employees downloaded books from LibGen beginning years earlier, understood that the source was unlawful, obscured LibGen references by describing datasets as “Books1” and “Books2,” and later deleted its LibGen datasets because of legal concerns.[2:1]

Some of the underlying discovery record remains redacted.

That means the strongest defensible description is:

Verified from OpenAI’s own filing: LibGen book compilations were used in training GPT-3 and GPT-3.5.[1:6]

Plaintiffs’ allegation based on discovery: OpenAI knowingly sourced books from a pirate library, concealed that history and deleted the datasets because of legal concerns.[2:2]

Not yet established by Judge Stein: that those actions constituted copyright infringement, that concealment occurred for the purpose plaintiffs allege, or that the acquisition necessarily falls outside fair use.

That last distinction is the heart of the case.

Copyright law does not ask only one “AI training” question

Public discussion often compresses the dispute into a single sentence:

Is training AI on copyrighted books fair use?

The litigation is more complicated because several potentially infringing acts may need to be analyzed.

The Copyright Act directs courts to consider four factors, including the purpose of the use, the nature of the work, the amount used and the effect on the market for the copyrighted work.[3]

And the Supreme Court emphasized in Andy Warhol Foundation v. Goldsmith that fair use depends on the specific challenged use. The same underlying copying can receive different treatment when put to different purposes.[4]

That makes it useful to separate at least four questions.

1. Was acquiring the books from LibGen fair use?

The author plaintiffs say no.

They argue that downloading copyrighted books from LibGen was a distinct use: OpenAI obtained full digital books without buying or licensing them and created collections that could subsequently be used in model development.[2:3]

OpenAI rejects the attempt to isolate acquisition from what came next. Its brief argues that obtaining the material was part of a broader, transformative process whose purpose was training models to learn statistical patterns in language, rather than distributing the plaintiffs’ books to readers.[1:7]

The 14-regurgitation statistic does little to resolve this dispute.

Even if ChatGPT never reproduced a single paragraph, the court would still have to determine whether the earlier copying was protected.

The opposite is also true: demonstrating that some output reproduced a passage would not automatically establish that every earlier training copy was unlawful.

2. Were the copies made for model training fair use?

This is related to acquisition but not necessarily identical to it.

A 2025 ruling involving Anthropic illustrates why.

In Bartz v. Anthropic, Judge William Alsup concluded that Anthropic’s use of copyrighted books to train its language models was fair use. But he treated copies acquired from pirate libraries for Anthropic’s broader permanent library differently, refusing to excuse that acquisition simply because some books were eventually used for model training.[5]

That decision is not binding on Judge Stein. It came from a federal district court in California.

It is nonetheless important because it demonstrates that a court can plausibly reach different fair-use conclusions about different copies of the same underlying books.

OpenAI takes a different view of the acquisition question and relies in part on another 2025 decision, Kadrey v. Meta, where Judge Vince Chhabria granted Meta summary judgment on the training-related fair-use claims presented by those plaintiffs.[6]

Neither California decision settles the law in New York.

3. Do ChatGPT outputs substitute for the books?

This is where OpenAI’s low-regurgitation evidence becomes much more powerful.

The Second Circuit’s 2015 Authors Guild v. Google decision upheld Google Books’ scanning of millions of copyrighted books for search and limited snippet display. A central consideration was that the system did not provide a meaningful substitute for reading the protected books themselves.[7]

That case matters especially here because Judge Stein sits in the Second Circuit, where Google Books is binding appellate precedent. Bartz and Kadrey are persuasive authorities, not binding ones.

If OpenAI can convince Stein that ordinary ChatGPT use almost never delivers meaningful protected passages from the authors’ books, that strengthens the analogy to the non-substitutive features that helped Google.

But the analogy has limits.

Google acquired books through participating libraries and created a search index with restricted snippet displays. The current litigation includes disputed allegations about acquisition from LibGen, generative outputs and an entirely new alleged market effect: machines capable of producing new books.

Those differences matter.

Targeted extraction and ordinary use are measuring different things

The author litigation also contains evidence that sounds inconsistent with the tiny real-world regurgitation rate until the experiments are separated.

OpenAI says the plaintiffs’ expert submitted more than 5.3 million targeted prompts attempting to induce its models to reproduce text from the asserted books. OpenAI says the longest contiguous passage produced through those experiments was 1,899 words from A Game of Thrones, about 0.62% of that book.[1:8]

OpenAI characterizes this testing as “brute force” experimentation that frequently supplied portions of a book and asked models to continue them.

That characterization comes from OpenAI, an interested party. The underlying expert materials and methodology should be evaluated independently as more of the sealed record becomes public.

But the two headline findings are not inherently contradictory.

A model can:

  1. almost never reproduce meaningful copyrighted text during ordinary observed use; and
  2. still be capable of producing longer passages when deliberately subjected to millions of carefully engineered attempts.

One measures frequency in observed usage.

The other measures extractability under targeted attack.

Calling either one the universal “ChatGPT regurgitation rate” would lose that distinction.

4. Can AI harm authors without quoting their books?

This may become the hardest part of the case.

OpenAI’s strongest market argument is straightforward: if ChatGPT almost never gives users meaningful excerpts from the plaintiffs’ books, consumers are unlikely to use ChatGPT as a free replacement for buying and reading those books.

That is not the only market-harm theory the plaintiffs are pursuing.

They argue that training on copyrighted books allows AI systems to generate enormous quantities of new literary content that competes for the same readers and sales as human-authored works—even where the resulting AI books do not reproduce identifiable passages from any particular training book.[2:4]

The plaintiffs call this market dilution.

They rely heavily on reasoning in Kadrey. Although Judge Chhabria ultimately ruled for Meta on the record before him, he treated the possibility that generative AI could enable large-scale production of competing works as a serious copyright concern.[6:1]

The plaintiffs now claim they have developed evidence that was missing in Kadrey, including research and market evidence concerning AI-generated books.[2:5]

OpenAI says the theory is legally wrong.

Its position is that copyright protects an author’s expression, not a right to be free from competition by new, noninfringing works. If an AI-generated novel does not reproduce protected expression from a John Grisham novel, OpenAI argues that lost sales to the new novel are not copyright market harm simply because the model learned from Grisham’s writing.[1:9]

That dispute cannot be resolved by counting verbatim passages.

The licensing-market fight is separate again

The plaintiffs have another factor-four argument: OpenAI allegedly harmed an existing or developing market in which AI companies pay copyright holders to license material for model training.

Their September brief says the record identifies more than 120 publicly disclosed licenses between content providers and AI companies and industry spending exceeding $1 billion.[2:6]

Those figures are presented by the plaintiffs and are not judicial findings.

Their legal theory is that companies are already demonstrating a willingness to pay for training rights, so uncompensated copying deprives authors of a legitimate licensing market.

OpenAI says that reasoning is circular.

If the underlying use is fair, OpenAI argues, copyright holders cannot manufacture market harm merely by creating a market in which they demand payment for permission that copyright law does not require.[1:10]

That is another reason the 14-instance statistic cannot decide the whole case.

Whether ChatGPT reproduces a book to a user and whether OpenAI should have paid to place that book into training data are analytically distinct questions.

What the 0.00007% figure genuinely helps OpenAI establish

Assuming Judge Stein accepts the methodology, the evidence appears especially helpful to OpenAI on one proposition:

ChatGPT does not appear to routinely function as a system for delivering verbatim copies of these books to ordinary users.

That matters.

If users could routinely type a title and receive chapter after chapter of the original book, the argument that ChatGPT substitutes for legitimate copies would be substantially stronger.

A rate this low cuts in the opposite direction.

It may also help OpenAI argue that its technical objective is not to publicly redistribute training books and that regurgitation is an incidental failure mode rather than the ordinary function of the product.

But reasonable inference has limits.

The statistic alone does not establish:

  • that OpenAI lawfully obtained every training book;
  • that downloading books from LibGen was fair use;
  • that every intermediate training copy was fair use;
  • that no specific ChatGPT output can infringe copyright;
  • that paraphrased or non-verbatim outputs create no copyright issue;
  • that AI-generated books create no economic harm;
  • that an AI-training licensing market is legally irrelevant; or
  • that the observed sample perfectly represents every model, version, user or prompting technique.

Those questions require additional legal and factual analysis.

The denominator deserves more attention than it is getting

OpenAI’s filing describes the expert search this way: 95 million outputs from 20 million ChatGPT conversation logs, producing 14 identified instances.[1:11]

It then calls the result 0.00007%.

The arithmetic shows why terminology matters:

  • 14 / 20,000,000 conversations = 0.00007%
  • 14 / 95,000,000 outputs ≈ 0.000015%

The latter number is actually smaller.

So the distinction does not undermine OpenAI’s basic claim that observed verbatim regurgitation was extremely rare.

It does undermine any headline or argument that treats “95 million outputs” and “0.00007%” as though one were simply calculated from the other.

The better formulation is:

OpenAI says its expert found 14 verbatim or near-verbatim instances while searching 95 million outputs contained in 20 million conversations, and describes the affected-conversation rate as approximately 0.00007%.

That is both more precise and more informative.

The plaintiffs’ 31-instance result may be just as important

OpenAI’s brief says the plaintiffs themselves examined a larger, overlapping dataset and identified 31 instances in 108 million conversation logs, which OpenAI calculates as approximately 0.00003%.[1:12]

If that characterization survives adversarial scrutiny, it makes the ordinary-use evidence more difficult to dismiss as merely a defendant-created experiment.

But caution remains necessary.

The public briefing does not yet expose every detail necessary to compare methodologies cleanly: definitions of near-verbatim similarity, sampling decisions, model coverage, time periods, minimum match lengths and other expert choices can materially affect results.

The proper conclusion is not that the regurgitation question has been scientifically settled.

It is that both sides appear to have encountered very low rates when looking at large bodies of actual conversation data, at least as OpenAI describes the current record.

This does not mean regurgitation is irrelevant to copyright

Another overcorrection would be to say that because regurgitation is rare, outputs no longer matter.

Individual outputs can still present individual copyright questions.

Copyright infringement is not calculated by averaging infringing and noninfringing activity across every interaction a product has ever generated.

An allegedly infringing output must still be assessed based on what it contains, how much protected expression it reproduces, substantial similarity and the legal theory being asserted.

A low system-wide frequency can be highly relevant to market substitution and the overall character of a service without immunizing every particular output.

The news-publisher case has different numbers

The consolidated litigation also contains claims from The New York Times and other news organizations.

OpenAI’s separate summary-judgment brief in those cases reports another set of figures. Its expert allegedly found 24 instances of verbatim regurgitation in 20 million ChatGPT conversation logs, or 0.00012%. OpenAI also reports different rates from analyses performed by various news plaintiffs’ experts.[8]

Those numbers reinforce an important point:

There is no single universal “ChatGPT regurgitation rate.”

The result depends on what copyrighted works are being searched for, which conversations or outputs are examined, the models involved, the definition of a match and the methodology used.

The book evidence should therefore be described as evidence about the asserted books in this litigation, not as proof that every copyrighted work appears at the same rate across every OpenAI model.

Why summary judgment changes the stakes

Until now, much of the public argument over generative AI copyright has occurred while courts were dealing with pleadings, discovery disputes or cases involving different companies and different records.

The OpenAI litigation has moved further.

OpenAI and the class plaintiffs have now submitted competing summary-judgment motions. The plaintiffs ask Stein to rule, among other things, that they have established infringement, that OpenAI’s challenged LibGen acquisition was not fair use, that further copying for LLM training was not fair use, and that certain distributions were not fair use.[2:7]

OpenAI asks the court to hold that its challenged use of the books qualifies as fair use.[1:13]

Summary judgment does not force Stein to choose one side’s entire theory.

Under Rule 56, a court can resolve an issue without trial only where there is no genuine dispute over a material fact and the moving party is entitled to judgment as a matter of law. Where the evidentiary record leaves a material factual dispute, Stein can deny summary judgment and leave that question for later proceedings.

So saying that “the judge can now decide fair use” is substantially right, but incomplete.

Stein can now decide fair-use questions on a developed record. He is not required to decide every disputed question at summary judgment.

More of the record is about to become public

The current briefs remain substantially redacted.

Under Judge Stein’s sealing schedule, the parties are supposed to publicly refile their summary-judgment briefs and Rule 56.1 factual statements by September 17, 2026, leaving unredacted material for which no party or third party continues to seek sealing.[9]

The underlying exhibits, declarations and expert reports operate on a different sealing schedule, so September 17 should not be interpreted as the date when everything becomes public.

Opposition briefs are due October 9, followed by replies on November 6.[9:1]

Those filings could materially change how strong either side’s factual case appears.

There is currently no basis to assume Stein has already decided the fair-use issue or that a ruling will arrive immediately after briefing.

So, does 0.00007% prove OpenAI’s training is fair use?

No.

It potentially proves something narrower and important.

If the evidence survives methodological scrutiny, it is strong evidence against the proposition that ordinary ChatGPT use commonly gives people verbatim substitutes for the plaintiffs’ books.

That helps OpenAI.

It is particularly relevant to direct market substitution and to OpenAI’s argument that its models were designed to generate new language rather than redistribute the books used during training.

But copyright law does not begin and end at the output screen.

The court still has to confront how particular books were acquired, whether acquisition and training should be analyzed together or separately, what intermediate copies were made, whether generative competition can constitute cognizable market harm, and whether the growing market for AI-training licenses legally matters.

The clearest conclusion from the current record is therefore neither:

OpenAI pirated books, so everything that followed was infringement.

Nor:

ChatGPT almost never quotes books, so training must be fair use.

The evidence supports a more precise conclusion:

OpenAI’s remarkably low claimed rate of ordinary-use book regurgitation may substantially weaken one theory of copyright harm—direct substitution through copied outputs—without resolving the separate legal questions surrounding source acquisition, model training and broader market effects.

Judge Stein is now being asked to decide how those pieces fit together.

And that, rather than the 0.00007% number alone, could determine whether this case becomes one of the most consequential U.S. copyright decisions of the generative-AI era.


Endnotes

References and Further Reading

Primary filings in the OpenAI litigation

Governing law and major precedents

Independent coverage

Editorial currency note: This article reflects the public record available through September 11, 2026. Judge Stein’s current schedule calls for additional portions of the parties’ summary-judgment briefs and Rule 56.1 statements to become public on September 17 where continued sealing has not been requested. Opposition and reply briefing may also materially change the evidentiary record or the parties’ legal arguments.

  1. OpenAI. “OpenAI’s Memorandum of Points and Authorities in Support of Motion for Summary Judgment.” U.S. District Court for the Southern District of New York, filed September 4, 2026. https://storage.courtlistener.com/recap/gov.uscourts.nysd.606655/gov.uscourts.nysd.606655.1221.0.pdf ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  2. Authors Guild and Class Plaintiffs. “Class Plaintiffs’ Memorandum of Law in Support of Motion for Partial Summary Judgment.” In re OpenAI, Inc. Copyright Infringement Litigation, U.S. District Court for the Southern District of New York, publicly filed September 5, 2026. https://authorsguild.org/app/uploads/2026/09/2026-09-05-Redaction-dckt-1881_0.pdf ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  3. U.S. Congress. “17 U.S.C. § 107 — Limitations on Exclusive Rights: Fair Use.” Legal Information Institute, Cornell Law School. https://www.law.cornell.edu/uscode/text/17/107 ↩︎

  4. U.S. Supreme Court. “Andy Warhol Foundation for the Visual Arts, Inc. v. Goldsmith, 598 U.S. 508.” May 18, 2023. https://www.supremecourt.gov/opinions/22pdf/21-869_87ad.pdf ↩︎

  5. U.S. District Court for the Northern District of California. “Bartz v. Anthropic PBC — Order on Fair Use.” June 23, 2025. https://www.courtlistener.com/docket/69058235/231/bartz-v-anthropic-pbc/ ↩︎

  6. U.S. District Court for the Northern District of California. “Kadrey v. Meta Platforms, Inc. — Order on Cross-Motions for Partial Summary Judgment.” June 25, 2025. https://law.justia.com/cases/federal/district-courts/california/candce/3:2023cv03417/415175/598/ ↩︎ ↩︎

  7. U.S. Court of Appeals for the Second Circuit. “Authors Guild v. Google, Inc., 804 F.3d 202.” October 16, 2015. https://www.courtlistener.com/opinion/3137282/authors-guild-v-google-inc/ ↩︎

  8. OpenAI. “OpenAI’s Memorandum in Support of Motion for Summary Judgment in the News Cases.” U.S. District Court for the Southern District of New York, filed September 4, 2026. https://storage.courtlistener.com/recap/gov.uscourts.nysd.612697/gov.uscourts.nysd.612697.1496.0.pdf ↩︎

  9. U.S. District Court for the Southern District of New York. “Summary-Judgment Briefing and Sealing Schedule in In re OpenAI Copyright Infringement Litigation.” September 2026 docket entries. https://dockets.justia.com/docket/new-york/nysdce/1%3A2026cv01947/659334 ↩︎ ↩︎

Cite this article

Published September 11, 2026

Think something here is wrong, incomplete, outdated, or insufficiently supported? You can challenge a factual claim, source, interpretation, missing context, or privacy issue.

Learn How the challenge process works


More to think on...