NYT v. OpenAI & Microsoft: The Fair-Use Trial

On this page10 sections

NYT v. OpenAI & Microsoft: The Fair-Use Trial

NYT v. OpenAI & Microsoft
Date:
December 27, 2023
Location:
Southern District of New York
Significance:
Most consequential copyright case in AI history; will define whether LLM training is fair use
NYT v. OpenAI & Microsoft

A copyright lawsuit filed by The New York Times in December 2023 alleging that OpenAI and Microsoft trained GPT models on millions of NYT articles and that GPT outputs regurgitate near-verbatim excerpts. The case is the highest-profile fair-use test of the LLM era.


The Background: How AI Models Are Trained

To understand the lawsuit, you have to understand how large language models like ChatGPT are trained, and why that training process raises copyright questions.

A large language model is, at its core, a statistical system that predicts what word comes next in a sequence of text. To build this system, you feed it enormous amounts of text — billions of words, drawn from books, articles, websites, and other sources. The model learns the statistical patterns in this text: which words tend to follow which other words, which topics are associated with which vocabulary, how sentences are structured, how arguments are made. When you ask the model a question, it generates an answer by predicting, one word at a time, what should come next, based on the patterns it has learned.

The text that is used to train the model is called the “training data” — the billions of words used to train a language model, and the central subject of the NYT copyright dispute. For the largest models, the training data is collected from across the internet. Companies like OpenAI use web crawlers — automated programs that systematically browse the web — to download billions of web pages. They also use existing datasets, like Common Crawl, which is a non-profit archive of web pages that has been built up over many years. The training data includes news articles, blog posts, books, academic papers, code repositories, social media posts, and many other kinds of text.

Much of this text is copyrighted. News articles are copyrighted by the news organisations that publish them. Books are copyrighted by their authors. Blog posts are copyrighted by their writers. The question is whether training a language model on copyrighted text, without the permission of the copyright holder, is legal.

This is the central question that the NYT v. OpenAI lawsuit is designed to answer. The answer will determine the future of the AI industry. If training on copyrighted text is legal, then AI companies can continue to build models using the enormous amounts of text available on the internet, without paying the creators of that text. If training on copyrighted text is not legal, then AI companies will need to license the text they use for training, which will be expensive and will give large publishers significant leverage over the AI industry.


The Lead-Up: Failed Negotiations

The New York Times Company did not sue OpenAI without warning. The conflict had been building for months.

In April 2023, the Times first reached out to Microsoft and OpenAI to raise intellectual-property concerns. The Times was concerned that its articles had been used to train OpenAI’s models without permission, and that the models could, under certain conditions, reproduce near-verbatim excerpts of those articles. The Times proposed a licensing deal: OpenAI would pay the Times for the right to use its articles for training, and the Times would provide access to its content through a formal API.

Negotiations continued through the summer and fall of 2023, but they failed. The parties could not agree on the terms. The Times wanted a deal that reflected the value of its content to OpenAI’s models. OpenAI, presumably, did not want to pay what the Times was asking. The negotiations broke down.

On August 25, 2023, the Times took a public step: it blocked OpenAI’s GPTBot web crawler (OpenAI’s crawler, which respects robots.txt — a standard web protocol that allows website owners to tell crawlers which pages they are allowed to access) from accessing its website. The Times also blocked the Common Crawl CCBot, which is the crawler used by the non-profit Common Crawl project. This was a signal that the Times was serious about its concerns, and it was also a pragmatic step to prevent future scraping. But it did not remove the Times’ articles that had already been scraped and included in OpenAI’s training data.

The blocking of GPTBot was a significant moment in the relationship between publishers and AI companies. Other publishers, including CNN and ABC News, also blocked GPTBot around the same time. The blocking was a public statement that publishers did not consent to having their content used for AI training, even if the legal question of whether consent was required had not yet been resolved.

By December 2023, the Times had decided to sue. The lawsuit was filed on December 27, 2023, in the Southern District of New York. The case number is 1:23-cv-11195. The presiding judge is Sidney H. Stein, a senior United States District Judge. The magistrate judge handling discovery is Ona T. Wang.


The Complaint

The Times’ complaint is a detailed, 69-page document that lays out the legal and factual basis for the lawsuit. The core allegations are straightforward: OpenAI trained its language models on millions of copyrighted New York Times articles without permission, and the models can, under certain conditions, reproduce near-verbatim excerpts of those articles. Microsoft is named as a co-defendant because it invested billions of dollars in OpenAI, built the supercomputers used to train OpenAI’s models, and distributes OpenAI’s technology through its Copilot and Bing Chat products.

The complaint makes several specific legal claims. It alleges direct copyright infringement — that OpenAI copied the Times’ articles when it trained its models. It alleges vicarious copyright infringement — that Microsoft had the right and ability to control OpenAI’s training, and financially benefited from it. It alleges contributory copyright infringement — that Microsoft knowingly encouraged OpenAI’s infringement. It alleges violations of the Digital Millennium Copyright Act (DMCA), specifically section 1202(b), which prohibits the removal of copyright management information — the argument being that OpenAI stripped the Times’ copyright notices and authorship information from the articles when it included them in the training data.

And it alleges common law unfair competition by misappropriation — the common-law tort of unauthorised commercial use.

The most striking part of the complaint is Exhibit J, a 127-page document that contains 100 examples of ChatGPT regurgitating near-verbatim excerpts of New York Times articles.

The examples include the Pulitzer-winning interactive feature “Snow Fall: The Avalanche at Tunnel Creek” (2012), restaurant reviews, opinion columns, news features, and product recommendation listicles from Wirecutter, the Times’ product-recommendation site. The exhibit is designed to demonstrate that OpenAI’s models have “memorised” the Times’ articles — that the articles are, in a sense, stored in the models, and can be retrieved with the right prompts.

The remedies the Times seeks are aggressive. The complaint asks for statutory damages of up to $150,000 per willfully infringed work. If the Times can show that tens of thousands of its articles were infringed, the damages could reach billions of dollars. The complaint also asks for an injunction — a court order barring OpenAI from using Times content for future training. And, most aggressively, it asks for the destruction of the trained models and training data that contain the Times’ content — a remedy that, if granted, could require OpenAI to retrain its models from scratch without the Times’ data.

This last remedy — destruction of the models — is unusual and has generated significant legal commentary. Some experts believe a court would be unlikely to actually order the destruction of a multi-billion-dollar AI model. Others note that the remedy is available under copyright law, and that the threat of it gives the Times significant leverage in any settlement negotiation. Bloomberg Law called the lawsuit an “existential threat” to OpenAI — capturing the case’s high stakes.


OpenAI’s Response

OpenAI responded to the lawsuit on January 8, 2024, with a blog post titled “OpenAI and journalism.” The response made four main arguments.

First, OpenAI argued that training on publicly available internet materials is fair use — a legal doctrine that allows the use of copyrighted material without permission under certain circumstances, for purposes like criticism, comment, news reporting, teaching, scholarship, or research. OpenAI argued that training a language model on internet text is a transformative use — it does not reproduce the text for its original purpose, but uses it to build a statistical model that can generate new text. OpenAI cited two key precedents: Authors Guild v. Google (2015), in which the Second Circuit Court of Appeals held that Google’s book-scanning project was fair use, and Authors Guild v. HathiTrust (2014), which reached a similar conclusion about a digital library.

Second, OpenAI argued that the Times could have blocked its crawler using robots.txt — a standard web protocol that allows website owners to tell crawlers which pages they are allowed to access. OpenAI’s GPTBot respects robots.txt, and the Times did not block GPTBot until August 2023. OpenAI argued that by not blocking GPTBot earlier, the Times had impliedly consented to having its content scraped.

Third, OpenAI argued that “regurgitation” — the verbatim reproduction of training data — is a “rare bug” that the company is working to eliminate. OpenAI said that the models are not designed to memorise and reproduce training data, and that the examples in the Times’ complaint were the result of unusual conditions.

Fourth, and most aggressively, OpenAI argued that the Times had “manipulated” the prompts to produce the regurgitation examples. OpenAI said that the Times had repeatedly prompted the model with lengthy excerpts of articles, in order to force it to reproduce the rest of the article. OpenAI characterised this as an artificial demonstration that did not reflect ordinary user behaviour.

The Times disputed this characterisation. The Times said that the prompts it used were not unusual, and that the regurgitation occurred even with relatively simple prompts. The court did not resolve this factual dispute at the motion-to-dismiss stage. It is a question of fact that will be resolved later in the case, through discovery and, if necessary, at trial.


The Motion to Dismiss and the March 2025 Ruling

On February 26, 2024, OpenAI filed a motion to dismiss the complaint. The motion made several arguments: that the contributory copyright infringement claims should be dismissed; that the DMCA claims should be dismissed; that the unfair competition claims should be dismissed; that the direct infringement claims were time-barred by the three-year statute of limitations; and that certain claims against Microsoft should be dismissed.

The motion to dismiss was fully briefed and argued over the course of 2024. On March 26, 2025, Judge Stein announced his ruling from the bench. The written memorandum opinion was published on April 4, 2025.

The ruling was a significant, if partial, victory for the Times. Judge Stein denied OpenAI’s motion to dismiss the direct infringement claims. He rejected OpenAI’s argument that the claims were time-barred, holding that the claims involving conduct occurring more than three years before the suit was filed could proceed. He denied the motion to dismiss the DMCA claims against OpenAI in the parallel Daily News and Center for Investigative Reporting actions. He granted the motion to dismiss the DMCA claims against Microsoft in all three complaints, for insufficient factual allegations. And he allowed the central copyright infringement theory to proceed.

The ruling was significant because it meant that the case would continue. The fair use question — the central legal issue — was not resolved at the motion-to-dismiss stage. Fair use is a fact-intensive inquiry that is typically reserved for summary judgment (a ruling without trial) or trial. The ruling meant that the case would move to discovery, and that the fair use question would be decided later, on a fuller record.

Eric Goldman, a law professor at Santa Clara University, reviewed the ruling and noted that it featured “a series of narrow holdings on issues such as DMCA applicability, statutes of limitations, and copyright preemption.” The ruling did not decide the big question — whether training on copyrighted web content is fair use — but it cleared the way for that question to be decided.


The Discovery Phase and the ChatGPT Log Controversy

After the motion to dismiss was denied, the case moved to discovery — the pre-trial phase in which the parties exchange evidence. Discovery in copyright cases is typically extensive, but in this case it raised an unusual and controversial issue: access to OpenAI’s ChatGPT conversation logs.

The Times argued that it needed access to ChatGPT’s logs to demonstrate that the model could regurgitate Times content in ordinary use, not just in response to “manipulated” prompts. The Times wanted to analyse millions of ChatGPT conversations to look for evidence of memorisation — a model’s storage of training text closely enough to reproduce it — and regurgitation.

OpenAI resisted. The company argued that the ChatGPT logs contained sensitive user data — the private conversations of hundreds of millions of users — and that producing them would be a massive privacy violation. OpenAI also argued that the logs were not relevant to the legal questions in the case, which were about training, not about how users interact with the model.

The discovery dispute made its way through the courts over the course of 2025. On November 22, 2024, Magistrate Judge Wang denied OpenAI’s motion to limit discovery. On May 13, 2025, Judge Wang issued a broad preservation order requiring OpenAI to retain all ChatGPT conversation logs. On June 26, 2025, Judge Stein heard oral argument on OpenAI’s objections to the preservation order and ultimately denied the objection, affirming the order.

On November 7, 2025, Magistrate Judge Wang ordered OpenAI to produce 20 million anonymised ChatGPT conversation logs — covering the period from December 2022 to November 2024 — to the news plaintiffs and class plaintiffs. The plaintiffs’ experts would analyse the logs for evidence of memorisation and regurgitation. OpenAI publicly opposed the order on privacy grounds, and the order triggered significant public debate about AI-user privacy.

The production of 20 million ChatGPT logs was a remarkable development. It meant that one of the largest AI companies in the world would be required to turn over a substantial sample of its users’ conversations — anonymised, but still potentially revealing — to a group of news organisations. The order raised serious questions about the privacy of AI conversations, and about the balance between the needs of litigation and the privacy interests of users. These questions are not unique to this case; they will recur in other AI-related litigation, and they may eventually require legislative attention.


The Parallel Licensing Strategy

While the lawsuit was making its way through the courts, OpenAI was pursuing a parallel strategy: signing licensing deals with major publishers. These deals were significant both commercially and strategically. Commercially, they gave OpenAI access to high-quality content for training and for real-time integration into ChatGPT. Strategically, they served as evidence that OpenAI was willing to pay for content, and that the unlicensed training that the Times was challenging was a temporary practice that had been replaced by a licensed model.

The deals were extensive. In December 2023, OpenAI signed a deal with Axel Springer (the German publisher of Bild and Politico) — the first major publisher to sign an OpenAI licensing deal. In the spring of 2024, OpenAI signed a deal with the Financial Times, reportedly worth $5 to $10 million. In May 2024, OpenAI signed deals with Vox Media and The Atlantic. Also in May 2024, OpenAI signed a much larger deal with News Corp — the parent company of the Wall Street Journal, the New York Post, and The Times of London — reportedly worth more than $250 million over five years. OpenAI also signed deals with Dotdash Meredith, Politico, the Associated Press, Le Monde, Prisa (the parent of El País), and Reddit.

The Times and the other suing publishers argued that these deals proved that OpenAI knew its unlicensed training was unlawful — if the training was legal, why would OpenAI pay for licenses? OpenAI argued the opposite: the deals showed that OpenAI was pursuing voluntary licenses in good faith, and that the unlicensed training that had occurred before the deals was a reasonable interpretation of fair use.

This argument cuts both ways, and the court has not yet ruled on it. But the licensing deals have had a significant effect on the industry. They have created a market for AI training data, and they have given large publishers leverage that they did not have before. They have also created a divide between publishers that have signed deals with OpenAI and publishers that have sued. The Times is in the latter category, along with the Daily News, the Chicago Tribune, the Orlando Sentinel, the Center for Investigative Reporting, and others.


The Broader Litigation Landscape

The NYT v. OpenAI lawsuit is not the only AI copyright case working its way through the courts. It is part of a broader wave of litigation that includes at least 35 copyright lawsuits against AI companies, according to a tracker maintained at chatgptiseatingtheworld.com.

The most prominent parallel cases include:

  • Authors Guild v. OpenAI (filed September 2023): A class action by 17 prominent authors, including John Grisham, Jodi Picoult, David Baldacci, and George R.R. Martin, alleging that OpenAI trained its models on their books without permission.

  • Doe v. GitHub (filed November 2022): A class action by open-source programmers against GitHub Copilot, Microsoft, and OpenAI, alleging DMCA violations for training on GPL-licensed code. The court dismissed most of the DMCA claims in 2024.

  • Concord Music Group v. Anthropic (filed October 2023): Universal, Concord, and ABKCO sued Anthropic over the use of more than 500 song lyrics in training its Claude models.

  • Getty Images v. Stability AI (filed January 2023 in the UK, February 2023 in the US): A lawsuit over the use of millions of Getty photos to train Stable Diffusion. (This case is the subject of the companion piece B80.)

  • Raw Story Media v. OpenAI and AlterNet v. OpenAI (2024): DMCA claims that were dismissed for failure to show concrete injury.

These cases are proceeding on different timelines and raising different legal questions, but they all turn on the same fundamental issue: whether training AI models on copyrighted content, without the permission of the copyright holder, is legal. The NYT v. OpenAI case is the most prominent of these cases, both because of the prominence of the Times and because of the aggressive remedies it seeks. But the outcome of any one of these cases will affect the others, and the collective outcome will shape the AI industry.

On December 6, 2024, the Judicial Panel on Multidistrict Litigation consolidated 12 of these copyright cases against OpenAI and Microsoft into a single multidistrict litigation (MDL) proceeding — MDL No. 3143 — before Judge Stein in the Southern District of New York. The MDL was formally transferred on April 11, 2025. The consolidation means that the cases will be handled together for pre-trial purposes, which will streamline the litigation and ensure consistent rulings on common legal questions.


The Stakes

The stakes in NYT v. OpenAI are enormous, both for the parties and for the broader AI industry.

If the Times wins, the consequences for OpenAI could be severe. Statutory damages of up to $150,000 per willfully infringed work, multiplied by potentially tens of thousands of Times works, could produce damages in the billions of dollars. An injunction barring the use of Times content for training could force OpenAI to retrain its models without the Times’ data — a costly and time-consuming process. And the most aggressive remedy — destruction of the trained models containing Times content — could theoretically require OpenAI to retrain GPT-4 and later models from scratch. Whether a court would actually order this is uncertain, but the threat is real.

A Times win would also set a precedent — at least in the Southern District of New York — that training on copyrighted web content without a license is not fair use. This precedent would upend the generative AI industry’s data-acquisition model. AI companies would need to license the text they use for training, which would be expensive and would give large publishers significant leverage. It would also create a multi-billion-dollar licensing market that would benefit large incumbent rights holders — the Times, News Corp, the major record labels, the major book publishers — at the expense of smaller creators and the open internet.

If OpenAI wins, the consequences would be equally significant. A ruling that training on copyrighted web content is fair use would dramatically reduce litigation risk across the AI industry. It would validate the data-acquisition model that the major AI companies have been using, and it would allow them to continue building models using the enormous amounts of text available on the internet without paying the creators of that text. It would also weaken the parallel suits — the Authors Guild case, the music publishers’ case, the GitHub Copilot case, the image cases — since most of them rely on the same fair-use denial theory.

A win for OpenAI would also have consequences for the nascent publisher-licensing market. If training on copyrighted web content is fair use, then the licensing deals that OpenAI has signed with News Corp, Axel Springer, the Financial Times, and others may have been unnecessary — OpenAI may have overpaid. This could cause the licensing market to collapse, or at least to shrink significantly.


Where the Case Stands

As of mid-2025, the case is in the active discovery phase. The core copyright infringement claims — direct, vicarious, and contributory — have survived the motion to dismiss. The DMCA claims against Microsoft have been dismissed, but the DMCA claims against OpenAI in the Daily News and Center for Investigative Reporting actions have survived. The fair use question has been reserved for summary judgment or trial. No trial date has been set.

The case is likely to take years to resolve. The discovery phase, which is already producing disputes over ChatGPT logs and other evidence, will continue through 2025 and into 2026. After discovery, the parties will likely file motions for summary judgment, which will ask the court to decide the case without a trial. If the summary judgment motions are denied, the case will go to trial. A trial could happen in 2026 or 2027, and an appeal could extend the case into 2028 or beyond.

In the meantime, the case is casting a long shadow over the AI industry. AI companies are watching it closely, because its outcome will determine the legal framework within which they operate. Publishers are watching it closely, because its outcome will determine whether they have leverage to extract licensing fees from AI companies. And policymakers are watching it closely, because if the courts do not resolve the question of AI training and copyright, Congress may need to step in with new legislation.

The case is, in a sense, the first big test of whether the copyright system — a system that was designed for a world of physical copies and discrete works — can adapt to a world of AI training and statistical models. The outcome will not just determine who pays whom. It will determine the relationship between the creators of text and the builders of AI systems, and it will shape the development of one of the most important technologies of the twenty-first century.


Further reading
  • The New York Times Company v. Microsoft Corporation — Complaint (filed 27 December 2023) — The primary source. The 69-page complaint, including the 127-page Exhibit J with 100 regurgitation examples, is the foundation of the case.
  • Judge Stein’s April 4, 2025 Memorandum Opinion — The written opinion denying OpenAI’s motion to dismiss in large part. Essential for understanding the legal state of the case.
  • CourtListener docket for 1:23-cv-11195 — The full docket of the case, with all filings and orders. courtlistener.com/docket/67507323
  • CourtListener MDL No. 3143 docket — The consolidated multidistrict litigation docket, covering 12 parallel copyright cases against OpenAI and Microsoft.
  • “OpenAI and journalism” — OpenAI blog, 8 January 2024. OpenAI’s official response to the lawsuit. openai.com/blog/openai-and-journalism
  • chatgptiseatingtheworld.com AI copyright lawsuit tracker — a regularly updated catalogue of AI copyright lawsuits. chatgptiseatingtheworld.com

Series Companions

This piece is part of Minds & Machines: Beyond the Series. The companion pieces B80 — Getty Images v. Stability AI (the image-AI parallel to this text-AI case, turning on the same fair-use question), B82 — Authors Guild v. Google (the 2015 book-scanning precedent OpenAI cites as its key fair-use authority), the main-series A26 — The Future of Creativity (the broader cultural context), and the main-series A23 — The Governance Gap (the regulatory context) cover the related milestones.

What would change if more people understood the story behind nyt v. openai? Who benefits from the current state of affairs, and who is left out? The conversation is worth having — with colleagues, with students, with anyone who uses technology without thinking about where it comes from.