The U.S. Department of Justice has issued a significant legal intervention in the ongoing copyright war between media giants and AI developers. By arguing that training Large Language Models (LLMs) on copyrighted data constitutes "fair use," the DOJ has provided a massive legal shield for companies like OpenAI and Microsoft.

The DOJ’s Argument: Training vs. Output

At the heart of the consolidated lawsuit involving The New York Times is a fundamental distinction between how an AI learns and what it produces. The New York Times alleges that millions of its articles were used without permission to train models such as GPT-4, causing billions of dollars in damages and creating products that directly compete with the newspaper.

However, the DOJ’s recent filing argues that copyright infringement occurs at the point of output, not the point of ingestion. The department contends that while entire works are copied during the training phase, these copies are never made publicly available. Furthermore, the DOJ asserts that the outputs of models like GPT-4 "often if not always lack substantial similarity" to the original source material. By separating the training process from the generated content, the DOJ is attempting to decouple the act of machine learning from the legal concept of market substitution.

The "Hemingway Analogy" and Human Creativity

To make the concept of machine learning relatable, the DOJ invoked a literary analogy involving author Joan Didion. The filing noted that as a teenager, Didion would copy Ernest Hemingway’s stories to analyze his sentence structure and learn his craft. The DOJ argues that if the law treated machine training as infringement, a writer like Didion could theoretically face liability every time she published new work, as her learning process would be inextricably linked to her subsequent writing.

The department further argues that imposing strict liability for AI training would paradoxically stifle the very creativity copyright law is intended to protect. With human beings increasingly using LLMs to draft and edit original works, the DOJ suggests that making training impermissible without massive licensing schemes would cripple AI-driven innovation.

The DOJ’s stance directly contradicts recent findings from the US Copyright Office. Former Register Shira Perlmutter had previously argued against a blanket "fair use" defense, noting that AI operates at a scale and speed far beyond human capability and often creates commercial products that compete directly with the original content creators.

The DOJ has gone on the offensive against this assessment, stating that Perlmutter’s report carries no binding legal authority and fails to account for established case law regarding case-by-case analysis. The department warns that imposing broad liability would effectively mandate a licensing regime that would render the development of advanced AI models legally and financially impossible.

Key Takeaways

  • Separation of Training and Output: The DOJ argues that copying text for training does not constitute infringement if the resulting model outputs do not show "substantial similarity" to the original works.
  • Protection of Innovation: The department warns that requiring licenses for all training data would stifle creativity and prevent the development of LLMs that assist in human content creation.
  • A Legal Bellwether: This intervention creates a high-stakes conflict between the DOJ and the US Copyright Office, setting the stage for a definitive court ruling on the future of generative AI.

ARTICLE: The U.S. Department of Justice filed an amicus brief in the New York Times’ copyright lawsuit, arguing that copying protected text to train large language models (LLMs) qualifies as fair use. The filing gives companies such as OpenAI and Microsoft a legal shield that could keep them out of a damages pool that the newspaper estimates in the billions.

Why the case matters now

The lawsuit consolidates several claims that AI developers fed millions of newspaper articles into their models without permission, then released products that compete directly with the Times’ own reporting. If a court rejects the DOJ’s fair-use argument, AI firms could face massive licensing bills or be forced to halt development of the next generation of conversational agents.

The DOJ’s core argument

The brief splits the AI workflow into two distinct stages. First, the model ingests large corpora of text; second, it generates responses to user prompts. The department says infringement only arises when a copyrighted work is reproduced in a way that the public can access. Because the training copies never leave the developer’s servers, the act of ingestion does not meet that threshold.

The DOJ also leans on the “substantial similarity” test, a long-standing copyright standard. It contends that the text output by models such as GPT-4 “often if not always” differs enough from any single source that a plaintiff cannot prove the required similarity. In short, the brief argues that the legal focus should be on the final output, not the internal learning process.

A literary analogy to illustrate the point

To make the technical argument more relatable, the filing invokes a story about a teenage writer who copied Ernest Hemingway’s stories to study his style. The DOJ says that if the law treated that learning exercise as infringement, the writer could be sued each time she published a new piece, even though her work was original. The department warns that extending the same logic to AI would “paralyze the very creativity copyright law was designed to protect.”

The stakes for AI developers and content creators

  • Innovation risk: Requiring licenses for every piece of text used in training could raise costs so high that building state-of-the-art models becomes financially untenable.
  • Creative workflow: Human writers increasingly rely on LLMs for drafting, editing, and brainstorming. If the training process were deemed illegal, those tools could disappear, reshaping how content is produced across industries.
  • Market impact: The newspaper industry argues that AI models that can answer questions or generate news-like text erode the value of original reporting, potentially siphoning advertising revenue and subscriptions.

A recent report from the U.S. Copyright Office, authored under former Register Shira Perlmutter, pushed back against a blanket fair-use defense. The office highlighted three factors that set AI apart from human learning: the sheer volume of material copied, the speed at which it is processed, and the fact that the resulting models can be packaged and sold as commercial products that directly compete with the source creators.

The DOJ rebuts those points, noting that the report does not carry binding legal authority and that existing case law already requires a fact-by-fact analysis of fair use. It warns that imposing “broad liability” would effectively force the industry into a licensing regime that “would render the development of advanced AI models legally and financially impossible.”

What could happen next

The district court handling the Times case will have to weigh the DOJ’s fair-use position against the Copyright Office’s findings. A ruling in favor of the DOJ would set a precedent that training data can be harvested without explicit permission, provided the model’s outputs stay clear of substantial similarity. A decision siding with the newspaper could trigger a wave of licensing negotiations, potentially reshaping the economics of AI research.

Bottom line

The DOJ’s filing turns the question of whether AI training is fair use into a high-stakes courtroom battle. The outcome will determine whether developers can continue to build powerful language models on existing text or must seek costly licenses for every piece of material they ingest. The decision will ripple through the tech sector, the media industry, and the broader creative economy.