US DOJ Supports OpenAI, Argues AI Training Data Restrictions Threaten American Prosperity

iconCryptoBriefing
Share
AI summary iconSummary
The US Department of Justice filed a brief in a Manhattan court supporting OpenAI and Microsoft against The New York Times in a copyright case. The DOJ claims training AI models on copyrighted data falls under fair use and limiting it could hurt American prosperity and national security. This marks the first time the government has taken a formal stance in such a case, linking AI progress to broader economic and CFT (Countering the Financing of Terrorism) concerns. The move impacts liquidity and crypto markets as AI becomes central to financial infrastructure.

The US Department of Justice just planted its flag in the most consequential copyright battle of the AI era. In a brief filed in Manhattan federal court, the DOJ argued that training large language models on copyrighted material qualifies as fair use, and that restricting it would threaten American prosperity, economic mobility, and national security.

The filing landed in the consolidated lawsuit between The New York Times and OpenAI/Microsoft, a case that has been grinding through the courts since December 2023. It marks the first time the federal government has formally intervened in any of the sprawling copyright disputes surrounding AI training data.

The government’s argument

The DOJ’s core claim is straightforward: when an LLM ingests copyrighted text during training, it is not copying that text in any meaningful competitive sense. The outputs do not reproduce or substitute for the original works. In the government’s framing, this process is “extraordinarily transformative,” a legal term of art that carries real weight in fair use analysis.

Advertisement

The brief went further than just defending the mechanics of AI training. It argued that clamping down on data access would “significantly hamper the progress of science and useful arts,” language that deliberately echoes the constitutional purpose of copyright law itself.

Associate Attorney General Stanley E. Woodward Jr., Assistant Attorney General Brett Shumate, and senior counsel Michael Weisbuch all contributed to the filing, signaling this was not some mid-level staff memo.

The national security card

Perhaps the most politically charged element of the DOJ’s argument is its explicit framing of AI training as a national security issue. The brief contends that the United States has a “profound national interest” in maintaining a dominant domestic AI industry, and that overly restrictive copyright interpretations could hand competitive advantages to foreign adversaries.

China was the obvious, if sometimes unstated, subtext. The argument essentially boils down to: if American AI companies cannot train on high-quality English-language data, their Chinese counterparts will not face similar constraints, and the US will fall behind in a technology race with enormous military, economic, and geopolitical stakes.

Wider implications for copyright and AI

The DOJ’s brief does not just apply to the NYT case. Its reasoning extends to related copyright disputes involving other publishers, authors, and creators who have sued AI companies over training data practices. By articulating a broad fair use rationale, the government is effectively laying down a legal template that could shape the outcome of dozens of pending cases.

The key legal question remains unresolved: does the transformative nature of LLM training outweigh the potential market harm to copyright holders? Courts have historically applied a four-factor balancing test for fair use, weighing the purpose and character of the use, the nature of the copyrighted work, the amount used, and the effect on the market for the original.

The DOJ’s position leans heavily on the first and fourth factors. Training is transformative because it does not reproduce the source material in recognizable form. And LLM outputs do not directly compete with the articles, books, or creative works they were trained on, at least in the government’s telling. Publishers would obviously dispute that second point, arguing that AI-generated summaries and answers do cannibalize the audience for original journalism and writing.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.