US Department of Justice backs fair use for AI training in landmark copyright case
The US Department of Justice filed a brief siding with AI companies in the New York Times v. OpenAI/Microsoft copyright case, arguing that training LLMs on copyrighted material qualifies as fair use The DOJ draws a sharp legal distinction between copying for training purposes and model outputs, noting that training copies are never publicly distributed and outputs "often if not always lack substantial similarity" to originals The DOJ invokes a Joan Didion analogy, arguing that requiring payment
Analysis
TL;DR
- The US Department of Justice filed a brief siding with AI companies in the New York Times v. OpenAI/Microsoft copyright case, arguing that training LLMs on copyrighted material qualifies as fair use
- The DOJ draws a sharp legal distinction between copying for training purposes and model outputs, noting that training copies are never publicly distributed and outputs "often if not always lack substantial similarity" to originals
- The DOJ invokes a Joan Didion analogy, arguing that requiring payment whenever creators draw on prior works would stifle the creativity copyright law aims to protect
- The filing directly challenges the US Copyright Office's earlier report that rejected blanket fair use for AI training, dismissing it as carrying no binding legal authority
- Political overtones emerge as former Copyright Register Shira Perlmutter was fired by the Trump administration after her report, with Democrats alleging the dismissal was tied to her refusal to legitimize AI training on copyrighted works
Why It Matters
This DOJ filing represents the most significant federal government intervention in favor of AI companies in an ongoing copyright battle that could reshape the entire generative AI industry. A ruling favoring the DOJ's fair use interpretation would lock in the current training paradigm for billions of dollars in AI development, while a rejection could force companies into costly licensing deals or fundamentally alter how models are built.
Technical Details
- The case centers on allegations that OpenAI and Microsoft used millions of New York Times articles without permission to train GPT-4 and competing products, with the Times seeking billions in damages and model destruction
- The DOJ's fair use argument hinges on the four-factor test, particularly emphasizing that training copies are non-expressive and never made publicly available, and that model outputs lack substantial similarity to source works
- The DOJ directly contests the US Copyright Office's 2023 report, arguing it ignored case law on case-by-case fair use analysis and misidentified the types of market harm recognized under copyright statute
- The Joan Didion/Hemingway analogy is used to frame AI training as analogous to human creative learning processes, though critics note the scale difference between individual study and corporate mass reproduction
- The filing references the Kadrey ruling to argue against treating a learning process and subsequent creative output as a single infringing use
Industry Insight
- The DOJ's intervention signals strong administrative support for the AI industry's current training practices, but fair use remains an affirmative defense decided case-by-case—this does not establish a legal precedent, only strengthens the companies' position in litigation
- AI companies should prepare for continued legal pressure regardless of this filing; the scale argument raised by the Copyright Office and critics will likely dominate future courtroom debates and legislative efforts
- The political dimension—Perlmutter's firing and the Trump administration's pro-AI stance—suggests copyright policy may become increasingly partisan, creating uncertainty for long-term planning and encouraging companies to pursue licensing deals as a hedge against shifting legal winds
Disclaimer: The above content is generated by AI and is for reference only.