The Justice Department has weighed in on the copyright fight over generative AI, telling a court that training models on copyrighted text is fair use. In a brief filed Sept. 1, it called OpenAI's training "exceedingly transformative" and warned that forcing companies to license all ingested material would make AI training impermissible and cede advantage to foreign adversaries.

Eight days later, the NSA, CISA, and FBI issued a joint advisory accusing six Chinese labs—including DeepSeek, Moonshot AI, and Alibaba—of running "industrial-scale distillation campaigns" against American frontier models. Distillation trains a smaller model to mimic a larger one, and the agencies described the activity as malicious, targeted, and central to Beijing's AI strategy.

As the Lawfare piece points out, distillation is essentially a specialized form of training. Both processes involve using another party's output without permission, yet the government treats American training as transformative innovation and Chinese distillation as a national security threat. The article asks whether there is a legal or rational basis for that difference, and how the law should handle mass harvesting of data—whether human-written or machine-generated.