Big AI's habit of training on other people's work is getting harder to ignore. In the copyright fight between The New York Times, OpenAI, and Microsoft, unsealed court documents quote a Microsoft director calling the practice an "astonishing theft" and possibly the "largest theft of labor in human history" — though those statements appear in plaintiffs' briefs and have not been ruled on. Microsoft and OpenAI dispute the allegations, arguing that using public articles and books to train models is fair use.
Internal Microsoft documents cited in the same case reportedly warn that generative AI could "significantly disrupt the employment" of the people whose data trains it, and that LLMs are "a product that destroys its supply chain." OpenAI's corporate representative also reportedly testified that the company made no effort to detect or remove paywalled content from training datasets, and cofounder Greg Brockman was quoted responding "ah, nice" when told about bypassing firewalls. These are allegations, not final findings.
Separately, the Ninth Circuit handed GitHub, Microsoft, and OpenAI a narrow win in the Doe v. GitHub case, ruling that generating code without copyright-management information is not necessarily the same as removing or altering it. The court did not decide that training on open-source code is always fair use. The Register's opinion piece argues that Big AI's real business model is to take the work and keep the money, and that most users never check the sources behind AI answers.