TechCrunch reported that U.S. law currently lacks a clear conclusion on whether AI models can be trained using copyrighted books. Existing rulings indicate that courts do not necessarily consider the act of "training" itself as infringement, but companies may still face substantial damages if the training materials are sourced from pirated channels.
The Anthropic case sends complex signals.
The article mentions that last year, U.S. Judge William Alsup ordered Anthropic to pay $1.5 billion in copyright settlement fees in a class-action lawsuit brought by authors. However, this ruling did not directly invalidate the act of training AI itself; the judge’s penalty focused on Anthropic’s acquisition of books from an illegal “shadow library.”
According to the court’s reasoning in this case, a model trained on vast amounts of text is more akin to “reading and absorbing content” than simply copying works. The article cites intellectual property attorney Cathy Gellis, who notes that this interpretation is generally more favorable to AI companies, as copyright law primarily restricts copying, rather than directly prohibiting reading, using, or accessing copyrighted content.
The dispute centers on fair use.
The report notes that U.S. copyright law, enacted in 1976, leaves courts to apply outdated rules to new issues posed by today’s generative AI. The key issue in most current cases is whether AI training constitutes “fair use,” particularly whether such use is sufficiently “transformative.”
Courts typically consider several factors, including the purpose of use, the nature of the work, the amount used, and the impact on the original market. The article cites attorney Jason Henderson’s view that if a company uses others’ content to train a model for the purpose of directly creating a competing product, courts are generally more cautious; if there is no direct competition, there is greater room for support.
Whether it constitutes competition remains the focus.
The article cites the case of Thomson Reuters v. Ross Intelligence, in which the court ruled that the latter’s use of the former’s content to develop an AI legal platform did not constitute fair use, partly because the two products were in direct competition.
In contrast, while authors may argue that chatbots generating new content based on their works undermines their income, this argument has not yet achieved a decisive victory in court. That is, whether AI outputs are sufficient to substitute for the original author’s market remains a key issue in future cases.
AI-generated content also has another layer of issues.
The article also notes that whether AI training is legal and whether AI-generated content can be copyrighted are two distinct issues. Previously, in the case of Thaler v. Perlmutter, a U.S. court ruled that works entirely generated by AI cannot be granted copyright protection.
This also raises new practical issues: how to define the proportion of human creation when a work is partially assisted by AI, a question that may continue to spark debate in the future. TechCrunch believes that, with many cases still under review, U.S. courts are unlikely to provide a unified answer in the short term, but existing rulings have already begun to influence how AI companies train their models and use data.
