The New York Times' copyright lawsuit against OpenAI and Microsoft has recently revealed additional unredacted materials. The new documents show that both companies internally discussed multiple times how AI training's reliance on news content could directly impact publishers' traffic, revenue, and employment foundations.
Internal documents mention "theft."
According to court filings, a Microsoft executive privately described AI scraping as "theft." In another internal memo from January 2023, Microsoft’s Director of Applied Science, Brent Hecht, referred to it as "an unprecedented, astonishing theft" and wrote the phrase "the largest labor theft in human history."
The document also shows that OpenAI internally discussed similar risks. Nick Turley, who leads ChatGPT, stated in internal communications that chatbot-like products pose a "survival threat" to publishers, and that this substitutability will grow stronger as model capabilities improve.
News website traffic is under pressure
Legal documents state that Microsoft's own data shows Copilot's "answer engine" reduces user clicks to original news websites. Compared to traditional Bing searches, clicks to The New York Times' domain dropped by as much as 93%.
In a January 2024 internal presentation, Hecht referred to this trend as a "doom loop." The document stated that if the commercial foundation of content providers is weakened, both the model itself and the entire network ecosystem will suffer.
This year, Microsoft CEO Nadella also stated in his testimony that anyone wishing to use content behind a paywall for training or retrieval-augmented purposes should obtain authorization. He added that if he had known at the time that OpenAI had scraped content behind paywalls for training, Microsoft could have required OpenAI to retrain its model.
Scraping method and data scale exposure
The newly disclosed materials also describe how the two companies obtain news content. The lawsuit alleges that OpenAI and Microsoft built their training datasets through large-scale web scraping, including some content from Bing’s index and a substantial amount of web data from Common Crawl.
- The medium-term training data includes over 91,692 copies of news articles.
- nytimes.com has over 2 million documents in a dataset
- Project Mango includes at least 160,903 unique work copies.
The document also alleges that OpenAI employees discussed how to bypass The New York Times paywall undetected and remove copyright notices prior to training, to prevent the model from outputting relevant copyright information to users.
Additional information: Some underlying evidence remains sealed; the newly disclosed content primarily comes from court filings submitted by The New York Times. Neither OpenAI nor Microsoft responded to TechCrunch’s request for comment.
