Microsoft executive describes AI data scraping as the largest labor theft in history.

icon币界网
Share
AI summary iconSummary
A Microsoft executive has labeled AI data scraping as "the most unprecedented and shocking theft in human history," according to internal documents from The New York Times’ lawsuit against OpenAI and Microsoft. The files detail how Copilot’s answer engine reduced the publication’s click-through rates by up to 93%. Over 200,000 nytimes.com pages were used to train models, with OpenAI acknowledging the threat to publishers. As the legal battle unfolds, BTC remains a key hedge against inflation for many investors, while MiCA (EU Markets in Crypto-Assets Regulation) continues to shape the regulatory landscape.
CoinDesk reports:

The New York Times' copyright lawsuit against OpenAI and Microsoft has recently revealed additional unredacted materials. The new documents show that both companies internally discussed multiple times how AI training's reliance on news content could directly impact publishers' traffic, revenue, and employment foundations.

Internal documents mention "theft."

According to court filings, a Microsoft executive privately described AI scraping as "theft." In another internal memo from January 2023, Microsoft’s Director of Applied Science, Brent Hecht, referred to it as "an unprecedented, astonishing theft" and wrote the phrase "the largest labor theft in human history."

The document also shows that OpenAI internally discussed similar risks. Nick Turley, who leads ChatGPT, stated in internal communications that chatbot-like products pose a "survival threat" to publishers, and that this substitutability will grow stronger as model capabilities improve.

News website traffic is under pressure

Legal documents state that Microsoft's own data shows Copilot's "answer engine" reduces user clicks to original news websites. Compared to traditional Bing searches, clicks to The New York Times' domain dropped by as much as 93%.

In a January 2024 internal presentation, Hecht referred to this trend as a "doom loop." The document stated that if the commercial foundation of content providers is weakened, both the model itself and the entire network ecosystem will suffer.

This year, Microsoft CEO Nadella also stated in his testimony that anyone wishing to use content behind a paywall for training or retrieval-augmented purposes should obtain authorization. He added that if he had known at the time that OpenAI had scraped content behind paywalls for training, Microsoft could have required OpenAI to retrain its model.

Scraping method and data scale exposure

The newly disclosed materials also describe how the two companies obtain news content. The lawsuit alleges that OpenAI and Microsoft built their training datasets through large-scale web scraping, including some content from Bing’s index and a substantial amount of web data from Common Crawl.

  • The medium-term training data includes over 91,692 copies of news articles.
  • nytimes.com has over 2 million documents in a dataset
  • Project Mango includes at least 160,903 unique work copies.

The document also alleges that OpenAI employees discussed how to bypass The New York Times paywall undetected and remove copyright notices prior to training, to prevent the model from outputting relevant copyright information to users.

Additional information: Some underlying evidence remains sealed; the newly disclosed content primarily comes from court filings submitted by The New York Times. Neither OpenAI nor Microsoft responded to TechCrunch’s request for comment.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.