Microsoft exec called AI scraping ‘the largest theft of labor in human history,' new unredacted filings reveal

Newly unsealed court filings show Microsoft privately called OpenAI's data practices "theft" while both companies scraped paywalled Times content, built datasets from it, and warned internally it would gut publishers.

Microsoft exec called AI scraping ‘the largest theft of labor in human history,' new unredacted filings reveal

TL;DR

  • New unredacted court filings reveal Microsoft executives privately described AI data scraping as "theft."
  • OpenAI's leadership acknowledged its AI models pose an "existential threat" to publishers and journalists.
  • The companies allegedly bypassed paywalls, scraped content en masse, and stripped copyright notices.
  • Microsoft's Copilot "answer engine" caused a drastic drop in The New York Times' domain click-through rates, described as a "doom loop."
  • Microsoft CEO Satya Nadella stated that paywalled content should be licensed for AI training.
  • Internal communications describe publishers facing an "existential threat" from AI chatbots that are "largely substitutive."
  • OpenAI's mid-training datasets contained over 91,692 copies of works from NYT, Daily News, and Center for Investigative Reporting.
  • Microsoft provided training data to OpenAI through initiatives like Project Taxi and Project Mango.
  • OpenAI employees allegedly devised methods to circumvent paywalls undetected.
  • Researchers allegedly removed copyright notices to prevent models from outputting them.