deniz.in

Markets

Weather

Loading weather

· via TechCrunch

Microsoft exec called AI scraping 'the largest theft of labor in human history,' filings show

Unredacted NYT v. OpenAI/Microsoft filings quote a Microsoft exec calling AI scraping 'the largest theft of labor in human history' and detail paywall workarounds and massive-scale copying.

Microsoft exec called AI scraping 'the largest theft of labor in human history,' filings show

Newly unredacted material in the copyright case The New York Times brought against OpenAI and Microsoft contains some of the sharpest language yet from the defendants' own ranks, TechCrunch reports. In an internal memo dated January 2023, Brent Hecht, Microsoft's director of Applied Science, described mass AI scraping as "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." Senior figures at OpenAI, meanwhile, characterized chatbot products as an "existential threat" to the publishers and journalists whose work trained them.

One caveat TechCrunch raises: much of the new detail comes from the Times' own legal brief rather than the underlying exhibits, which remain sealed, and the quoted passages appear without their full original context.

Microsoft's own 'doom loop' warning

According to the filing, Microsoft's internal data showed that its Copilot "answer engine" drove click-through rates for the Times' domain down by as much as 93 percent compared with conventional Bing search. A January 2024 presentation written by Hecht reportedly described this dynamic as a "doom loop" that would "hurt the performance of our models and the entire web at the same time."

Another Microsoft document quoted in the brief concedes that it is "highly unusual" for a finished product to threaten the economic foundations of its essential suppliers, adding that this is the situation the company created for its LLM business and its "content supply chain." A separate internal assessment warned of a "real risk" that generative AI could significantly disrupt the employment of the very people whose data trained the foundation models.

What executives said under oath

Microsoft CEO Satya Nadella, deposed earlier this year, testified that "anything that is paywalled should be licensed by anyone who wants to use it" for grounding or training, and said that had he been aware OpenAI had trained on paywalled information, he would have invoked Microsoft's right to require OpenAI to retrain its models. He also agreed that chatting with a bot can substitute for visiting the underlying source website.

On OpenAI's side, Nick Turley, head of ChatGPT, wrote internally that publishers face an "existential threat" from chatbots that are "largely substitutive" and "will get more and more substitutive as they get better." OpenAI President Greg Brockman described the models as "excellent at news."

The scale of the copying

The filing offers a first look at the sheer volume involved. OpenAI's mid-training datasets alone reportedly contain more than 91,692 copies of works published by the Times, the Daily News and the Center for Investigative Reporting, while a dataset derived from Common Crawl included more than two million documents from nytimes.com alone.

The two companies also exchanged data directly, per the filing: OpenAI handed Microsoft its entire GPT-3 training dataset to evaluate how to embed OpenAI's models in commercial products, and Microsoft supplied training data back through efforts code-named Project Taxi and Project Mango. The Project Mango material was allegedly assembled into a dataset containing copies of at least 160,903 unique works from the news publishers.

Paywalls and copyright notices

The unsealed passages describe how the content was obtained. When OpenAI researcher Nick Ryder told Brockman about a "hack to get around nytimes paywall," Brockman reportedly replied "ah nice." Internal datasets such as WebText and WebText2 leaned disproportionately on scraped news, millions of articles were pulled from Common Crawl, and employees allegedly stripped copyright notices from training data before it reached the models — partly, as the filing puts it, because researchers "wouldn't want model outputting" such notices to users.

TechCrunch notes that neither OpenAI nor Microsoft responded to requests for comment.

Fair use under strain

Whether AI firms may legally train on copyrighted material remains unsettled, but courts have so far been largely receptive to fair use arguments, and earlier this month the Trump administration filed a brief defending OpenAI's unlicensed use of copyrighted works for training. The new admissions matter because they speak directly to one pillar of the fair use test: the requirement that the use not substitute for or harm the market for the original. Internal data showing collapsing click-through rates, plus executives describing chatbots as substitutive under oath, could complicate that defense.

Why it matters

The filings place the defendants' internal assessments in direct conflict with their legal posture. Statements from Microsoft's own scientists and its CEO about theft, licensing obligations and substitution could become pivotal evidence on market harm — the question on which several pending AI copyright cases may ultimately turn. For publishers, the documents provide the most detailed accounting yet of how much of their output entered training corpora and how paywalls and copyright notices were allegedly worked around. And for the web at large, Microsoft's "doom loop" framing — answer engines starving the very sources they depend on — lays out stakes that extend well beyond this single lawsuit.

  • #copyright
  • #openai
  • #microsoft
  • #fair-use
  • #new-york-times
  • #ai-training

Related posts