deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

Microsoft VP reportedly calls AI scraping 'largest theft of labor' in court filing

A Microsoft vice president described AI training-data scraping as the largest theft of labor in history in unsealed court filings, an admission that could surface across more than 30 pending copyright suits.

Microsoft VP reportedly calls AI scraping 'largest theft of labor' in court filing

What the filing says

A Microsoft vice president has described the scraping of creative work for AI training data as "the largest theft of labor in human history," according to a dev.to report citing TechCrunch. The dev.to post, published September 18, 2026, says the remark appears in unredacted discovery documents from ongoing copyright litigation that TechCrunch reported on a day earlier. Because the language surfaced in a court filing rather than a keynote or press release, the post argues it functions as an acknowledgment of scale rather than marketing rhetoric, even if Microsoft's lawyers introduced it to position the company defensively.

A product stack built on contested data

The statement lands awkwardly against Microsoft's commercial reality. The same company sells GitHub Copilot, Microsoft 365 Copilot and Azure OpenAI Service to millions of business customers, and each of those products depends on training corpora assembled largely without explicit licensing. The dev.to article points to Common Crawl — which it describes as holding over 170 billion tokens and petabytes of web content — as a foundation for many large language models, including the ones Microsoft commercializes through its OpenAI partnership.

Microsoft's answer to legal exposure is its Copilot Copyright Commitment, announced in September 2023 and updated in 2025. The pledge indemnifies commercial customers against copyright claims arising from Copilot outputs and covers legal fees and damages, but it excludes free-tier users, intentional infringement, and patent, trademark and trade secret claims. As the dev.to analysis frames it, the commitment is a risk-transfer mechanism rather than a licensing regime: it shields buyers while compensating no creators.

The litigation pile-up

More than thirty major copyright lawsuits against AI companies are working through US federal courts, with plaintiffs including the Authors Guild, Getty Images, the New York Times and Universal Music Group. The dev.to post highlights two in particular: Getty's case against Stability AI, which it says seeks $1.8 trillion in damages over 12 million scraped images, and the New York Times suit against OpenAI and Microsoft, which alleges millions of articles were copied without permission. The core question across these cases is whether fair use shields training at scale.

The vice president's statement matters here because, in the post's framing, it is an admission against interest. Plaintiffs can cite it to argue that Microsoft understood the practice to be unlawful, which could support findings of willfulness — a factor that raises potential damages. The post also points to knock-on effects: rising insurance costs and investors demanding clearer risk assessments, with hypothetical damages in the trillions making the calculation existential for vendors.

Transparency rules complicate the picture

Regulation is converging on the same problem from another direction. According to the dev.to article, the EU AI Act took effect in August 2026, and its Article 53 requires general-purpose AI providers to publish a detailed summary of training data for any model placed on the EU market — a mandate that binds Microsoft, OpenAI, Google and Anthropic alike. Providers cannot simply name Common Crawl; the post argues they must explain what the dataset contains and whether rights were cleared. That collides with trade secrecy norms, since companies treat training data composition as a competitive advantage, and abandoning the EU market is not a realistic option.

Creators are no longer fragmented

The human side of the story is quantified in a survey by the Authors Guild and the Creators' Rights Alliance, cited by dev.to: of more than 5,000 writers, artists, musicians and photographers polled in 2025, 78 percent said their work had been used in AI training without authorization. The post describes a creative class that now coordinates across disciplines, funds litigation, testifies before Congress and engages with the EU regulatory process. The "labor theft" framing resonates, the author argues, because it centers human effort: models ingest the accumulated output of careers, learn style, voice and technique, and then compete with the very people whose work formed them.

Why it matters

An internal acknowledgment from a senior executive at one of the world's largest AI vendors that scraping amounts to historic theft is evidence, not advocacy. It strengthens plaintiffs arguing willfulness in the pending US cases, pressures Microsoft's own fair-use bet, and sits uneasily beside an indemnification program that absorbs risk without paying creators. Combined with EU transparency mandates and increasingly organized rightsholders, it shifts leverage toward licensing negotiations. And as the dev.to piece notes, the contradiction is not Microsoft's alone: every major AI player operates on the same foundation, so the eventual reckoning will be industry-wide.

  • #microsoft
  • #copyright
  • #ai-training-data
  • #litigation
  • #eu-ai-act

Related posts