Microsoft exec said AI scraping was ‘the largest theft of labor in human history,’ lawsuit filings reveal

Scraping copyrighted content is still being debated

by · TechRadar

News By Craig Hale Published 18 September 2026

(Image credit: TestPlanet)

Share this article 0 Join the conversation Follow us Add us as a preferred source on Google Newsletter Subscribe to our newsletter


  • Microsoft exec describes AI scraping as "astonishing theft" of labor
  • Copilot reportedly reduced NYT click-throughs by 93% compared with Bing
  • Trump administration supports scraping so that the US can retain its AI dominance

In an earlier January 2023 memo uncovered in recently unredacted court filings, Microsoft Director of Applied Science Brent Hecht accused AI scraping of being an "astonishing theft of unprecedented proportions" (via TechCrunch).

With generative AI posing a "real risk" of disrupting the employment of the same people who unwillingly provided that training data, Hecht went on to describe scraping as "the largest theft of labor in human history."

This comes from a 2023 New York Times case against Microsoft and OpenAI, when the publication accused the AI giants of copying and using its copyrighted work without permission.

Latest Videos FromTechRadarWatch full video here:

AI scraping described as mass theft

Microsoft's own internal research reportedly found that Copilot could reduce click-throughs to The New York Times by as much as 93% compared with its own conventional search engine, Bing.

Ultimately, the company has fundamentally admitted that LLMs could undermine the foundations of their own content supply chain, with Hecht introducing the theory of a "doom loop."

Importantly, the company's CEO, Satya Nadella, even admitted that had he known OpenAI was scraping data found behind paywalls, he would have exercised Microsoft's rights to require OpenAI to retrain its models, excluding this protected data.

As for OpenAI, the ChatGPT maker allegedly removed copyright notices from material before it entered the large datasets so that the models wouldn't re-generate those copyright notices for users.

Are you a pro? Subscribe to our newsletter

Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed!

Contact me with news and offers from other Future brandsReceive email from us on behalf of our trusted partners or sponsors