Microsoft Executive Calls AI Training "Major Theft of Labour in Human History"

In Short

A Microsoft executive's leaked remarks call AI data scraping an unprecedented theft, reigniting the copyright battle between publishers and tech giants.

Chatbots like ChatGPT, Claude and Gemini owe their vast knowledge to enormous quantities of text pulled from across the internet
X

Chatbots like ChatGPT, Claude and Gemini owe their vast knowledge to enormous quantities of text pulled from across the internet

Font size
FOLLOW ON Google News

Not everyone inside the world's biggest AI companies is comfortable with how their models learned to be so smart. As it turns out, some of the harshest criticism of the industry's data-scraping practices has come from within its own ranks - and it's now surfacing in court.

Chatbots like ChatGPT, Claude and Gemini owe their vast knowledge to enormous quantities of text pulled from across the internet. The catch is that much of that material wasn't necessarily up for grabs. Authors, publishers and news organisations have repeatedly accused OpenAI, Anthropic and other AI labs of using their work without asking - or paying. Now it appears some insiders agree. A Microsoft executive reportedly described the practice as the "largest theft of labour in human history."

The remark belongs to Brent Hecht, Microsoft's director of applied science, and surfaced in newly unsealed court filings tied to the copyright lawsuit brought by The New York Times and other news outlets against Microsoft and OpenAI. Discussing the use of online content to train AI systems, Hecht didn't mince words. "Millions of people around the world will soon consider large models 'hoovering up' all their work to be an astonishing theft of unprecedented proportions," he wrote. He added that "almost no one intended for content they created to be used in this fashion, nor are they compensated for its use."

Microsoft and OpenAI, for their part, continue to lean on a fair-use defence, arguing that training AI models on copyrighted material transforms the content rather than reproducing it outright. The news organisations disagree, insisting that AI products end up substituting for the very journalism they were trained on.

That substitution concern runs through much of the filing. Rather than clicking through to a news site, users increasingly just ask a chatbot and get their answer instantly. Even Microsoft CEO Satya Nadella acknowledged this shift in his deposition, saying that talking to chatbots has "substituted giving you the information right there on the website on the AI platform versus needing to go to the underlying source."

The implications for publishers, the news organisations argue, are troubling - AI models are trained on journalism, yet the resulting chatbots may reduce the traffic that journalism depends on. The filing points to Microsoft's own data showing click-through rates for The New York Times and Daily News domains were 83 to 93 percent lower on Copilot's "answer engine" compared to traditional Bing Search. It also cites OpenAI's ChatGPT head Nick Turley, who called the products "largely substitutive" and predicted they would grow "more and more substitutive as they get better."

Elsewhere in the filings, OpenAI co-founder Greg Brockman is quoted praising how well the models handled news content, noting they were "particularly good at predicting text of news articles" and "very good at any news task."

The New York Times filed its lawsuit against Microsoft and OpenAI in 2023, roughly a year after ChatGPT's debut. Since then, other outlets - including the New York Daily News, The Intercept and the Center for Investigative Reporting - have joined the case. While Microsoft and OpenAI continue to stand by their fair-use argument, the news plaintiffs are now pushing the court for a ruling in their favour on the copyright claims.

Kahekashan is a passionate technophile with a keen eye for cutting-edge gadgets, emerging technologies, and everything in the digital realm. Raised in a Defence family with strong values and a background in literature, she has consistently pursued excellence in every endeavour. Her last full-time assignment involved content writing with the Indian School of Business.

Next Story
Share it