Unsealed NYT filings: Microsoft called AI scraping 'theft of labor'
Two stories today about AI meeting records that already exist — one in a federal courtroom, one inside the UN's statistics division.
Newly unredacted filings in The New York Times' three-year copyright suit show Microsoft's own leadership privately described AI training on scraped content as theft, while OpenAI's executives conceded the product competes with the publishers who supplied it. In a January 2023 internal memo, Brent Hecht, Microsoft's director of applied science, called the practice "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history," according to the TechCrunch read of the filing. A January 2024 Microsoft presentation by Hecht described a "doom loop" in which the company's own Copilot answer engine cut click-through rates to the Times' domain by as much as 93% versus traditional Bing search — a decline the document said would "hurt the performance of our models and the entire web at the same time." On the OpenAI side, head of ChatGPT Nick Turley wrote internally that publishers face an "existential threat" from a product that is "largely substitutive" and "will get more and more substitutive as they get better"; president Greg Brockman described the models as "excellent at news."
The operational detail is worse for the fair-use defense than the quotes. The filing describes OpenAI delivering the entire GPT-3 training dataset to Microsoft, training data flowing back through initiatives called Project Taxi and Project Mango, and a Mango-derived dataset containing copies of at least 160,903 unique works from the news publishers. It says OpenAI employees planned to circumvent paywalls without detection — researcher Nick Ryder messaged Brockman about a "hack to get around nytimes paywall," and Brockman replied "ah nice" — and that researchers deliberately stripped copyright notices from training data because they "wouldn't want model outputting" them. Microsoft CEO Satya Nadella testified that paywalled material "should be licensed by anyone who wants to use it" and that he would have required OpenAI to retrain had he known. Fair use turns on whether the use substitutes for the original and harms its market; internal documents conceding substitution, market harm, and knowledge of the paywall are not what a defendant wants unsealed. The caveat that matters: much of the new material comes from the Times' own brief, not the underlying exhibits, which remain sealed — and quotes arrive without their original context. Microsoft and OpenAI did not comment.
The United Nations is handing Google the job of making the world body's statistics readable by AI agents, and the numbers behind the decision are an indictment of how badly models handle authoritative data today. A UNICEF benchmark run across more than 133,000 responses on global development indicators — covering GPT-4o, GPT-4o-mini, Claude Sonnet 4.5 and Haiku 4.5, and Gemini 2.5 Flash and 2.0 Flash — produced an average accuracy score of 21.2%, the agency's chief statistician told reporters. About three in five answers contained no usable number at all, usually because the model hedged; when the same questions were re-run on the same model versions two days later, the models that gave a number both times agreed with themselves only about half the time. Consistent wrongness would at least be predictable — this is noise wearing a confidence interval.
The fix is structural: 26 UN entities have committed to a UN-governed instance of Data Commons, with data from nearly 20 available at launch, and the goal is 80% of the UN system's statistical datasets on the platform by 2027. Google.org put in $2 million in capacity-building funding and technical support, and the platform keeps provenance for each statistic so an answer can be traced back to its source. Google demonstrated an agent pulling multiple indicators through MCP to generate charts and written analysis without a human assembling the datasets. The honest caveat came from Google's own Data Commons lead: giving a model authoritative data does not make its conclusions authoritative, and a human should review outputs before citing them. The demand is already there — UNICEF says AI assistants now account for roughly one in 10 visits to its data site, with ChatGPT referrals up 67% year over year.
What to watch: whether the Times' unredacted material survives a sealing fight if Microsoft and OpenAI move to strike it, and whether any UN agency publishes an accuracy re-test once its data is agent-readable.
If a model can cite its source but still mangle the number, is provenance enough — or does the UN need to publish model accuracy scores alongside the data? Tell us in the comments.
Sources: TechCrunch — Microsoft exec called AI scraping 'the largest theft of labor in human history' · Techmeme — NYT court filing coverage · TechCrunch — UN turns to Google to make its global data ready for AI agents · Google — Google and UN system launch new global data platform · UN DESA — UN Data: AI-ready gateway to trusted UN System data