Snorkel AI triples to $3.5B as labs buy environments, not labels
The day's money went to the unglamorous half of AI: the datasets and simulated environments that make a model trainable, and the diplomats who want a say in what it is allowed to do.
Snorkel AI raised a $350 million Series E at a $3.5 billion valuation, led by Insight Partners and S32 — nearly triple the $1.3 billion it was worth when it raised $100 million seventeen months ago. The company says its annualized revenue run-rate has reached $375 million, an eighteen-fold increase in twelve months, and the number is the interesting part. Snorkel sells reinforcement-learning environments and finished datasets rather than rented human hours, so what it pays its experts sits in cost of goods sold instead of being booked as headline revenue — the accounting distinction between a data vendor and a staffing agency dressed as one.
CEO Alex Ratner's framing for the round is what he calls Data 2.0: the first mile of AI data was volume, and the last mile is a curriculum of hard problems that a senior engineer would need days to solve, built to resist models that would rather cheat the environment than solve it. Snorkel's answer is a loop it calls an RSI engine for data — hundreds of specialized agents running quality control alongside human reviewers, with the humans' feedback then used as weak supervision to improve the agents. The company reports those agents lift QC efficiency by more than 50 percent and review accuracy by 15-plus points against a human-plus-off-the-shelf-LLM baseline, and that scaled human supervision delivers a 2x accuracy gain over a general frontier model.
Those figures are self-reported and unaudited, which matters when the whole pitch is that this work is too hard for humans alone. But the direction of the market is not in doubt: every lab shipping a cheaper frontier model this month — OpenAI's half-price GPT-6 Sol and Luna, Anthropic's Opus 5.5 — is buying post-training data and reward environments at a rate that shows up as revenue here. Snorkel also said it will expand its Open Benchmarks Grants program, the funding behind Terminal-Bench, OSWorld 2.0 and Agents' Last Exam, and will push research money toward data and environments for alignment and safety.
Sam Altman, Dario Amodei, Yoshua Bengio and Hugging Face's Clément Delangue will brief the UN Security Council on AI on Wednesday, in a session France initiated under its rotating presidency. Amodei will attend remotely, according to Bloomberg. The French concept note puts the agenda on the malicious use of AI and the risk of losing control of advanced models — and DeepSeek and Moonshot AI have also been invited to make statements, which is the detail that makes this more than a photo opportunity.
Why it matters: twenty-four hours after President Trump used the General Assembly to reject any globalist scheme to control AI — we covered that address and its counter-positions — the same building hosts the opposite premise, with American lab chiefs, a UN panel co-chair and two Chinese labs in the same room two days before Trump meets Xi Jinping. Nothing binding comes out of a Security Council briefing. What comes out is a public record of what each party was willing to say on it.
A one-prompt benchmark site called LLM AssBench hit the Hacker News front page by asking twenty frontier models to build the same interactive 3D figure in a single HTML file, and publishing the results side by side, date-stamped. Each entry is a self-contained Three.js scene you can rotate in the browser — Claude Opus 5.5 and GPT-6 Sol are both on it, hours after their launches, next to Fable 5.1, Gemini 3.1 Pro and Grok 4.6.
It is not a benchmark in any statistical sense: one prompt, one run per model, no scoring rubric. It is still a more honest read on coding ability than most launch-day charts, because you can look at the output and judge it yourself instead of trusting a vendor's table. Visual tasks are also the ones where a model cannot hide behind a plausible-sounding paragraph.
What to watch: whether the Security Council session produces anything beyond statements, and whether Snorkel's Data 2.0 economics survive contact with the labs building their data engines in-house.
If you can rotate the output and judge it yourself, does a benchmark table still tell you anything — or is visual verification the only evaluation that hasn't been gamed yet? Tell us in the comments.
Sources: TechCrunch · Snorkel AI — Data 2.0 and the research era of AI data · Unite.AI · Bloomberg · TNW · Quartz · LLM AssBench · Hacker News discussion