AI data startup Micro1 hits $500M run rate
Two signals this morning point at the same pressure point in the AI economy: the scramble for training data has become a billion-dollar business of its own.
Micro1, a four-year-old AI training-data startup, grew its gross annual run rate from $100 million to $500 million in just eight months, according to a person familiar with the company. The startup keeps roughly 60–70% of that after paying its contract annotators, putting its net run rate between $150 million and $200 million — real money, though still behind Mercor (about $2 billion gross this summer) and Handshake ($1 billion earlier this year).
The boom is a direct side effect of the frontier labs' hunger for unique, high-quality training data: Micro1 hires domain experts — doctors, lawyers, scientists — on contract to label and evaluate model outputs, and it is increasingly generating synthetic data with no human in the loop, which it can resell to multiple customers at 80–90% gross margins. That resale model has sparked controversy, with critics arguing off-the-shelf datasets handed to Chinese AI developers help close the gap with US models; founder Ali Ansari says Micro1 does not sell to Chinese model makers. We covered the demand surge from the other side earlier this month — Nvidia pursues a $20B stake in Mercor shows how seriously the chip giant is taking the data-supply chain.
ChatGPT's web search has quietly started leaning hard on the site: operator, according to tracking from the optimization firm Promptwatch. The share of ChatGPT Search queries that restrict results to a specific domain jumped from a steady 0.3–0.5% to 16–17% on August 8, right as OpenAI rolled out GPT-5.6's "Sol" search mode for Plus and Pro users, which it said would give "more focused answers."
Simon Willison, who surfaced the data, notes the shift lines up with OpenAI's vague Aug 6 announcement rather than any public operator change — and that Promptwatch also saw ChatGPT lean far less on Reddit as a source around Aug 18. For publishers and the emerging "GEO" (generative-engine-optimization) industry, it's a clear signal that where ChatGPT looks for answers is now a moving target set largely behind closed system prompts.
What to watch: whether OpenAI formalizes domain-restricted search into a publisher-facing control, or keeps tuning it in the dark.
Is the training-data bottleneck about to mint a new tier of billion-dollar suppliers — or will synthetic data make the whole category obsolete? Tell us in the comments.
Sources: TechCrunch · Simon Willison · Promptwatch