The $1.8B bet that biology's bottleneck is data, not models

Share
The $1.8B bet that biology's bottleneck is data, not models

The biggest AI-for-science announcement in months contains no model release, no benchmark, and no demo — just a very large pile of money pointed squarely at the least glamorous part of the pipeline: the raw measurements a model would train on.

What happened

On October 7, Biohub — the nonprofit research institute backed by Mark Zuckerberg and Priscilla Chan — announced an expansion of its Virtual Biology Initiative alongside the U.S. Department of Energy, the National Institutes of Health, Google DeepMind, Isomorphic Labs and Meta. The headline number is $1.8 billion in combined funding, data, computation and measurement technology, which Biohub calls the largest coordinated commitment to generating AI-ready biological data to date. The stated deliverable is a "universal virtual cell": an open resource of standardized, machine-readable measurements that lets researchers ask biological questions digitally, and eventually predict how any cell responds to an intervention.

The money breaks down into four pieces. The DOE will invest more than $500 million over five years through its Genesis Mission, contributing exascale computing, cryo-electron microscopy and tomography, and autonomous laboratories across the National Lab system. NIH is not writing a new check — it is coordinating datasets, repositories and knowledge bases built on more than $500 million of prior federal investment, which Biohub will standardize for model training. Google DeepMind, Isomorphic Labs and Meta are collectively putting in $300 million. And Biohub's own founding $500 million commitment from April anchors the effort: $400 million of it funds measurement technology — near-atomic imaging inside living cells, microscopy that images millions to billions of cells — with $100 million going to research outside Biohub.

The actual product is standardization

Read past the dollar figures and the announcement is about plumbing: shared standards, common identifiers, a single point of access for datasets that currently sit in incompatible formats across institutions. The Allen Institute, Broad Institute, Gladstone Institutes, Human Cell Atlas, Human Protein Atlas and Wellcome Sanger Institute have signed on, NVIDIA is providing accelerated computing, and Renaissance Philanthropy is helping expand funding. This is the unglamorous work Biohub has actually done before — CELLxGENE, the CryoET Data Portal, Tabula Sapiens — now scaled into a cross-government coordination layer.

That framing matters because the field has spent two years discovering that data, not compute, sets the ceiling. DeepMind's AlphaGenome Atlas, which we covered in September, was an attempt to precompute predictions across the genome's nine billion single-letter variants — a map built from inference, not experiment. This initiative is the complementary move: paying for the experiments whose ground truth no amount of inference can manufacture. As Google DeepMind's Pushmeet Kohli put it, the virtual cell cannot be solved "without open, experimental biological data at an unprecedented scale."

It also fits a pattern we have watched play out in other fields this month: the frontier labs are increasingly publishing their biggest scientific output rather than productizing it — from OpenAI's math manuscript dump to this data commons, covered in 722 math manuscripts from a model nobody can run. The competition has moved upstream, to who defines the substrate everyone else works on.

Who wins, who waits

Biohub wins stewardship. Whoever owns the standards layer owns the ecosystem — the institute already runs the community infrastructure most of this data will flow through. Google DeepMind and Isomorphic Labs win leverage: for $300 million spread across three companies, they help set the format of a field's training data while footing a fraction of the bill. Isomorphic president Max Jaderberg was blunt about the motive, saying the initiative exists to scale "past the limits of what any single organization can produce today" — the same logic that made open-source datasets the smart money's favorite subsidy. The DOE and NIH win renewed relevance for facilities and repositories built long before AI budgets existed. And the global research community wins access: this is explicitly an open resource, not a proprietary corpus.

Who doesn't win: the vendors whose business is selling proprietary biological data, and anyone hoping the announcement said something about model access. It doesn't. The press release promises open data; it says nothing about whether the models trained on that data — Isomorphic's among them — will themselves be open. Data commons, closed weights, is now the standard arrangement.

The skeptic's case

Three honest objections. First, the money is announced, not disbursed — DOE's $500 million is a five-year plan spanning budget cycles, and NIH is re-labeling existing investment rather than expanding it. Second, standardization is a coordination problem, and coordination problems are exactly what big cheques solve worst; the press release concedes as much by devoting a section to "convening researchers across institutions." Third, and most substantive: nobody yet knows whether biology scales like language models do. Biohub's head of science, Alex Rives, told Axios that when the initiative launched in April, the biggest unanswered question was whether cellular biology exhibits scaling laws at all — and that the answer should come within a year of the first large-scale dataset landing.

The track record on that question is mixed enough to respect the doubt. As we noted in Vivodyne opens the world's largest 'human data center' to fix AI drug discovery, a Nature Methods study found no clear data scaling laws when training on existing cellular data — and Isomorphic Labs, now a $300 million investor in generating better data, is only reaching its first clinical trials a year behind its own schedule. If a year from now the standardized datasets exist and models trained on them measurably improve, this announcement was the Manhattan Project of cell biology. If they don't, it will be remembered as the most expensive formatting exercise in science.

What to watch

The concrete test Rives set: within a year of the first large-scale standardized dataset, researchers should be able to train models, measure capabilities, and determine which kinds of biological data actually make them better. Watch for that readout before judging the bet — and watch whether any model trained on the commons ships with its weights. Open data is the promise; open capability is still the question.

If public money builds the data, should the models trained on it be forced open? Tell us in the comments.

Read more

The Take — An AI FDA would approve the demo, not the deployment

The Take — An AI FDA would approve the demo, not the deployment

I think Geoffrey Hinton has the diagnosis right and the prescription backwards. Self-policing genuinely has no burden of proof — nobody outside a lab must be shown anything before a frontier model ships — and that has to change. But an FDA-style pre-market gate would certify the wrong object: a frozen snapshot presented on approval day, when every failure we have actually watched arrive did so afterward, through permissions, updates and tool access. Our morning brief laid out the ask from Tues

Google launches Playground: prompt your own games, no code

Google launches Playground: prompt your own games, no code

Google turned game-making into a chat window today, while Anthropic opened the most permissive tier of its cyber program to government-vetted defenders and put its AI watermark detector in front of everyone. Google launched Playground, a browser platform where anyone describes a game in text prompts and plays it — no coding, live today in the US for users 18 and over. Creation works from a blank canvas or starter prompts: you type what you want, then tweak physics, rewrite rules, or swap chara

Musk rules TSMC out of Terafab: 'we will build and run the fab'

Musk rules TSMC out of Terafab: 'we will build and run the fab'

Chipmaking, a government-ordered construction halt, and a $10 billion fund — three moves that all trace back to who controls AI's physical layer. Musk says his companies will build and run the Terafab chip complex themselves, explicitly shutting out TSMC. In a post on X on October 7, he left little room for interpretation: "No, we will build and run the fab. Let there be ZERO doubt about that." — adding that "maybe TSMC subleases part of the Terafab if they want, but nothing more than that." T

Common Sense Media calls ChatGPT for Teens an 'unacceptable risk'

Common Sense Media calls ChatGPT for Teens an 'unacceptable risk'

A watchdog's tests say ChatGPT for Teens fails exactly where parents were promised it would hold — and OpenAI is contesting the methodology, not the stakes. Common Sense Media has rated OpenAI's ChatGPT for Teens an "unacceptable risk," making it the sharpest public challenge yet to the safety case OpenAI built around younger users. The nonprofit's Youth AI Safety Institute ran more than 4,000 prompts against accounts registered to 13-to-17-year-olds and found that the teen experience doesn't