Mark Zuckerberg's nonprofit Biohub just corralled $1.8 billion from U.S. taxpayers and Big Tech to build AI that can simulate living cells — and the companies putting up a fraction of the cost get a full year of exclusive access to the biological data before the public sees a byte of it.
The Virtual Biology Initiative is the largest coordinated push yet to generate AI-ready biological data. The ambition: model cells so accurately that scientists can run experiments digitally instead of in a lab. The problem: the money trail shows American citizens funding the bulk of it while private giants who censored them reap the first-mover advantage.
Here is who is paying what. The U.S. Department of Energy committed more than $500 million over five years for lab measurement, modeling, and computing, drawing on exascale supercomputers and X-ray and neutron scattering through its Genesis Mission program. The National Institutes of Health is contributing datasets and repositories built with more than $500 million in prior federal funding. Biohub itself pledged $500 million back in April — $400 million for new cell-measurement tools and $100 million for outside research.
And the private sector? Google DeepMind, Isomorphic Labs, and Meta combined are chipping in just $300 million. Nvidia is providing computing and software. Renaissance Philanthropy is helping raise more. The Allen Institute, Broad Institute, Gladstone Institutes, and the UK's Wellcome Sanger Institute are also participating, along with the Human Cell Atlas and Human Protein Atlas consortia.
That is over $1 billion in taxpayer money versus $300 million from the tech giants. Yet Biohub's head of science, Alex Rives, told Axios that commercial funders get one year of exclusive access to the data before it becomes public. "We have to have some incentive for commercial players to be a part of this, and the embargo period creates that," Rives said. Government-funded work will carry no such restriction, Rives told Reuters, but that distinction does little for the public when the data those companies are hoarding was built on a foundation of taxpayer dollars.
The scientific bet is straightforward. Current cell datasets cover hundreds of millions of cells. Biohub researchers say useful predictive models will require data on billions, eventually trillions, of cells and cellular states. "We need to capture the language of biology; we need to capture the language of the cell. And that doesn't exist today," Rives told Reuters. The initiative plans to collect data on how cells respond to drugs, genetic modifications, and environmental changes using spatial transcriptomics, advanced microscopy, cryo-electron tomography, and large-scale cellular screening.
TechStartups.com framed the announcement as a breakthrough for computational biology, emphasizing the promise of digital experiments that could unlock new disease understanding and cures. TNW was more forthcoming, reporting the exclusive-access embargo and noting that Biohub plans to approach drug companies and philanthropies next — meaning more private players could buy into the data head start.
The first dataset is expected in roughly a year. Accurate predictive models could emerge within five years, Biohub says.
The question no one at the press conference answered: if the public funds the science and the public provides the biological samples, who owns what the AI learns about your cells — and who profits from selling that knowledge back to you as medicine?








