Technology

Tech Giants Bet $300 Million on Shared Biology Data for a Virtual Cell

Martin HollowayPublished 53m ago4 min readBased on 9 sources
Reading level
Tech Giants Bet $300 Million on Shared Biology Data for a Virtual Cell
source:chanzuckerberg.com

Google DeepMind, Meta and Isomorphic Labs are putting $300 million into Biohub as part of a $1.8 billion initiative to build AI datasets for biological research. The Verge

Biohub is the nonprofit biomedical research organization founded by Mark Zuckerberg and Priscilla Chan in 2016. Its stated goal is to help fight disease by building a "virtual cell," a computer model researchers can use to run simulations.

The $300 million from the three technology contributors sits inside the larger $1.8 billion effort. The focus is datasets for biology, not a single model or product release.

A public-private funding stack

The U.S. Department of Energy will invest more than $500 million over the next five years to support the Biohub virtual-cell effort. The National Institutes of Health will contribute datasets, repositories and knowledge bases built with more than $500 million in earlier federal investment. The Verge

That structure combines new private money with earlier public spending. NIH is bringing existing data assets from that prior investment, alongside DOE funding scheduled across five years.

Biohub itself has committed $500 million to the Virtual Biology Initiative. Biohub

The Virtual Biology Initiative is building open, global datasets to power predictive models of the cell. In Biohub's description, the emphasis is on open datasets as shared infrastructure, rather than private collections held inside one lab or company.

Priscilla Chan is co-founder and co-CEO of the Chan Zuckerberg Initiative. She earned a bachelor's degree in biology from Harvard University and earned her doctor of medicine from the University of California, San Francisco, where she also completed her pediatrics residency. Mark Zuckerberg is co-founder and co-CEO of the Chan Zuckerberg Initiative. He is founder and CEO of Meta and studied computer science at Harvard University before moving to Palo Alto, California, in 2004.

Datasets as the constraint

The initiative centers on AI datasets for biological research, with predictive models of the cell as the later output.

Cell modeling is often limited by data, not by algorithms. Perturbation data, or records of what happens when a cell is disturbed, plus imaging, transcriptomics (measurements of gene activity) and proteomics (measurements of proteins), are split across different repositories, test methods and cell types. Models trained on a narrow slice of that data often perform poorly on new cell types.

Biohub previously unveiled an AI world model for drug discovery, a system designed to learn general patterns it can apply to new tasks. Reuters That world model comprises open-source AI models.

A separate CZI and NVIDIA collaboration addresses the data-processing layer. On Oct. 28, 2025, CZI published a Technology article titled "CZI and NVIDIA Accelerate Virtual Cell Model Development for Scientific Discovery". The collaboration aims to scale biological data processing to petabytes of data spanning billions of cellular observations to enable next-generation model development. On the same date, CZI's newsroom listed R&D World coverage titled "CZI, NVIDIA Expand Virtual Cell Push With Open Models and Benchmarks".

At that scale, curation, harmonization, metadata discipline and benchmark design become central. That means cleaning data, aligning formats across labs, keeping careful records of how each dataset was made, and designing tests that measure biological validity and not just held-out prediction loss, or performance on data withheld during training.

Commercial models meet open infrastructure

Isomorphic Labs comes to this effort from the applied side. It is an artificial intelligence-driven drug discovery company. It is backed by Google. Reuters

Isomorphic Labs raised $2.1 billion to scale AI-driven drug discovery. That raise, reported in May, was described around scaling discovery operations, separate from the Biohub dataset commitment now being made jointly with Google DeepMind and Meta.

The broader context here comes from earlier platform shifts. Open datasets and open-source models lower entry costs for academic labs and startups, while well-capitalized companies build proprietary pipelines on top. In my view, this is the most useful way to read the $300 million commitment. It does not erase the distinction between open science and commercial drug development. It funds the layer both depend on.

For teams that will actually use these resources, three practical questions matter more than the headline total. The first is licensing and access terms for the global datasets. Open in name can still mean friction in practice if access controls, use restrictions or compute requirements limit reuse. The second is benchmark governance. Predictive models of the cell will need independent, versioned benchmarks tied to experimental validation. The third is durability. A five-year DOE tranche and a $500 million Biohub commitment provide runway, while long-term maintenance of repositories and knowledge bases is a separate problem from initial dataset creation.

One risk worth flagging is coordination. The combination of DOE, NIH, Biohub, Google DeepMind, Meta and Isomorphic Labs puts unusual technical and institutional weight behind one idea, the virtual cell as a simulatable object. That weight helps with standards and scale. It also raises coordination costs, because harmonizing data across federal repositories, nonprofit programs and industry labs is organizational work before it is modeling work.

I have seen this pattern before while watching my kids grow up as the consumer internet turned mobile and then cloud-native. Each transition looked obvious only in retrospect. The unglamorous part held steady. Infrastructure was built quietly, then new applications arrived quickly. If the virtual cell effort succeeds on these terms, researchers could test ideas in computer simulations before committing to expensive lab experiments, and drug discovery teams could iterate faster because the underlying data layer is shared and maintained.