Biohub, the U.S. Department of Energy, the National Institutes of Health, Google DeepMind, Isomorphic Labs and Meta have expanded the Virtual Biology Initiative, an effort to generate biology data suitable for artificial intelligence models. Announced on October 7, 2026, the collaboration represents a coordinated commitment of $1.8 billion in funding, existing data resources, computing capacity and measurement technology.
The initiative aims to measure how cells respond to a much wider range of interventions and conditions, then standardize those observations for models that attempt to predict biological processes. Biohub says the resulting resources are intended to form an open foundation for the research community. The announcement does not present a validated treatment or a clinical result; it funds data generation and model-development infrastructure.
How the $1.8 billion total is assembled
Biohub’s founding $500 million commitment, announced in April, anchors the expanded program. According to the organization, $400 million supports new biological measurement technologies and $100 million funds research outside Biohub. The Department of Energy will contribute more than $500 million over five years for cell research, data collection, imaging, modeling and computation.
The National Institutes of Health will coordinate relevant datasets, repositories and knowledge bases built through more than $500 million in earlier federal investment. That portion is described as the alignment of resources created through prior spending, rather than a new cash commitment in this announcement. Google DeepMind, Isomorphic Labs and Meta are jointly investing another $300 million.
Standardized data for a “virtual cell”
The program is intended to make data produced by different institutions work together through shared standards and identifiers. Capabilities in the Department of Energy’s national laboratories—including supercomputers, X-ray and neutron scattering facilities, cryo-electron microscopy and autonomous laboratories—will provide measurement and computing resources. NIH will connect national biomedical repositories and research programs already developing common data standards and AI-ready resources.
Reuters reported that the first dataset is expected within about a year and that functional predictive models are targeted within five years. Those dates are goals, not evidence that the models will reliably predict disease biology. The complexity of cellular systems and the challenge of harmonizing measurements across laboratories remain central scientific and technical tests for the project.
Open-science goals and access questions
Biohub describes the planned output as an open resource for the research community. Reuters and Axios reported that commercial funding partners will receive an initial period of access before the data becomes public. Because the official announcement does not spell out every access condition, licensing details and the practical length of any early-access period will be important to monitor as datasets are released.
Scientific organizations including the Allen Institute, Broad Institute, Human Cell Atlas, Human Protein Atlas and Wellcome Sanger Institute are also participating in coordination. NVIDIA will support the effort with accelerated computing infrastructure and technical expertise. The initiative’s value will therefore depend on more than the headline funding total. Data quality, openness, interoperability across laboratories and validation of model predictions against experiments will determine whether the shared infrastructure produces reliable scientific tools.
If successful, the initiative could let researchers prioritize some experiments digitally before moving into the laboratory. That remains a research objective rather than a demonstrated outcome. The next concrete milestones will be the publication of the first datasets, clear access terms and evidence that models trained on the material can reproduce or predict biological observations outside the datasets used to build them.
