· via Hacker News – Front Page (native)
Biohub, DOE and NIH commit $1.8B to open AI-ready biological data
Biohub, the U.S. Department of Energy and the National Institutes of Health have committed $1.8 billion in funding, data and computation toward an open resource for training predictive AI models of biology.

A $1.8 billion push for AI-ready biology
Biohub, the U.S. Department of Energy and the National Institutes of Health have announced a $1.8 billion commitment to generate biological data that is purpose-built for AI training. Unveiled on October 7, 2026, the package spans funding, data, computation and new measurement technology, and Biohub describes it as the largest coordinated commitment to AI-ready biological data to date. Everything produced is meant to become an open resource for the research community.
The effort expands the Virtual Biology Initiative, first announced in April 2026. Its long-term goal is a predictive "virtual cell": a model accurate enough that researchers could ask a biological question, simulate an intervention and get a reliable answer digitally. Biohub's head of science, Alex Rives, called building a virtual cell "one of the most important challenges for the next era of science," arguing it demands data generation coordinated at a scale no single organization can manage.
Where the money comes from
According to the announcement, the total combines new funding with data and infrastructure contributed in kind:
- DOE will invest more than $500 million over five years through the Genesis Mission, a cross-agency program drawing on exascale supercomputing, X-ray and neutron scattering, cryo-electron microscopy and tomography, and autonomous laboratories across the National Laboratory system.
- NIH, through its Bio Genesis Mission, will fold in datasets, repositories and knowledge bases built with more than $500 million in prior federal investment, including repositories catalogued by the National Library of Medicine and the National Center for Biotechnology Information, plus Common Fund programs already assembling biological atlases and shared data standards. Biohub will work with NIH to standardize these datasets for model training.
- Biohub's founding $500 million anchors the initiative. Of that, $400 million supports measurement technology: cryo-electron tomography that resolves near-atomic detail inside cells, microscopy able to image millions to billions of cells in living tissue, and engineering tools for perturbing biology at molecular, cellular, tissue and whole-organism levels. The remaining $100 million funds research outside Biohub.
- Google DeepMind, Isomorphic Labs and Meta are together putting $300 million into the technologies and multimodal datasets behind the initiative.
Standards are the hard part
The announcement stresses that volume alone is not the deliverable. Biohub is building a unification layer of shared standards, common identifiers and a single point of access, so datasets produced on different instruments in different disciplines can be combined for training. It points to a decade of prior open-data projects, including the Tabula Sapiens cell atlas, OpenCell, Zebrahub and infrastructure such as CELLxGENE and the CryoET Data Portal, as the template for this coordination.
On the science side, the initiative plans to expand measurements of how cells respond to interventions across far more cell types and conditions than have been studied so far. A group of institutions experienced in running large international collaborations since the Human Genome Project era, including the Allen Institute, Broad Institute, Gladstone Institutes, Human Cell Atlas, Human Protein Atlas and Wellcome Sanger Institute, have joined to help organize the community. NVIDIA is contributing accelerated computing infrastructure, domain-specific software and technical expertise, while Renaissance Philanthropy is helping raise additional funding for data generation.
Why it matters
For AI in biology, the binding constraint is data rather than model architecture. Biological measurements are expensive, inconsistent across labs and rarely collected with training in mind, which is why virtual cell models have lagged language models trained on web-scale text. A large, standardized, openly available corpus of perturbation data, covering cells measured across many types and conditions, is the input such models have been missing.
The openness may be the most consequential detail. Google DeepMind, Isomorphic Labs and Meta are funding data that academic labs and rival companies will equally be able to train on, which could blunt data-moat dynamics and let smaller players compete on modeling. NIH's Nicole Kleinstreuer argued that cell models with enough biological complexity to predict how a cell responds to an intervention could shorten timelines to medical breakthroughs compared with laboratory experiments alone.
There are caveats worth keeping in view. The $1.8 billion figure blends new spending with previously funded NIH datasets and contributed computation, the initiative is only six months old and still taking shape, and multi-agency coordination has a long history of friction. But the strategic signal is clear: the next bottleneck in AI-driven biology is data generation, and a large coalition has just committed to producing it in the open.
- #ai
- #bioinformatics
- #open-data
- #research-funding
- #machine-learning