Signal

Meta, Google DeepMind, and Isomorphic Labs invest $300M and the US invests $500M+ in Zuckerberg-backed Biohub to build open biology datasets for AI training

First reported by Reuters ·

The signal ●●●● Compiled by AI from Reuters, Techmeme, Biohub, Axios, Quartz and 2 more
Why you might care

Researchers can now access a vast, open repository of AI-ready biological data, reducing the cost and time required for biological modeling and experimentation.

What happened

Biohub, in partnership with the U.S. Department of Energy (DOE) and the National Institutes of Health (NIH), is launching an expanded international effort to create accessible, AI-ready biological datasets. Announced on October 7, 2026, the initiative involves a total investment of $1.8 billion in funding, data, computation, and measurement technology. The DOE will contribute over $500 million through its Genesis Mission, focusing on lab measurements, modeling, and computation using exascale supercomputing and advanced imaging techniques. The NIH will leverage over $500 million in prior federal investment, standardizing existing datasets for AI training through its Bio Genesis Mission, aiming to double the pace of biomedical innovation. Additionally, Google DeepMind, Isomorphic Labs, and Meta are collectively investing $300 million to develop the necessary technologies and multi-modal datasets. This coordinated effort aims to build predictive AI models of biology, ultimately accelerating scientific discovery and the development of treatments for human diseases by enabling digital experimentation and the creation of a virtual cell.

What it means

The significant public and private investment in open biological datasets signals a strategic pivot towards democratizing AI research in life sciences. By pooling resources from major tech players like Google DeepMind, Meta, and Isomorphic Labs alongside government bodies like the DOE and NIH, the initiative aims to overcome data silos and proprietary limitations that have historically hindered rapid progress. This move not only accelerates the development of predictive AI models for biology but also sets a precedent for future large-scale, collaborative scientific endeavors in the age of AI, emphasizing data standardization and accessibility.

This initiative directly impacts the speed and scope of AI-driven drug discovery and disease research, enabling the creation of more accurate virtual cell models and accelerating the translation of scientific insights into clinical applications. The emphasis on open data commons means that smaller research institutions and individual scientists will have greater access to high-quality datasets, potentially leveling the playing field in biomedical innovation. The coordinated effort, involving advanced measurement technologies and exascale computing, suggests a future where complex biological systems can be understood and manipulated digitally, transforming the timelines for medical breakthroughs.

AI-written summary. May contain errors.