The US government and Google joined a $1.8 billion push for AI biology data led by Mark Zuckerberg’s Biohub. The project uses massive datasets to build predictive models for drug discovery. Dr. Priscilla Chan said, "We have always held this as a community asset, not just for one group."

Meta, Google DeepMind and drug discovery startup Isomorphic Labs are jointly investing $300 million. The Department of Energy will invest more than $500 million over five years in laboratory measurement, modeling and computation. The National Institutes of ​Health will coordinate datasets and repositories built with more than $500 million in earlier federal funding, which Biohub will standardize ​for AI training.

Sign up here.

The commitments follow the $500 million that Biohub — a philanthropic venture of Meta CEO Mark ⁠Zuckerberg and his wife, Dr. Priscilla Chan — put into the project in April.

The investments fund the Virtual Biology Initiative, which ​aims to measure how cells respond to changes across far more conditions than scientists have so far studied, and use that ​data to build predictive models that could compress drug development timelines that currently take years.

"Biology has been just sort of a clever discovery-based science until this point," Chan said in an interview. "We have always held this as a community asset, not just for one group, so that ​it can build upon itself over time.”

The datasets will eventually be released publicly, but companies that fund them get a head ​start, according to Biohub's head of science Alex Rives.

"With commercial funders we have embargo periods where there's a period of time where the ‌groups ⁠can work on the data, and then it becomes available as a public scientific resource," Rives said.

The arrangement is how Biohub is drawing private money into a project it describes as open science. Rives said the government-funded work running in parallel will carry no such restrictions, and that Biohub plans to approach pharmaceutical companies and philanthropies next.

Current cell datasets run to hundreds of ​millions of cells, Rives said, ​while an accurate predictive model ⁠will require billions and eventually trillions. Biohub's goal is to close that gap.

"We need to capture the language of biology, we need to capture the language of the cell. And that ​doesn't exist today," he said.

The data will come from techniques including spatial transcriptomics, which maps ​molecular activity inside ⁠intact tissue, and screens that record how cells respond to changes in environment. Much of it has never been generated in a coordinated way.

Rives said the work would normally take decades and that the partners aim to compress it into five years, with a ⁠first dataset ​ready in about a year. He expects accurate predictive models within five ​years.

Other AI labs are pursuing similar initiatives. Anthropic doubled down on its biology efforts with a wet lab, while the OpenAI Foundation started a more than $125 million ​grant program to fund biological and medical datasets for AI research.

Reporting by Krystal Hu in San Francisco; Editing by Kevin Buckland