The first grants pay for drug-absorption benchmarks and, at UNC, for measuring the targets on cancer tumors directly instead of predicting them from sequencing.
The OpenAI Foundation has announced Public Data for Health, committing more than $125 million in first grants to universities and nonprofits to build scientific datasets and make them widely available to researchers.
The foundation says some valuable datasets may never be created or shared because individual institutions lack the incentives or the resources.
One award goes to OpenADMET, for open data and blinded contests testing whether models can forecast how small molecules move through the body.
At the University of North Carolina, the money will pay for direct measurements of tumor targets and immune responses across hundreds of tumors and several types of cancer, with the results published in de-identified form.
A $500,000 grant to CTD Commons will test whether records from failed or shelved drug programs can be saved before the companies close, and made publicly available.
A CTD brings together a drug's development records, including animal toxicology, manufacturing details and exchanges with the FDA, only a sliver of which ever reaches published papers. The foundation says the files could show developers what regulators required of similar drugs.
Last year, clinical-trials policy analyst Ruxandra Teslo proposed bidding for such records, normally treated as trade secrets, in biotech bankruptcy proceedings. The grant went to 1Day Sooner, a clinical-trial volunteer advocacy group she advises.
Josh Morrison, who cofounded the group and is its president, told MIT Technology Review that non-exclusive copies of a company's datasets might be obtained for a few tens of thousands of dollars each. It holds three so far, two donated by Lumen Bioscience, a biotech that had used Chapter 11 to study a rival's drug development. Two bids this year were turned down.
About 70% of the money and time in drug development goes into the clinical stage, Teslo said, and that part is close to a black box, especially for small biotechs. She hopes the files could help train AI systems to guide developers through approval.