We built an agent harness that assembles evidence across the biomedical ecosystem and paired it with a model that converts that evidence into probabilities of clinical success, outperforming the leading academic models on trial prediction.
Clinical trials are fundamentally inference problems
Every clinical trial has an implicit forecast. The actions of companies and markets track an implied probability that a therapy will succeed. The fundamental question is given everything knowable about a program at a point in time, what is the probability that it advances?
Existing models estimate this probability largely from structured datasets. Yet the state of a therapeutic program is not contained in any single database. It is distributed across papers, conference abstracts, patents, regulatory filings, prior clinical programs, and the broader history of the target, disease, and drug. The problem has been one of practical infeasibility: the evidence exists, but assembling and interpreting it at scale has historically required prohibitive amounts of manual effort.
Agents change information collection
AI agents have made long-running, open-ended investigation tractable.
Instead of operating over a fixed dataset, agents can adaptively search across the biomedical ecosystem: navigating interfaces, discovering relevant sources, extracting findings, and following new lines of inquiry as evidence accumulates. Work that may have required weeks of manual diligence can now be performed in hours.
More importantly, information can also be synthesized effectively. Trial success depends on how diverse pieces of evidence combine into an integrated view of disease biology, pharmacology, clinical precedent, competitive programs, trial design, sponsor history, and regulatory context, among other factors. Agents are uniquely suited to constructing these higher-order representations by continuously placing new information into broader biomedical and regulatory contexts.
The result is a substantially richer representation of the evidence informative of drug development success.
Evidence to probability
To guide decisions, the resulting evidence must be concretized into quantitative estimates of trial success. Our approach uses agents to generate data to train a deep-learning model: agents generate evidence while a model learns the statistical relationship between that evidence and subsequent clinical outcomes.
Empirical performance
The objective is to estimate the probability that a clinical program transitions to the next stage of development, e.g., a phase 1 trial progressing to phase 2.
We evaluated the platform on historical trials using a sandboxed backtesting environment with data temporally restricted such that no knowledge of a trial's outcome could enter the evidence. The model was trained on a 100k trial corpus and evaluated on a held-out set of 10k trials.
Our model outperforms the leading academic baselines on AUROC and AUPRC scores across all phases.
It should be noted these comparisons are between models evaluated on different test sets, so we choose a randomized test set for evaluation.
Significance and further implications
Clinical development is fundamentally a problem of capital allocation under extreme uncertainty. Biopharma must decide which programs to fund, which trials to run, which assets to acquire, and when to stop investing. Investors face a related problem: estimating whether the probability implied by available evidence differs from the probability reflected in an asset's price. In both cases, a predictive probability of success is the critical parameter in all decision making.
More broadly, the entire process of drug discovery and development can be viewed as a sequence of experiments to reduce uncertainty about whether a therapy works in patients: preclinical studies, animal models, and clinical trials each contribute information toward that question. AI promises a complementary source of information by integrating a broader share of the existing biomedical record and learning which patterns in that evidence have historically been associated with clinical outcomes. Better use of the biomedical record is another means of derisking development by transforming dispersed evidence into estimates of clinical success that help determine where capital should be committed and which experiments should be run.
If you're interested in working with us on clinical trial prediction, please reach out to founders@atlasdiscovery.bio.
