Atlas Discovery
← all posts

Clinical Trial Prediction with Agents

We built an agent harness that assembles evidence across the biomedical ecosystem and paired it with a model that converts that evidence into probabilities of clinical success, outperforming the leading academic models on trial prediction.

Clinical trials are fundamentally inference problems

Every clinical trial has an implicit forecast. The actions of companies and markets track an implied probability that a therapy will succeed. The fundamental question is given everything knowable about a program at a point in time, what is the probability that it advances?

Existing models estimate this probability largely from structured datasets. Yet the state of a therapeutic program is not contained in any single database. It is distributed across papers, conference abstracts, patents, regulatory filings, prior clinical programs, and the broader history of the target, disease, and drug. The problem has been one of practical infeasibility: the evidence exists, but assembling and interpreting it at scale has historically required prohibitive amounts of manual effort.

Agents change information collection

AI agents have made long-running, open-ended investigation tractable.

Instead of operating over a fixed dataset, agents can adaptively search across the biomedical ecosystem: navigating interfaces, discovering relevant sources, extracting findings, and following new lines of inquiry as evidence accumulates. Work that may have required weeks of manual diligence can now be performed in hours.

More importantly, information can also be synthesized effectively. Trial success depends on how diverse pieces of evidence combine into an integrated view of disease biology, pharmacology, clinical precedent, competitive programs, trial design, sponsor history, and regulatory context, among other factors. Agents are uniquely suited to constructing these higher-order representations by continuously placing new information into broader biomedical and regulatory contexts.

The result is a substantially richer representation of the evidence informative of drug development success.

Evidence to probability

To guide decisions, the resulting evidence must be concretized into quantitative estimates of trial success. Our approach uses agents to generate data to train a deep-learning model: agents generate evidence while a model learns the statistical relationship between that evidence and subsequent clinical outcomes.

Empirical performance

The objective is to estimate the probability that a clinical program transitions to the next stage of development, e.g., a phase 1 trial progressing to phase 2.

We evaluated the platform on historical trials using a sandboxed backtesting environment with data temporally restricted such that no knowledge of a trial's outcome could enter the evidence. The model was trained on a 100k trial corpus and evaluated on a held-out set of 10k trials.

Bar charts of AUROC by phase transition against four academic baselines. Our model leads all three panels: 0.820 for Phase I to II versus Trial-Bench 0.782, AUTOCT 0.753, SPOT 0.660, and HINT 0.576; 0.799 for Phase II to III versus Trial-Bench 0.771, HINT 0.645, AUTOCT 0.639, and SPOT 0.630; and 0.858 for Phase III to IV / FDA approval versus Trial-Bench 0.741, HINT 0.723, SPOT 0.711, and AUTOCT 0.702.
Figure 1. AUROC by phase transition on the held-out set of 10k trials. Our model outperforms Trial-Bench, AUTOCT, SPOT, and HINT at every stage: 0.820 for Phase I → II, 0.799 for Phase II → III, and 0.858 for Phase III → IV / FDA approval.
Bar charts of AUPRC by phase transition against the same four baselines. Our model leads all three panels: 0.770 for Phase I to II versus AUTOCT 0.710, SPOT 0.689, Trial-Bench 0.579, and HINT 0.567; 0.721 for Phase II to III versus SPOT 0.685, HINT 0.629, AUTOCT 0.512, and Trial-Bench 0.510; and 0.885 for Phase III to IV / FDA approval versus SPOT 0.856, HINT 0.811, AUTOCT 0.697, and Trial-Bench 0.638.
Figure 2. AUPRC by phase transition on the same held-out set. Our model again leads at every stage: 0.770 for Phase I → II, 0.721 for Phase II → III, and 0.885 for Phase III → IV / FDA approval.
Two bar charts of average scores across phases. Average AUROC: our model 0.826, Trial-Bench 0.765, AUTOCT 0.698, SPOT 0.667, and HINT 0.648. Average AUPRC: our model 0.792, SPOT 0.743, HINT 0.669, AUTOCT 0.640, and Trial-Bench 0.576.
Figure 3. Average AUROC and AUPRC across the three phase transitions. Our model averages 0.826 AUROC and 0.792 AUPRC, ahead of every baseline on both metrics.

Our model outperforms the leading academic baselines on AUROC and AUPRC scores across all phases.

It should be noted these comparisons are between models evaluated on different test sets, so we choose a randomized test set for evaluation.

Significance and further implications

Clinical development is fundamentally a problem of capital allocation under extreme uncertainty. Biopharma must decide which programs to fund, which trials to run, which assets to acquire, and when to stop investing. Investors face a related problem: estimating whether the probability implied by available evidence differs from the probability reflected in an asset's price. In both cases, a predictive probability of success is the critical parameter in all decision making.

More broadly, the entire process of drug discovery and development can be viewed as a sequence of experiments to reduce uncertainty about whether a therapy works in patients: preclinical studies, animal models, and clinical trials each contribute information toward that question. AI promises a complementary source of information by integrating a broader share of the existing biomedical record and learning which patterns in that evidence have historically been associated with clinical outcomes. Better use of the biomedical record is another means of derisking development by transforming dispersed evidence into estimates of clinical success that help determine where capital should be committed and which experiments should be run.

If you're interested in working with us on clinical trial prediction, please reach out to founders@atlasdiscovery.bio.

Subscribe to our work

Occasional notes on what we're building.

← all posts