Firms that build medicines organize themselves around three constraints: regulation, the cost of experimental capital, and the knowledge work that coordinates the two. AI shifts the cost of knowledge work to the cost of inference, which makes a differently organized firm possible, one driven by agents that can develop drugs for the long tail of diseases where the biology is tractable but market size has never justified the cost of thinking about how to treat them.
Existing repurposing benchmarks cannot tell reasoning apart from recall, because the drug-disease pairs they hold out already sit in the pretraining corpus. RepurposingBench builds its evaluation sets from pairs that no model recovers under direct probing, plus ten disclosed only after the knowledge cutoff. Weighted recall lands at 35% to 41% on the rediscovery set, and every model scores under 9 out of 100 on the prospective set.
One of our drug repurposing candidates is undergoing preclinical testing in partnership with Alliance to Cure Cavernous Malformation. Our research agents found an FDA-approved drug with over a decade of human safety data as a candidate for CCM, a disease affecting 18 to 24 million people with no approved cure.
Every clinical trial has an implicit forecast, but the evidence behind it is scattered across papers, patents, filings, and prior programs. We built an agent harness that assembles that evidence and paired it with a model that converts it into probabilities of clinical success: it beats the leading academic baselines on AUROC and AUPRC at every phase transition, reaching AUROC 0.858 for Phase III to approval.
Clinical care requires more than retrieving medical knowledge. ClinicBench, drawn from 500,000+ real patient records, asks whether an agent can forecast the next 30 days of a patient's care. The best of 9 frontier models scores 51.5 of 100, and all of them are better at generating plausible continuations than at predicting the care that actually happened.
Knowledge graphs freeze biomedical evidence into fixed relations before the question is known. We built a harness of specialized agents that instead builds the representation each hypothesis requires, and backtested it on a corpus frozen before the answer existed: it put all three SGLT2 inhibitors in the top 3 for heart failure, recovering dapagliflozin at rank 2, where TxGNN placed it 4,388th of 7,957.
A foundational model of patient biology predicts response to ustekinumab in ulcerative colitis from baseline biopsies alone at AUROC 0.76 — and would have let the UNIFI trial enroll 458 fewer patients at matched statistical power.