Atlas Discovery
← all posts

Agents for Drug Repurposing

The Problem of Drug Repurposing

Drug repurposing is fundamentally a mapping problem: Given a set of drugs and a set of diseases, the objective is to identify the drug-disease connections most likely to produce clinical benefit.

Each potential connection can be represented as an edge whose weight is the probability that a representative patient with the disease will respond favorably to the drug. Repurposing should not be treated simply as the retrieval of plausible associations. It is a probabilistic inference problem of estimating the likelihood of clinical success from the joint distribution of evidence surrounding a drug, a disease, and the patients who connect them.

This evidence spans multiple biological scales. At the molecular level, a drug perturbs targets, pathways, and regulatory networks. At the cellular and tissue levels, these perturbations alter disease-relevant states and physiological processes. At the patient level, genetic background, disease subtype, pharmacokinetics, comorbidities, and prior treatments influence whether those effects translate into clinical benefit. The strength of a drug-disease pair therefore depends on two kinds of evidence:

  1. What effects does the drug produce at each biological scale?
  2. What changes at each scale are associated with therapeutic response in the disease?

The Limits of Knowledge Graph Repurposing

Traditional computational approaches to drug repurposing have been dominated by knowledge graphs. These methods represent drugs, diseases, genes, proteins, pathways, phenotypes, and other entities as nodes connected by fixed relationships. Drug repurposing becomes a link-prediction problem: the model estimates whether an unobserved drug-disease edge is supported by paths through intermediate biological entities.

The limitation is not the theoretical expressivity of the graph representation but the fixity of the schema. It is a compressed representation of a subset of biomedical knowledge. Biomedical facts must first be converted into discrete entities related by consistent edge types. As a result, information that cannot be expressed within the chosen ontology is unavailable to downstream inference.

Crucially, the construction transforms evidence into typed relations before the question is known, discarding the biological context that is largely determinative of whether an edge actually holds between entities. For instance, the effect of a protein may depend on the cell type in which it is expressed, the tissue environment, or the disease state. A statement such as "protein A activates pathway B" may be correct in one biological context though misleading in another.

Perhaps most importantly, though heretofore discussed, is the point that knowledge graphs work over already established knowledge, such that there is no discovery, only new walks. They do not naturally perform the open-ended analyses required to generate new evidence or hypotheses beyond alternative ontologies of existing knowledge.

A Harness for Drug Repurposing

Advances in large language models have enabled long-running agentic systems that can plan, delegate subtasks, use computational tools, evaluate intermediate results, and coordinate multistep experiments. Rather than reasoning over a single static representation of biomedical knowledge, agents can construct task-specific representations of a disease and its biology as they work. They can read papers, query databases, run machine-learning models, analyze datasets with bioinformatics pipelines, and revise their hypotheses in response to the resulting evidence. Context need not be pre-encoded, because it can be reconstructed on demand at whatever granularity a specific hypothesis requires.

We built a harness that coordinates a fleet of specialized agents and gives them access to hundreds of tools. Unlike a knowledge graph, an agentic system can generate novel evidence during inference by combining existing observations with analyses performed during the search itself. These agents can read scientific literature, retrieve molecular, pharmacological, physiological, and clinical evidence, run and build predictive models, and coordinate multistep computational experiments. For example, one group of agents may identify disease-associated pathways from genetic and transcriptomic evidence. Another may retrieve drugs predicted to perturb those pathways. Additional agents may evaluate target engagement, chemical similarity, tissue exposure, pharmacokinetics, toxicity, and clinical plausibility. A final synthesis process can produce a detailed, evidence-backed report for each drug-disease pair.

The result is an active discovery system that searches across molecular, cellular, tissue, physiological, and clinical scales, generates candidate mechanisms, and ranks drug-disease pairs by their estimated probability of success.

Evaluation

The flexibility that makes agentic systems useful also makes them difficult to evaluate. A system designed to access nearly all available biomedical information may recover simply through direct retrieval of the known answer or memorization from the model's pretraining corpus.

We address these problems with an auxiliary controlled evaluation harness that allows the platform to be backtested on historical repurposing cases. For each benchmark, we select a drug-disease relationship that was validated after a specified cutoff date. The agents are then placed in an environment in which retrieval is restricted to evidence published before that cutoff, such that any sources that explicitly disclose the eventual repurposing relationship are excluded. To prevent pre-training leakage, we map key terms such as the drug, disease, organization, authors, and trial names to nonspecific identifiers, e.g., dapagliflozin to "d-001," such that the agents have no direct knowledge of the drug repurposing outcomes. We also inspect the agents' queries, retrieved documents, and reasoning traces for evidence that the answer was recalled rather than discovered.

This makes the task whether the harness can recover the subsequently validated candidate using only contemporaneously available biomedical knowledge with safeguards to reduce pre-training leakage.

Case Study: Dapagliflozin for Heart Failure

We first evaluated the platform on heart failure, with the goal of recovering dapagliflozin as a repurposing candidate.

Dapagliflozin was originally approved to improve glycemic control in patients with type 2 diabetes. It inhibits sodium-glucose cotransporter 2, or SGLT2, in the proximal renal tubule, reducing the reabsorption of filtered glucose and increasing urinary glucose excretion. The DAPA-HF trial later evaluated dapagliflozin in patients with heart failure and reduced ejection fraction. The treatment reduced the risk of worsening heart failure or cardiovascular death, with benefits observed in patients both with and without type 2 diabetes. The effect therefore could not be explained solely by improved glycemic control.

The mechanisms underlying the heart-failure benefit of SGLT2 inhibitors remain incompletely resolved. Proposed explanations include natriuresis and osmotic diuresis, altered renal hemodynamics, preservation of kidney function, changes in cardiac loading conditions, improved cardiac energetics, and modulation of cellular stress and inflammatory pathways.

The historical motivation for testing SGLT2 inhibitors in heart failure was driven by cardiovascular-outcome trials in patients with type 2 diabetes, which showed unexpectedly large reductions in hospitalization for heart failure. We were interested in whether the harness could not only recover the candidate but support its prediction with alternative molecular, cellular, or physiological evidence.

We prompted the harness with heart failure as the target disease and evaluated the output of a ranked list of the top 20 candidates on recall and evidential backing. The harness ranked dapagliflozin 2nd out of all possible drugs, and it identified other SGLT2 inhibitors canagliflozin and empagliflozin at 1st and 3rd.

Ranked list of 10 heart-failure repurposing candidates by composite evidence-weighted score, with canagliflozin at rank 1, dapagliflozin highlighted at rank 2, and empagliflozin at rank 3.
Figure 1. Recovery of known heart-failure repurposing candidates, over a corpus frozen at 2018-12-31. All ten candidates were marketed by the cutoff and none was established heart-failure practice. Dapagliflozin was recovered at rank 2, with the other two SGLT2 inhibitors, canagliflozin and empagliflozin, at ranks 1 and 3.

We also compared this to the state of the art repurposing model TxGNN, a GNN trained on an over 100,000-node knowledge graph, which ranked dapagliflozin at 4,388 out of 8,000 drugs when queried for heart failure.

Two-panel comparison of three SGLT2 inhibitor ranks: the agent platform places canagliflozin, dapagliflozin and empagliflozin at 1, 2 and 3 out of 10, while TxGNN places them at 3,331, 4,388 and 3,529 out of 7,957.
Figure 2. The same three SGLT2 inhibitors as ranked by TxGNN, a graph neural network trained on a knowledge graph of over 100,000 nodes. The agent platform placed canagliflozin, dapagliflozin and empagliflozin at 1, 2 and 3 within a post-triage pool of 10; TxGNN placed them at 3,331, 4,388 and 3,529 within its pool of 7,957 drugs.

Rather than relying on a single drug-target-disease path, the agents constructed a multiscale argument linking the drug's renal, metabolic, cardiovascular, and molecular effects to mechanisms implicated in heart failure.

Layered mechanism graph for heart failure running from disease through symptom, physiology, organ, tissue, region, cell state, and pathway to drug target, with the SGLT2 recovery path highlighted through renal sodium and water retention, the kidney, the proximal tubule, and the SGLT2 sodium/glucose transporter.
Figure 3. A subgraph of the 80-node heart-failure mechanism DAG, running from disease through symptom, physiology, organ, tissue, region, cell state, and pathway to target. Orange nodes lie on the SGLT2 inhibitors' recovery path to congestion; teal nodes are other nearby mechanism nodes. Target nodes show the candidate class recovered there. Select the figure to open it at full size.

The effects at each scale were evidenced by original experiments the agents performed. Most notably, it constructed a disease-relevant transcriptomic signature and tested whether available drug-induced expression profiles reversed that state. It also integrated multiple pharmacology databases to build a target-selectivity and occupancy profile, revealing both strong SLC5A2 engagement and a potentially important SLC5A4 off-target effect. Additional analyses included chemical-similarity searches, candidate off-target identification, protein structure and interaction analysis, pathway enrichment, tissue-specific expression scoring, and structured comparisons of clinical efficacy and safety evidence. These tool calls distinguish the system from purely literature synthesis, as it moves into novel evidence produced by new analysis of existing data.

Limitations

The principal limitation of this test is pre-training leakage. Retrieval can be controlled directly by imposing a knowledge cutoff, inspecting the reasoning traces for retrieved source disclosing the eventual dapagliflozin-heart-failure relation. The contents of the model's weights cannot be controlled in the same way. Mapping the drug, disease, organizations, authors, and trial names to nonspecific identifiers prevented the agents from resolving the candidate at the outset, such that its identity would have to be inferred from its properties, though this does not establish that the model was unaware of dapagliflozin, heart failure, or DAPA-HF, and it is unlikely that it was. The substitution removes the shortest path to the answer, forcing the system to arrive at the candidate through mechanism and evidence. Some transfer of general knowledge nonetheless persists, and much of it is appropriate, as familiarity with heart-failure pathophysiology and renal handling of sodium and glucose is what one would expect of a good repurposing engine.

Conclusion

The important result is not simply that the harness recovered a known drug-disease relationship, but that it could reconstruct the reasoning behind it: coordinating various tools, generating new evidence, preserving uncertainty, and producing an auditable mechanistic argument from temporally restricted information. Larger benchmarks and prospective validation will be required to robustly establish discovery performance. However, the underlying possibility is now evident: drug repurposing may become less a process of reasoning over standardized relations than one of continuously reorganizing the knowledge, data, and models that already exist into new therapeutic hypotheses. In that future, the agent is not merely a search engine or predictor, but an adaptive scientific system that builds the representation each problem requires and turns the expanse of biomedical knowledge into new experiments, explanations, and medicines.

Subscribe to our work

Occasional notes on what we're building.

← all posts