Atlas Discovery
← all posts

ClinicBench: Evaluating Frontier Models on Clinical Tasks

Large language models have transformed how many knowledge professionals work, yet their impact on medicine, particularly clinical decision-making, has remained limited. Clinical care depends on more than retrieving medical knowledge: it requires the ability to integrate complex, heterogeneous longitudinal patient data and make reliable predictions under uncertainty.

To advance agents in clinical settings, we need new sources of data to post-train models and evaluations of agent performance to provide training signals. In this work, we provide an evaluation of frontier agents' ability to reason over a patient's medical history to predict health outcomes and recommend next steps compared to an actual physician's decisions. ClinicBench evaluates whether an AI agent can act within a patient's evolving clinical workflow, recommending investigations, treatments, procedures, and follow-up based on prior patient data.

ClinicBench does not capture whether an alternative recommendation might be more thorough or clinically useful, but it provides a grounded measure of how closely today's models track real-world care decisions: a necessary step toward building agents that can ultimately improve upon them.

ClinicBench's data is drawn from a subset of real-world data across 500,000+ patient records from hospital partners around the world. The subset we focus on in this bench is in Non-Small Cell Lung Cancer (NSCLC) and rectal cancer across different points in the patient journey.

Evaluation design

The model receives a patient's clinical record up to a defined cutoff. This includes diagnoses, prior treatments, laboratory results, procedures, imaging, and clinical notes.

Example patient record

Patient history and clinical context:

Clinical history before the forecast cutoff:

The model access is cut off at this point and it is asked to predict the clinical actions taken during the next 30 days. This includes labs, imaging, treatments, procedures, and specialist recommendations, among others.

What happened during the forecast window:

Responses are compared with a rubric derived from the patient's held-out clinical record. Models can receive partial credit for recovering components of the expected workflow.

Evaluation method

We tested 9 frontier models, producing a set of recommended actions with which to measure predictive performance and consistency. Using patient-record inspection, biomedical, and web tools, each model inspected the patient's state at the forecast cutoff and predicted the next set of recommended clinical actions. We use the ground truth data of up to the next 30 days to check this.

ClinicBench evaluation loop: the patient record up to the forecast cutoff, including diagnosis and stage, clinical timeline, notes, labs, medications, procedures, and imaging, feeds a clinical agent that acts and observes over patient-evidence tools (record search, labs, imaging, timeline) and biomedical research tools (Open Targets, ChEMBL, PubMed, drug labels), produces a next-actions forecast of labs, imaging, and treatments, and a rubric judge scores it from 0 to 100 against the held-out record.
Figure 1. The ClinicBench evaluation loop. The agent receives the patient record up to the forecast cutoff, acts and observes in a tool-enabled world of patient-evidence and biomedical research tools, and submits a next-actions forecast that a rubric judge scores from 0 to 100 against the held-out record.

Benchmark performance

Models demonstrated varied performance on the benchmark, which scores how accurately an agent's predicted clinical actions match the truth dataset across semantic meaning, clinical keywords and specificity, and timing or urgency of the suggested action. The agents are also penalized for missing truth actions and predicting additional unsupported predictions. We tried using Fable 5 as well, but could not due to prompt denials.

Horizontal bar chart of overall benchmark scores for 9 frontier models: GPT-5.6 Sol 51.52, Kimi K3 47.08, Claude Sonnet 5 44.98, Muse Spark 1.1 43.91, Claude Opus 4.8 43.37, Grok 4.5 42.80, GPT-5.6 Terra 41.96, Gemini 3.5 Flash 39.03, and GPT-4.1 26.33.
Figure 2. Overall benchmark score (mean ± variation) for the 9 frontier models evaluated. GPT-5.6 Sol leads at 51.52; no model reaches 52 of 100.

We found that models struggle most with recommending treatments, perhaps owing to the higher dimensionality of the potential actions.

Heatmap of the percentage of valid model outcomes containing each action type, consultation, diagnostic, imaging, and treatment, across rectal cancer and NSCLC tasks, with consultation mostly above 85% and treatment as low as 17.8%.
Figure 3. Percentage of valid model outcomes containing each action type across the rectal cancer and NSCLC tasks. Consultations are recovered far more reliably than treatments.

We decomposed the lost scores of the model outcomes into categories of incorrect, missing, or extra actions predicted. We looked at the specific case of the example patient story and found that the models were a) under-specifying treatment regimens, like predicting zoledronic acid without gefitinib for NSCLC, and b) suggesting a lot of clinically plausible extras like brain MRI, biomarker testing, and maintenance immunotherapy.

Stacked horizontal bar chart decomposing each model's lost score points into incorrect, missing, and extra predicted actions, with each model's overall score shown at right.
Figure 4. Decomposition of each model's lost score points into incorrect, missing, and extra predicted actions, next to its overall score.

We also measured the average number of tool calls by each model for each run, per task. We find that the number of tool calls does not correlate with the model performance in any way for these niche domain tasks. Interestingly, Kimi, Sol, and Terra relied mainly on the patient record.

Stacked horizontal bar chart of mean tool calls per scored attempt, split into patient evidence, biomedical tools, and provenance tools, with mean calls ranging from 9 to 38 alongside each model's score rank.
Figure 5. Mean tool calls per scored attempt, split by tool category, alongside each model's score rank. Call volume does not track performance: the two top scorers averaged 9 calls, while the two lowest averaged 38 and 32.

Our evaluation suggests that current models are considerably better at generating plausible clinical continuations than at forecasting realized clinical trajectories. We used a scoring mechanism to award both similar predictions, and more so exact keyword-matched predictions. Distinguishing them is essential for evaluating agents intended to operate in real clinical settings.

This is a first look at the shape of the problem. We're expanding ClinicBench along the dimensions that matter for real care: curating larger datasets, expanding to more indications across oncology and beyond, more points in the patient journey, and richer modalities including imaging, pathology, and genomics.

If you're interested in partnering with our research and data, please reach out to founders@atlasdiscovery.bio.

Subscribe to our work

Occasional notes on what we're building.

← all posts