Large language models have transformed how many knowledge professionals work, yet their impact on medicine, particularly clinical decision-making, has remained limited. Clinical care depends on more than retrieving medical knowledge: it requires the ability to integrate complex, heterogeneous longitudinal patient data and make reliable predictions under uncertainty.
To advance agents in clinical settings, we need new sources of data to post-train models and evaluations of agent performance to provide training signals. In this work, we provide an evaluation of frontier agents' ability to reason over a patient's medical history to predict health outcomes and recommend next steps compared to an actual physician's decisions. ClinicBench evaluates whether an AI agent can act within a patient's evolving clinical workflow, recommending investigations, treatments, procedures, and follow-up based on prior patient data.
ClinicBench does not capture whether an alternative recommendation might be more thorough or clinically useful, but it provides a grounded measure of how closely today's models track real-world care decisions: a necessary step toward building agents that can ultimately improve upon them.
ClinicBench's data is drawn from a subset of real-world data across 500,000+ patient records from hospital partners around the world. The subset we focus on in this bench is in Non-Small Cell Lung Cancer (NSCLC) and rectal cancer across different points in the patient journey.
Evaluation design
The model receives a patient's clinical record up to a defined cutoff. This includes diagnoses, prior treatments, laboratory results, procedures, imaging, and clinical notes.
Example patient record
Patient history and clinical context:
- Diagnosis: Stage IV poorly differentiated squamous-cell carcinoma of the right upper-lobe lung.
- Pathology: Endobronchial biopsy favored invasive squamous non-small-cell lung cancer.
- Immunohistochemistry: Diffusely p40 positive and TTF1 negative.
- Metastatic disease: Brain, liver, and bone metastases, including C1 vertebral involvement.
- Previous radiotherapy:
- Whole-brain radiotherapy: 25 Gy in 10 fractions
- Lung radiotherapy: 25 Gy in 10 fractions
- Forecast cutoff: Feb 13, 2023.
- Prediction period: The following 30 days, through March 12, 2023.
Clinical history before the forecast cutoff:
- Oct 27, 2022: Cycle 1: Paclitaxel, carboplatin, and zoledronic acid administered, plus PET-CT scan images for:
- Brain metastases
- Liver metastases
- Bone metastases, including the C1 vertebra
- Nov 17, 2022: Cycle 2 (same as cycle 1)
- Dec 10, 2022: Cycle 3 (same as cycle 1)
- Jan 1, 2023: Cycle 4 (same as cycle 1)
- Jan 22, 2023: Cycle 5 (same as cycle 1)
- Feb 13, 2023: Cycle 6
- Paclitaxel: 160 mg IV
- Carboplatin: 250 mg IV
- Zoledronic acid: 3 mg IV
- Feb 13, 2023: Laboratory assessment
- Hemoglobin: 6.5 g/dL
- Platelets: 247 × 10³/µL
- Total leukocyte count: 9.25 × 10³/µL
- Creatinine: 0.8 mg/dL
- Total bilirubin: 0.73 mg/dL
The model access is cut off at this point and it is asked to predict the clinical actions taken during the next 30 days. This includes labs, imaging, treatments, procedures, and specialist recommendations, among others.
What happened during the forecast window:
- Feb 21 - 27: Days 1-14
- Feb 27: A post-chemotherapy PET-CT was performed to assess treatment response.
- Recommended tests: serum creatinine tests and random blood sugar level tests.
- Feb 28 - March 3: Days 15-21
- March 3: Medical Oncology reviewed the PET-CT.
- No subsequent spike in recovery rate.
- Subsequent treatment was planned: zoledronic acid, gefitinib, and oral metronomic chemotherapy with methotrexate and cyclophosphamide.
- March 4 - 11: Days 22-29
- March 11: The first cycle of the subsequent regimen was administered: zoledronic acid, gefitinib, and oral metronomic chemotherapy.
Responses are compared with a rubric derived from the patient's held-out clinical record. Models can receive partial credit for recovering components of the expected workflow.
Evaluation method
We tested 9 frontier models, producing a set of recommended actions with which to measure predictive performance and consistency. Using patient-record inspection, biomedical, and web tools, each model inspected the patient's state at the forecast cutoff and predicted the next set of recommended clinical actions. We use the ground truth data of up to the next 30 days to check this.
Benchmark performance
Models demonstrated varied performance on the benchmark, which scores how accurately an agent's predicted clinical actions match the truth dataset across semantic meaning, clinical keywords and specificity, and timing or urgency of the suggested action. The agents are also penalized for missing truth actions and predicting additional unsupported predictions. We tried using Fable 5 as well, but could not due to prompt denials.
We found that models struggle most with recommending treatments, perhaps owing to the higher dimensionality of the potential actions.
We decomposed the lost scores of the model outcomes into categories of incorrect, missing, or extra actions predicted. We looked at the specific case of the example patient story and found that the models were a) under-specifying treatment regimens, like predicting zoledronic acid without gefitinib for NSCLC, and b) suggesting a lot of clinically plausible extras like brain MRI, biomarker testing, and maintenance immunotherapy.
We also measured the average number of tool calls by each model for each run, per task. We find that the number of tool calls does not correlate with the model performance in any way for these niche domain tasks. Interestingly, Kimi, Sol, and Terra relied mainly on the patient record.
Our evaluation suggests that current models are considerably better at generating plausible clinical continuations than at forecasting realized clinical trajectories. We used a scoring mechanism to award both similar predictions, and more so exact keyword-matched predictions. Distinguishing them is essential for evaluating agents intended to operate in real clinical settings.
This is a first look at the shape of the problem. We're expanding ClinicBench along the dimensions that matter for real care: curating larger datasets, expanding to more indications across oncology and beyond, more points in the patient journey, and richer modalities including imaging, pathology, and genomics.
If you're interested in partnering with our research and data, please reach out to founders@atlasdiscovery.bio.
