Clinical Evals

How MedHELM works

Stanford CRFM · 121 tasks · updated August 17, 2026

MedHELM is Stanford CRFM's holistic evaluation for clinical tasks, first released in February 2025. It runs models across 121 tasks drawn from 31 datasets, organized by a taxonomy that clinicians validated: 5 categories and 22 subcategories. The Stanford-led maintainers run every model themselves and refresh the board on a roughly quarterly cadence. The headline metric is mean win rate, which behaves differently from a percentage score.

publisherStanford CRFM / HAI and multi-institution collaborators
released2025-02
size121 tasks / 31 datasets
scalemean win rate 0-1, higher better

The task taxonomy

MedHELM starts from a map of clinical work rather than from whichever datasets happened to be available. Clinicians validated a taxonomy of 5 categories and 22 subcategories, and 121 tasks across 31 datasets were assembled to fill it. Coverage is the design goal: the board is trying to say something about the range of clinical work, not to find one hard task and rank models on it.

Tasks pulled from 31 different datasets do not all report on the same scale, so averaging their raw scores would produce a meaningless number. That constraint is what leads to the aggregation method.

How mean win rate works

On each task, a model is compared head to head against every other model evaluated on that task. Its win rate for the task is the fraction of those comparisons it wins. Mean win rate is the average of those per-task win rates, reported from 0 to 1.

This makes results comparable across tasks with different scales, and it also makes them relative. A mean win rate describes a model's position inside the evaluated cohort, not how much of the work it got right. Add a strong model to the cohort and every other number falls, even though no model changed.

Reading the board

The current board is version 5.0.0, last updated May 14, 2026. With a quarterly cadence, a recently released model may simply not have been run yet; the Claude 5 family and GPT-5.6 have no rows. Absence from the table is a scheduling fact, not a result. And because the metric is cohort-relative, comparing a model's number across two versions of the board tells you about the cohort as much as about the model.

Limits

Mean win rate cannot be read as accuracy and cannot be carried across versions, which is the most common misuse of this board. The refresh cadence means the table lags model releases by up to a quarter. Broad coverage also trades depth for breadth: 121 tasks spread over 22 subcategories leaves few tasks per subcategory, so any single subcategory result rests on thin ground. The underlying datasets were built for their own purposes and carry their own labeling conventions, which the taxonomy organizes but does not make uniform.

Current MedHELM standings, and how they line up against other clinical benchmarks, are indexed at clinicalbenchmarks.ai.

The grading machinery behind pages like this one is covered in the rubric grading and model graders guides.