Clinical Evals

How the Artificial Analysis Healthcare & Medical Index works

Artificial Analysis · 4-benchmark composite · updated August 17, 2026

Artificial Analysis runs models through its own harness and publishes index scores for particular kinds of work. The Healthcare & Medical Index, released in August 2026, is its health view, and it covers 27 of the 159 models the site tracks. The index is assembled from general capability benchmarks rather than from tasks written by clinicians. This site therefore files it as a capability index, not a clinical benchmark.

publisherArtificial Analysis
released2026-08
size27 of 159 models scored; composite of 4 underlying benchmarks
scaleindex score, higher better

The five weighted components

The index is a weighted average of five components. Medical & Health Knowledge carries 35 percent of the weight, Agentic Knowledge Work 25 percent, Non-Hallucination 15 percent, Reasoning 15 percent, and Agentic Customer Interaction the remaining 10 percent. The weights are the design: a gain on medical knowledge moves the index three and a half times as much as the same gain on customer interaction.

Those components are drawn from four underlying benchmarks: AA-Omniscience, GDPval-AA v2, Humanity's Last Exam, and tau3-Banking. None of the four was built for clinical work. The index takes health-relevant portions of general capability suites and reweights them into a healthcare view.

How the rows are produced

Artificial Analysis runs the evaluations itself instead of collecting numbers reported by model vendors, so every row on the page went through the same components under the same weighting. Coverage is limited to models it has run, which is 27 of the 159 it tracks.

The published page differentiates only its top scores numerically. Below the leading group you cannot read exact values off the page, so small gaps in the middle of the table are not visible. Treat the index as a coarse ordering rather than a number you can subtract.

What a score does and does not mean

A high score says a model is strong on knowledge, reasoning, hallucination avoidance, and agentic task work, measured on slices chosen for health relevance. It does not say the model was tested on patient conversations, clinical notes, or a hospital record system. For those questions, read a purpose-built board such as HealthBench Professional for physician-written rubrics or PhysicianBench for agent work inside electronic health records.

Averaging also hides its parts. A model can sit high on the index while being weak on the single component that matters to your use.

Limits

The index inherits the limits of its sources, and its sources are general benchmarks. Nothing in it is graded by physicians against clinical rubrics, and nothing in it happens in a clinical environment. The weights are an editorial judgment about what health work requires, and a different weighting would reorder the same underlying runs. Coverage stops at 27 of 159 tracked models, and numeric differentiation stops after the top scores, so ties and near-ties in the middle of the table cannot be resolved.

Current standings for this index, and for the clinical boards tracked next to it, are at clinicalbenchmarks.ai.

The grading machinery behind pages like this one is covered in the rubric grading and model graders guides.