Clinical Evals

How MAST works

ARISE AI Research Network · 6-benchmark composite · updated August 17, 2026

MAST, the Medical AI Superintelligence Test, is a composite rather than a task set. It aggregates six existing clinical benchmarks into a single percentage. It is published by the ARISE AI Research Network, a multi-institutional group, and covers 11 models. The board is marked as a preview and was last updated August 15, 2026.

publisherARISE AI Research Network (multi-institutional)
released2026-08
sizecomposite of 6 component benchmarks; 11 models
scalepercentage composite, higher better

What goes into the composite

MAST spans six areas of clinical work: diagnostic reasoning, management reasoning, safety, multimodal images, multimodal radiology, and agentic capability. The six components are First, Do NOHARM v2, SCT-Bench, MedAgentBench v2, PhysicianBench, ReXrank Mini, and CPC-Bench. Each component keeps its own construction and scoring; MAST supplies the aggregation, not the tasks.

One component has published detail. First, Do NOHARM v2 measures how often and how severely a model's consultation recommendations could cause harm, across 1,100 primary-care-to-specialist consultation cases in 10 specialties, with 12,747 expert annotations on 4,249 management options. Its v1 study found potential for severe harm in up to 24.6 percent of directly applied recommendations, and errors of omission accounted for more than 80 percent of the severe cases.

How a composite behaves

A composite moves for reasons a single benchmark does not. When one component saturates, the total stops responding to progress in that area. When a component is added, removed, or rerun, every row can shift while no model has changed. A single number also cannot tell you whether a model is safe but weak on radiology or the reverse.

That is why component-level publication matters. MAST currently breaks out only First, Do NOHARM v2, so a score tells you where a model sits overall but not what drove it. Read it as an index, and go to the components when you need to act on the result.

Reading a preview board

The board carries a preview label, and the publishers say scores may move before the full release. That is a statement about the measurement rather than about the models. Until the full release lands, gaps between adjacent rows are the least durable part of the table.

Limits

MAST is a preview covering 11 models, and the weighting behind the composite is not something you can currently audit from the published board. Five of the six components are not broken out, so a score cannot be reconstructed or traced to the area that produced it. Composites also inherit every limitation of their parts, including grader choices and task selection, while hiding those choices behind one number. ARISE publishes MAST and also hosts First, Do NOHARM v2, so one organization sits on both sides of that component. None of this makes the ordering wrong; it makes it provisional, which is what the preview label says.

Current MAST standings, and the component benchmarks tracked separately, are indexed at clinicalbenchmarks.ai.

The grading machinery behind pages like this one is covered in the rubric grading and model graders guides.