How MedXpertQA (MM) works
Tsinghua University · 2,000 multimodal questions · updated August 17, 2026
MedXpertQA (MM) is the multimodal subset of MedXpertQA, a 4,460 question benchmark released by TsinghuaC3I at Tsinghua University in January 2025. The MM subset holds 2,000 questions that pair a clinical image with a stem and a set of options. It arrived as exam style medical question sets were saturating: MedQA and similar sets passed 95 percent by 2025 and were retired by most trackers. Multiple choice keeps grading automatic, while the images and the difficulty keep the score from pinning.
| publisher | TsinghuaC3I (Tsinghua University) |
|---|---|
| released | 2025-01 |
| size | 2,000 multimodal questions (MM subset) |
| scale | percentage accuracy 0-100, higher better |
What a question contains
Images span X-ray, histology, dermatology, and charts, and the questions cover 17 specialties. The model has to interpret image and text together, then pick an option. This is still recognition and reasoning inside a fixed answer set, closer to a board exam than to a consult. What it adds over text only medical QA is that the answer cannot be reached from the stem alone.
Scoring, and what it buys you
The metric is percentage accuracy from 0 to 100. There is no rubric and no judge model, which removes a large source of variance: two people running the same questions against the same model should land close to the same number, and anyone can rerun the set.
The cost of that simplicity is what the format cannot see. Multiple choice rewards elimination, gives no credit for the quality of an explanation, and says nothing about what a model would write when no options are supplied.
Where the current numbers come from
The current frontier table for MM is vendor reported. It comes from Meta's Muse Spark launch evaluation and is mirrored display only on benchlm.ai, meaning the numbers are republished rather than rerun. Vendor numbers use the vendor's own configuration and often differ by points from independent runs of the same benchmark, so read a vendor table as one lab's measurement rather than a neutral standing. The Text subset has no comparable current table, so MM is the half with a live frontier picture.
Limits
Provenance is the main caution: a display only mirror of a vendor evaluation cannot be audited, and it covers whatever models that vendor chose to run. The question set is fixed and has been available since January 2025, so training contamination is a possibility that accuracy alone cannot detect. Multiple choice also diverges from deployed use, where a model produces free text and no option list narrows the space. A single accuracy number across 17 specialties hides which specialties or image types a model is weakest in.
Current MedXpertQA (MM) numbers are tracked on the results index at clinicalbenchmarks.ai.
The grading machinery behind pages like this one is covered in the rubric grading and model graders guides.