Clinical Evals

How WHBench works

academic team · 47 scenarios · updated August 17, 2026

WHBench asks a narrow question: how well do models handle women's health. Three independent researchers, Maurya, Govindgari, and Kumar, published it in April 2026 (arXiv 2604.00024). It contains 47 expert-crafted scenarios across 10 topics, and the study scored 3,100 responses from 22 models. It is a research study with expert validation, not a live leaderboard.

publisherIndependent researchers (Maurya, Govindgari, Kumar)
released2026-04
size47 scenarios / 3,100 scored responses across 22 models
scalemean normalized percentage 0-100, higher better

The rubric

Every response is graded against a 23-criterion rubric. The criteria cover clinical accuracy, safety, equity, and adherence to current guidelines, so a response is judged on more than whether its final recommendation is right.

Scores are reported as a mean normalized percentage from 0 to 100. A model's number is the share of available rubric credit it earned, averaged across scenarios, rather than a count of criteria met.

The failure modes it targets

The scenarios were written to catch specific errors: outdated guidelines, unsafe omissions, dosing errors, and equity blind spots. Omission is the one worth understanding. A model can produce an answer that is accurate line by line and still lose rubric credit for leaving out a warning a clinician would have given.

Rubric grading is built for exactly that. HealthBench does the same thing with criteria weighted from -10 to +10, where harmful content subtracts points. Grading against a written checklist catches what a fluent answer leaves out; grading on overall impression does not.

What the score covers

A WHBench number describes performance in one domain. It is not a general medical quality score and it does not transfer to other specialties. That is the purpose of a domain benchmark: broad boards average over many topics and can hide a weak area, and this one isolates the area. Read it next to the general rubric boards, not instead of them.

Limits

The model set was frozen in March 2026, before GPT-5.6 and the Claude 5 family shipped, so the published table cannot tell you anything about current models. Treat it as evidence about a method and a failure profile rather than a ranking. The scenario count is small at 47, so a few scenarios can move a model's mean. And one research team wrote both the scenarios and the rubric, with no stated plan to refresh, so nothing in the design corrects for a blind spot in the scenario set itself.

clinicalbenchmarks.ai tracks WHBench with the other rubric-graded boards, and it is where to check what has been scored since.

The grading machinery behind pages like this one is covered in the rubric grading and model graders guides.