MedPIC-Bench · medication-safety reasoning
When the patient changes,
does the answer change?
A model can recognize a medication risk yet miss when that warning stops applying. MedPIC-Bench tests this distinction by changing patient information. Explore the authors’ published results and Arcophos’s analysis of what they measure.
Study published · source-checked 28 September 2026
Historical results · how we checked ↗Compare the 28 published models
Paper · Table 2 ↗Selected model
Gemini-3.1-Pro
Fixed cases → changed patient information
17.7 percentage points lower on counterfactual questions.
Does the warning still apply?
These two operations cover 139 of the 183 counterfactual questions.
89 pairs reported in the paper. Pair links are absent from the public dataset, so this score is reproduced from the authors’ table. All bars use a 0–100% scale.
28 of 28 models · percentages, higher is better
Select a model name to update the comparison.
| Model | GF | CF | Activation | Deactivation | Pair | Overall |
|---|---|---|---|---|---|---|
| Proprietary | 87.7 | 69.9 | 74.6 | 77.9 | 48.3 | 80.7 |
| Proprietary | 83.5 | 63.4 | 70.4 | 61.8 | 38.2 | 75.6 |
| Proprietary | 81.7 | 57.9 | 66.2 | 55.9 | 32.6 | 72.4 |
| Proprietary | 81.3 | 56.8 | 67.6 | 50.0 | 32.6 | 71.7 |
| General | 79.6 | 56.8 | 62.0 | 55.9 | 30.3 | 70.7 |
| General | 78.2 | 55.7 | 67.6 | 51.5 | 31.5 | 69.4 |
| General | 77.5 | 54.6 | 70.4 | 45.6 | 30.3 | 68.5 |
| Medical-specific | 63.7 | 53.0 | 60.6 | 54.4 | 25.8 | 59.5 |
| Proprietary | 78.2 | 52.5 | 63.4 | 42.6 | 29.2 | 68.1 |
| Medical-specific | 67.6 | 51.4 | 76.1 | 32.4 | 23.6 | 61.2 |
| Proprietary | 74.6 | 48.1 | 62.0 | 35.3 | 21.3 | 64.2 |
| General | 73.9 | 48.1 | 54.9 | 47.1 | 24.7 | 63.8 |
| Medical-specific | 69.0 | 47.5 | 59.2 | 44.1 | 20.2 | 60.6 |
| General | 66.5 | 47.5 | 52.1 | 48.5 | 24.7 | 59.1 |
| General | 61.3 | 47.0 | 63.4 | 32.4 | 16.9 | 55.7 |
| Proprietary | 60.6 | 44.8 | 63.4 | 27.9 | 20.2 | 54.4 |
| Medical-specific | 63.7 | 42.1 | 63.4 | 23.5 | 14.6 | 55.2 |
| Medical-specific | 54.6 | 37.2 | 49.3 | 30.9 | 11.2 | 47.8 |
| Medical-specific | 59.9 | 36.6 | 53.5 | 29.4 | 12.4 | 50.7 |
| Medical-specific | 50.0 | 35.5 | 49.3 | 20.6 | 13.5 | 44.3 |
| General | 44.7 | 33.3 | 53.5 | 10.3 | 5.6 | 40.3 |
| Medical-specific | 37.7 | 33.3 | 36.6 | 41.2 | 11.2 | 36.0 |
| General | 44.4 | 33.3 | 54.9 | 10.3 | 5.6 | 40.0 |
| General | 44.0 | 32.8 | 57.7 | 7.4 | 3.4 | 39.6 |
| Medical-specific | 48.2 | 32.2 | 36.6 | 38.2 | 14.6 | 42.0 |
| Medical-specific | 50.7 | 31.7 | 39.4 | 32.4 | 6.7 | 43.3 |
| Medical-specific | 45.4 | 30.6 | 46.5 | 10.3 | 5.6 | 39.6 |
| Medical-specific | 51.8 | 28.4 | 42.3 | 20.6 | 5.6 | 42.6 |
GF and CF score different task sets. A smaller gap can reflect low accuracy on both. Group labels follow the paper and do not control for model size or training. All 28 rows: source Table 2.
Three readings of a correct answer
01 · 284 questions
Guideline-following · GF
Choose the medication-safety answer for a fixed patient case.
02 · 183 questions
Counterfactual · CF
Answer cases where a controlled change in patient information changes which rule applies.
03 · 89 linked pairs, paper-reported
Pair accuracy
Both cases in a linked comparison must be correct. Each was presented to the model separately.
Answers can contain several options. A question is correct only when the selected set exactly matches the answer key. A plausible explanation does not earn partial credit. Read the task and scoring definitions ↗
A warning needs a condition
The useful question is whether patient information governs the decision. In the study, every model scores lower on counterfactual questions than on guideline-following questions. Pair accuracy adds a stricter check: can it get both sides of a contrast right?
Read activation and deactivation together. The task tests whether a model can apply a warning and withdraw it when the relevant condition is absent. Frequent warnings alone do not establish appropriate medication-safety reasoning.
Authors’ findings · §4.2 ↗Coverage shapes the result
352 of the 467 questions carry an older-adult population label. The other questions concern pediatrics or pregnancy. These are constructed vignettes drawn from selected medication rules, not a representative sample of care.
The benchmark probes conditional decisions in English multiple-choice questions. It does not measure prescribing in practice, patient outcomes, or a model’s ability to gather a complete history.
Explore all six coverage dimensions ↗An independent reading of the authors’ benchmark
MedPIC-Bench was introduced by Zhitian Hou, Yuhang Liu, Pengkai Wang and colleagues, affiliated with The Hong Kong Polytechnic University, InfiX.ai and Sun Yat-sen University. Health Evals is an Arcophos analysis of their public release. We checked the published scores against the paper; we have not rerun the models. The review date does not represent a new measurement.
Looking for the wider collection? Browse the healthcare benchmark index ↗