Health Evals

MedPIC-Bench · medication-safety reasoning

When the patient changes,
does the answer change?

A model can recognize a medication risk yet miss when that warning stops applying. MedPIC-Bench tests this distinction by changing patient information. Explore the authors’ published results and Arcophos’s analysis of what they measure.

467 questions28 models89 linked pairs reportedOpen dataset · CC BY 4.0

Study published · source-checked 28 September 2026

Historical results · how we checked ↗

Compare the 28 published models

Paper · Table 2 ↗

Selected model

Gemini-3.1-Pro

Fixed cases → changed patient information

Guideline-following · 284 questions87.7%
Counterfactual · 183 questions69.9%

17.7 percentage points lower on counterfactual questions.

Does the warning still apply?

Risk activation · 71 questions74.6%
Risk deactivation · 68 questions77.9%

These two operations cover 139 of the 183 counterfactual questions.

Pair accuracy · both linked cases correct48.3%

89 pairs reported in the paper. Pair links are absent from the public dataset, so this score is reproduced from the authors’ table. All bars use a 0–100% scale.

28 of 28 models · percentages, higher is better

Select a model name to update the comparison.

MedPIC-Bench published model accuracy, Table 2, 4 August 2026. GF means guideline-following; CF means counterfactual; Pair requires both linked cases correct.
ModelGFCFActivationDeactivationPairOverall
Proprietary87.769.974.677.948.380.7
Proprietary83.563.470.461.838.275.6
Proprietary81.757.966.255.932.672.4
Proprietary81.356.867.650.032.671.7
General79.656.862.055.930.370.7
General78.255.767.651.531.569.4
General77.554.670.445.630.368.5
Medical-specific63.753.060.654.425.859.5
Proprietary78.252.563.442.629.268.1
Medical-specific67.651.476.132.423.661.2
Proprietary74.648.162.035.321.364.2
General73.948.154.947.124.763.8
Medical-specific69.047.559.244.120.260.6
General66.547.552.148.524.759.1
General61.347.063.432.416.955.7
Proprietary60.644.863.427.920.254.4
Medical-specific63.742.163.423.514.655.2
Medical-specific54.637.249.330.911.247.8
Medical-specific59.936.653.529.412.450.7
Medical-specific50.035.549.320.613.544.3
General44.733.353.510.35.640.3
Medical-specific37.733.336.641.211.236.0
General44.433.354.910.35.640.0
General44.032.857.77.43.439.6
Medical-specific48.232.236.638.214.642.0
Medical-specific50.731.739.432.46.743.3
Medical-specific45.430.646.510.35.639.6
Medical-specific51.828.442.320.65.642.6

GF and CF score different task sets. A smaller gap can reflect low accuracy on both. Group labels follow the paper and do not control for model size or training. All 28 rows: source Table 2.

Three readings of a correct answer

01 · 284 questions

Guideline-following · GF

Choose the medication-safety answer for a fixed patient case.

02 · 183 questions

Counterfactual · CF

Answer cases where a controlled change in patient information changes which rule applies.

03 · 89 linked pairs, paper-reported

Pair accuracy

Both cases in a linked comparison must be correct. Each was presented to the model separately.

Answers can contain several options. A question is correct only when the selected set exactly matches the answer key. A plausible explanation does not earn partial credit. Read the task and scoring definitions ↗

A warning needs a condition

The useful question is whether patient information governs the decision. In the study, every model scores lower on counterfactual questions than on guideline-following questions. Pair accuracy adds a stricter check: can it get both sides of a contrast right?

Read activation and deactivation together. The task tests whether a model can apply a warning and withdraw it when the relevant condition is absent. Frequent warnings alone do not establish appropriate medication-safety reasoning.

Authors’ findings · §4.2 ↗

Coverage shapes the result

352 of the 467 questions carry an older-adult population label. The other questions concern pediatrics or pregnancy. These are constructed vignettes drawn from selected medication rules, not a representative sample of care.

The benchmark probes conditional decisions in English multiple-choice questions. It does not measure prescribing in practice, patient outcomes, or a model’s ability to gather a complete history.

Explore all six coverage dimensions ↗

An independent reading of the authors’ benchmark

MedPIC-Bench was introduced by Zhitian Hou, Yuhang Liu, Pengkai Wang and colleagues, affiliated with The Hong Kong Polytechnic University, InfiX.ai and Sun Yat-sen University. Health Evals is an Arcophos analysis of their public release. We checked the published scores against the paper; we have not rerun the models. The review date does not represent a new measurement.

Looking for the wider collection? Browse the healthcare benchmark index ↗