CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — evaluation 7 upvotes

ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

QUESTION — How can we measure whether report-conditioned medical vision-language models actually rely on input images versus text reports?

The study introduces ModaLens, a paired image-swap audit to measure how the availability of a radiology report alters image sensitivity in medical vision-language models. Testing MedGemma-27B on 3,199 paired MIMIC-CXR cases from 293 patients, images were swapped while holding the report and question constant. Under explicit answering instructions, the model's generated answer changes on 4.26 percent of trials with the report and 20.94 percent without it, representing a paired increase of 16.7 points. This demonstrates that report availability significantly reduces image sensitivity, a direction that also replicates across two other model lineages.

MedGemma-27B is evaluated on 3,199 paired MIMIC-CXR cases from 293 patients across 14 questions per case.

Under explicit instructions, the generated answer changes on 4.26 percent of trials with the report and 20.94 percent without it, a paired increase of 16.7 points.

The original prompt with a lowercase first-token readout gives 4.70 percent against 17.07 percent.

sebasmos · 14 Sept 2026 read the original ↗
↑