A record count is not a patient count
Several questions refer to the same image. [1]
Our inference: resampling question rows alone can overstate the diversity of independent visual evidence.
Independent benchmark analysis / Original 2018 paper and archive
VQA-RAD is small enough that its construction matters as much as its headline score. Clinician-generated questions concern individual radiology images, and the release contains different forms of related questions. We unpack the frequently repeated question totals, the selected test questions and the original manual scoring protocol. The resulting analysis focuses on the unit of independence: a new wording of a question is useful robustness evidence, but is not automatically a new patient or a new image. Our historical results retain their original scoring context instead of being presented as contemporary model performance.
01 / What is being tested?
Data origin. MedPix teaching cases; questions and answers produced and reviewed by clinical participants. [1][2][3]
Single-image task, spanning head, chest and abdomen.
Data Records; Methods [1]Free-form and rephrased elements reported in Data Records.
Data Records [1]Includes 1,267 framed questions alongside 1,515 free-form and 733 rephrased.
Technical Validation [1]300 free-form + 151 paraphrases; a question-based selection.
Evaluation of VQA Systems [1]The paper uses one image per MedPix case.
[1]Question form and answer type are retained.
[1]Related question forms must remain visible in the manifest.
[1]The original paper uses manual interpretation including partial credit.
[1]Dataset anatomy
Original paper category.
Related wording, not independent new images.
Template form counted in the headline total.
The three paper categories sum to 3,515. Free-form plus rephrased sums to the 2,248 reported release records. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Higher is better
Predictions receive full, partial or no credit under manual review; open and closed question groups are reported separately.
Simple accuracy = awarded credit / number of questions in the selected group
Do not pool this manual partial-credit score with later exact-match or LLM-judge evaluations. [1][2]
03 / Measured evidence
Paper-reported results / selected rows
2018 manual evaluation of the closed-ended free-form subset within the original test; excludes paraphrase questions.
Paper-reported simple accuracy, not mean category accuracy or modern exact match. No uncertainty shown in this table.
Source: Table 2 caption, Simple Acc row (publisher URL /tables/3) [2][1]
04 / Our original analysis
Several questions refer to the same image. [1]
Our inference: resampling question rows alone can overstate the diversity of independent visual evidence.
The original test includes linked free-form and paraphrased questions. [1]
Our inference: consistency across a pair can be evaluated separately from correctness; consistently wrong answers should remain failures.
05 / Scope of the evidence
Performance does not establish reading of full image series or longitudinal studies. [1]
A test question should not be assumed to imply an unseen image. Audit image identities in the chosen release. [1]
Use the archived manifest and state transformations; headline counts alone do not identify a later dataset variant. [3]
Evidence trail
Lau et al. / Scientific Data. Original construction, released records, question variants, test selection and human scoring.
Lau et al. / Scientific Data. Selected original simple-accuracy measurements; these use the paper’s manual scoring.
Lau / US National Library of Medicine. Official dataset archive. Its OSF API license relationship resolves to CC0 1.0 Universal.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.