Independent benchmark analysis / Original 2018 paper and archive

VQA-RAD

Count images, question records and question variants separately.

VQA-RAD is small enough that its construction matters as much as its headline score. Clinician-generated questions concern individual radiology images, and the release contains different forms of related questions. We unpack the frequently repeated question totals, the selected test questions and the original manual scoring protocol. The resulting analysis focuses on the unit of independence: a new wording of a question is useful robustness evidence, but is not automatically a new patient or a new image. Our historical results retain their original scoring context instead of being presented as contemporary model performance.

01 / What is being tested?

The task, before the score.

input
A radiology image and an English question
output
An image-grounded short answer
unit
One image-question pair
setting
Original test: 300 selected free-form questions plus 151 corresponding paraphrases

Data origin. MedPix teaching cases; questions and answers produced and reviewed by clinical participants. [1][2][3]

Images
315

Single-image task, spanning head, chest and abdomen.

Data Records; Methods [1]
Released QA records
2,248

Free-form and rephrased elements reported in Data Records.

Data Records [1]
All question forms
3,515

Includes 1,267 framed questions alongside 1,515 free-form and 733 rephrased.

Technical Validation [1]
Original test
451 questions

300 free-form + 151 paraphrases; a question-based selection.

Evaluation of VQA Systems [1]
  1. 01

    Choose a single image

    The paper uses one image per MedPix case.

    [1]
  2. 02

    Ask an image-grounded question

    Question form and answer type are retained.

    [1]
  3. 03

    Apply the declared split

    Related question forms must remain visible in the manifest.

    [1]
  4. 04

    Score with the stated answer policy

    The original paper uses manual interpretation including partial credit.

    [1]

Dataset anatomy

Why the headline count exceeds the released record count

Free-form

Original paper category.

1,515 question forms[1]
Rephrased

Related wording, not independent new images.

733 question forms[1]
Framed

Template form counted in the headline total.

1,267 question forms[1]

The three paper categories sum to 3,515. Free-form plus rephrased sums to the 2,248 reported release records. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

Original manual simple accuracy

Higher is better

Predictions receive full, partial or no credit under manual review; open and closed question groups are reported separately.

Scoring definition

Simple accuracy = awarded credit / number of questions in the selected group

Do not pool this manual partial-credit score with later exact-match or LLM-judge evaluations. [1][2]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

Original closed-ended free-form results

2018 manual evaluation of the closed-ended free-form subset within the original test; excludes paraphrase questions.

Simple accuracy · %
050100
Reported
MCB_CLEF + RADTrained with ImageCLEF and VQA-RAD.
57.8%
SAN_CLEF + RADCombined training sources.
58.9%
MCB_RADVQA-RAD training.
60.6%
SAN_RADVQA-RAD training.
57.2%

Paper-reported simple accuracy, not mean category accuracy or modern exact match. No uncertainty shown in this table.

Source: Table 2 caption, Simple Acc row (publisher URL /tables/3) [2][1]

04 / Our original analysis

What follows from the design?

01

A record count is not a patient count

Published evidence

Several questions refer to the same image. [1]

Our interpretation

Our inference: resampling question rows alone can overstate the diversity of independent visual evidence.

02

Paraphrases enable a focused robustness question

Published evidence

The original test includes linked free-form and paraphrased questions. [1]

Our interpretation

Our inference: consistency across a pair can be evaluated separately from correctness; consistently wrong answers should remain failures.

03

The scoring rule can move the result

Published evidence

Original manual scoring accepts some partial or differently specified answers. [1][2]

Our interpretation

Our inference: document what constitutes equivalence before comparing a generative model with historical results.

05 / Scope of the evidence

Where this benchmark stops.

Selected single images

Performance does not establish reading of full image series or longitudinal studies. [1]

Question-level test construction

A test question should not be assumed to imply an unseen image. Audit image identities in the chosen release. [1]

Different downstream repackagings

Use the archived manifest and state transformations; headline counts alone do not identify a later dataset variant. [3]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Public official archive
License
CC0 1.0 Universal, confirmed via OSF license API
Conditions
This publication contains metadata and analysis only; consult the archive for assets.
[3]

Evidence trail

Read the originals.

  1. A dataset of clinically generated visual questions and answers about radiology images ↗

    Lau et al. / Scientific Data. Original construction, released records, question variants, test selection and human scoring.

  2. VQA-RAD Table 2: closed-ended free-form question results ↗

    Lau et al. / Scientific Data. Selected original simple-accuracy measurements; these use the paper’s manual scoring.

  3. Visual Question Answering in Radiology ↗

    Lau / US National Library of Medicine. Official dataset archive. Its OSF API license relationship resolves to CC0 1.0 Universal.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗