Benchmark analysis / 5 min read

VQA-RAD: reconcile 3,515 questions with 2,248 records

An original accounting of question forms, image reuse and the original VQA-RAD evaluation.

The short answer

VQA-RAD is frequently described through a single question count. The original paper gives several counts because it describes several units: images, released records, free-form questions, rephrasings and framed variants. Our analysis keeps those units separate. This is useful for anyone building an evaluation manifest, since a larger question total does not automatically mean more independent visual cases.

Make the counting unit explicit

The paper describes 1,515 free-form questions, 733 rephrasings and 1,267 framed questions. Those categories sum to the headline 3,515. Its Data Records section reports 2,248 elements, matching the free-form and rephrased categories together. The arithmetic is straightforward once the unit is named; neither total should be substituted silently for the other.

Our proposed data inventory begins with four separate fields: archive version, images, question records and question forms. Add a transformation log for any repackaging. If an adapter drops invalid rows or changes the representation of framed questions, its resulting count belongs to that adapter. A published benchmark name is not enough to identify the exact data that entered an evaluation.

A question split is not automatically an image split

The original test is selected from free-form questions and their corresponding paraphrases. That establishes which questions are held out; it does not by itself establish that every test image is absent from training. Multiple questions can refer to the same image, so checking only question identifiers misses an important relationship.

We recommend deriving an image-to-question index before interpreting generalization. For each test record, record whether its image and any linked question form appeared in development. This is a proposed audit, not a finding that every overlap invalidates every use. An evaluation of new questions about known images and an evaluation of unseen images can both be useful, but they support different claims.

Use paired wording to ask a narrower question

The paraphrase structure creates a natural robustness analysis. Hold the image constant and compare answers to linked question forms. Our suggested outcome categories are both correct, both incorrect, and disagreement. The distinction prevents a consistently wrong pair from being counted as a success merely because the answers match.

This analysis should retain the semantic relationship established by the source rather than treating every similarly worded question as a valid pair. Changes in requested specificity can change the correct answer. Inspect questionable pairings and record exclusions. A pair-level summary is an additional editorial analysis of the available structure; it should not be labeled as the official VQA-RAD score unless the source protocol defines it.

Keep historical scoring visible

The original paper manually evaluates answers and permits partial credit in some situations. Its simple accuracy and mean category accuracy are different summaries. Our selected panel uses the closed-ended free-form subset, excluding paraphrases, and preserves simple accuracy, model training condition and the Table 2 caption. The publisher serves that table at a URL ending in /tables/3. It does not reinterpret those measurements as later exact-match results.

For a modern replication, state whether answer normalization, synonyms, specificity and partial correctness are handled by rules, people or a model judge. Save the judged output alongside the raw response. Compare systems under one scoring policy, and use the original table as historical context. This makes the publication an inspectable analysis of measurement choices rather than a collection of superficially comparable percentages.

References & further reading

These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.

  1. A dataset of clinically generated visual questions and answers about radiology images ↗Lau et al. / Scientific Data. Original construction, released records, question variants, test selection and human scoring.
  2. VQA-RAD Table 2: closed-ended free-form question results ↗Lau et al. / Scientific Data. Selected original simple-accuracy measurements; these use the paper’s manual scoring.
  3. Visual Question Answering in Radiology ↗Lau / US National Library of Medicine. Official dataset archive. Its OSF API license relationship resolves to CC0 1.0 Universal.

Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.

Continue reading.

All guides →