The short answer
SLAKE combines visual questions, bilingual wording and structured medical knowledge. Those features make a single unqualified accuracy difficult to interpret. Our analysis treats release, language and knowledge access as three separate configuration choices. The original paper supplies one set of counts and baseline conditions; the maintainers explicitly warn that the cleaned SLAKE 1.0 release differs from it.
Identify the exact dataset release
A paper is a description of an experiment, while a maintained dataset can change after publication. SLAKE’s project page makes that distinction explicit through its cleaning notice. A later download should therefore not inherit every original count by assumption. Archive the release revision, file names and filters actually used for the run.
The author-linked dataset card also documents a validation filename change. That is a small example of why a working loader and a benchmark description need to be checked together. Our recommendation is to report counts after loading, after language selection and after exclusions. Keep those stages separate so a reviewer can distinguish a release change from a preprocessing decision or an accidental dropped record.
Separate visual evidence from external knowledge
The original benchmark distinguishes vision-only and knowledge-based questions. Some answers depend on what is visible, while others require a relationship supplied through the knowledge graph. The original experiments include configurations using that knowledge resource. Its availability is therefore part of the evaluated system, not background context that can be left unspecified.
Our inference is that aggregate accuracy cannot locate the source of an improvement. A change could improve image interpretation, question understanding, access to the graph or answer classification. Keep the task slices separate before attributing a gain to visual reasoning. A controlled follow-up can change one component while preserving question identities and the other inputs; it should be presented as a new experiment if actually performed.
Understand the supported generalization claim
The original image partition places all questions belonging to an image in the same split. That is a useful distinction from selecting question rows independently. However, the paper also constrains test answers to appear in training. The protocol therefore does not require a system to invent a completely unseen answer label at evaluation time.
These choices create a specific target: answer known-vocabulary questions about held-out images under a defined question distribution. Our analysis does not treat that as a weakness that makes the benchmark unusable. It is a boundary that helps select the right companion test. A product expected to generate long reports or handle new answer concepts needs evidence addressing those additional behaviors.
Report bilingual results without double-counting images
English and Chinese questions can refer to the same underlying images. Adding language-level question counts does not create additional independent image cases. Likewise, a bilingual paper total is not the denominator for an English-only evaluation. The count displayed beside a score should follow the language filter actually applied.
We recommend a compact run header naming the release, language, image split, question type and knowledge condition. Preserve open-answer and closed-answer outcomes when the protocol distinguishes them. This is our suggested comparison artifact, not an official submission requirement. Its purpose is to make future results commensurable and to stop version or input changes from being mistaken for progress in a particular model capability.
References & further reading
These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.
- SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering ↗Liu et al. / ISBI 2021. Original paper counts, image split, question families and reported VGG/SAN results.
- SLAKE official project page ↗Bo Liu and Xiao-Ming Wu / Med-VQA. Explicitly warns that the cleaned SLAKE 1.0 release differs from the paper.
- SLAKE author-maintained dataset card ↗BoKelvin / Bo Liu. Author-linked release declares CC BY 4.0 and documents validation filename change.
Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.