Knowledge access is part of the system
SLAKE has both vision-only and knowledge-based questions. [4]
Our inference: an overall gain cannot identify improved visual grounding unless the relevant slice is inspected.
Independent benchmark analysis / 2021 paper; cleaned SLAKE 1.0 tracked separately
SLAKE combines image questions with questions that require structured medical knowledge. That distinction makes it useful for inspecting what a multimodal model is actually being asked to contribute. We separate the image split, language setting and external-knowledge condition, then compare a small set of original vision-only baselines. We also preserve the authors’ warning that their cleaned release differs from the paper. A fair replication begins with that version choice. Our tools organize the published task families; they do not claim to measure a new model or infer clinical competence from a question-answer score.
01 / What is being tested?
Data origin. Selected medical images, physician-supervised questions and bilingual knowledge-graph relations. [4][5][6]
450 training, 96 validation and 96 test images.
§2.4 [4]9,849 training; 2,109 validation; 2,070 test.
Table 2 [4]2,603 English and 2,629 Chinese triplets.
§2.2–2.4 [4]The official page explicitly describes another cleaning pass.
Download notice [5]All questions attached to an image follow its partition.
[4]English and Chinese are distinct reported conditions.
[4]Separate vision-only from knowledge-based tasks.
[4]Original test answers are constrained to appear in training.
[4]Dataset anatomy
Original image partition.
Original image partition.
Original image partition.
Use the stated counts. The paper contains inconsistent ratio wording; 450/96/96 is unambiguous and sums to 642. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Higher is better
03 / Measured evidence
Paper-reported results / selected rows
2021 paper, English vision-only questions, original split.
These are different slices of one model, not three competing systems. Cleaned-release scores require a separate comparison.
Source: Table 4 [4]
04 / Our original analysis
SLAKE has both vision-only and knowledge-based questions. [4]
Our inference: an overall gain cannot identify improved visual grounding unless the relevant slice is inspected.
Images are partitioned while test answer classes occur in training. [4]
Our inference: unseen-image evaluation and unseen-answer evaluation are distinct; this protocol supports the former condition.
05 / Scope of the evidence
The original classification protocol is narrower than unrestricted clinical report generation. [4]
Knowledge-only shortcuts and visual grounding need separate investigations; no single aggregate establishes both. [4]
Paper, cleaned release and language-filtered subsets have different denominators. [5]
Evidence trail
Liu et al. / ISBI 2021. Original paper counts, image split, question families and reported VGG/SAN results.
Bo Liu and Xiao-Ming Wu / Med-VQA. Explicitly warns that the cleaned SLAKE 1.0 release differs from the paper.
BoKelvin / Bo Liu. Author-linked release declares CC BY 4.0 and documents validation filename change.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.