Independent benchmark analysis / 2021 paper; cleaned SLAKE 1.0 tracked separately

SLAKE

A visual answer and a knowledge answer test different things.

SLAKE combines image questions with questions that require structured medical knowledge. That distinction makes it useful for inspecting what a multimodal model is actually being asked to contribute. We separate the image split, language setting and external-knowledge condition, then compare a small set of original vision-only baselines. We also preserve the authors’ warning that their cleaned release differs from the paper. A fair replication begins with that version choice. Our tools organize the published task families; they do not claim to measure a new model or infer clinical competence from a question-answer score.

01 / What is being tested?

The task, before the score.

input
Radiology image, English or Chinese question, and optionally the provided knowledge graph
output
A short answer in the configured answer vocabulary
unit
One image-question pair
setting
Image-disjoint training, validation and test splits in the original paper

Data origin. Selected medical images, physician-supervised questions and bilingual knowledge-graph relations. [4][5][6]

Original images
642

450 training, 96 validation and 96 test images.

§2.4 [4]
Original bilingual QA
14,028

9,849 training; 2,109 validation; 2,070 test.

Table 2 [4]
Knowledge relations
5,232

2,603 English and 2,629 Chinese triplets.

§2.2–2.4 [4]
Release caveat
SLAKE 1.0 differs

The official page explicitly describes another cleaning pass.

Download notice [5]
  1. 01

    Split by image

    All questions attached to an image follow its partition.

    [4]
  2. 02

    Choose language

    English and Chinese are distinct reported conditions.

    [4]
  3. 03

    Route the question

    Separate vision-only from knowledge-based tasks.

    [4]
  4. 04

    Predict an answer class

    Original test answers are constrained to appear in training.

    [4]

Dataset anatomy

Original image split

Training

Original image partition.

450 images[4]
Validation

Original image partition.

96 images[4]
Test

Original image partition.

96 images[4]

Use the stated counts. The paper contains inconsistent ratio wording; 450/96/96 is unambiguous and sums to 642. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

Answer accuracy by task and language

Higher is better

The original classification evaluation separates vision-only from knowledge-based questions.

Scoring definition

Accuracy = correctly classified answers / evaluated questions in the specified slice

Fix release, language, question type, image split and access to the knowledge graph. [4][5]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

Original English vision-only baselines

2021 paper, English vision-only questions, original split.

Accuracy · %
050100
Reported
VGG + SAN: overallAll English vision-only questions.
72.73%
VGG + SAN: open-endedOpen-answer slice.
70.34%
VGG + SAN: closed-endedClosed-answer slice.
76.13%

These are different slices of one model, not three competing systems. Cleaned-release scores require a separate comparison.

Source: Table 4 [4]

04 / Our original analysis

What follows from the design?

01

Knowledge access is part of the system

Published evidence

SLAKE has both vision-only and knowledge-based questions. [4]

Our interpretation

Our inference: an overall gain cannot identify improved visual grounding unless the relevant slice is inspected.

02

The split supports one kind of generalization

Published evidence

Images are partitioned while test answer classes occur in training. [4]

Our interpretation

Our inference: unseen-image evaluation and unseen-answer evaluation are distinct; this protocol supports the former condition.

03

Cleaning changes the evaluated population

Published evidence

The maintainers explicitly distinguish SLAKE 1.0 from paper data. [5][6]

Our interpretation

Our inference: archive counts after language filtering instead of copying the bilingual paper total into an English-only report.

05 / Scope of the evidence

Where this benchmark stops.

Finite answer vocabulary

The original classification protocol is narrower than unrestricted clinical report generation. [4]

Mixed question requirements

Knowledge-only shortcuts and visual grounding need separate investigations; no single aggregate establishes both. [4]

Version-sensitive counts

Paper, cleaned release and language-filtered subsets have different denominators. [5]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Public author-linked release
License
CC BY 4.0 in author dataset-card metadata
Conditions
Attribute the creators and specify the cleaned release when using its data.
[6][5]

Evidence trail

Read the originals.

  1. SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering ↗

    Liu et al. / ISBI 2021. Original paper counts, image split, question families and reported VGG/SAN results.

  2. SLAKE official project page ↗

    Bo Liu and Xiao-Ming Wu / Med-VQA. Explicitly warns that the cleaned SLAKE 1.0 release differs from the paper.

  3. SLAKE author-maintained dataset card ↗

    BoKelvin / Bo Liu. Author-linked release declares CC BY 4.0 and documents validation filename change.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗