{"publication":"Medical Bench","url":"https://medicalbench.ai","publisher":"Arcophos","updated":"2026-09-28","provenance":"Independent analytical publication. Benchmark creation and experimental results belong to their cited authors. Reported results are source-version snapshots, not new Arcophos runs or a live leaderboard.","benchmarks":[{"slug":"vqa-rad","name":"VQA-RAD","shortName":"VQA-RAD","version":"Original 2018 paper and archive","creators":"Jason J. Lau, Soumya Gayen, Asma Ben Abacha and Dina Demner-Fushman","paperDate":"2018","headline":"Count images, question records and question variants separately.","summary":"VQA-RAD is small enough that its construction matters as much as its headline score. Clinician-generated questions concern individual radiology images, and the release contains different forms of related questions. We unpack the frequently repeated question totals, the selected test questions and the original manual scoring protocol. The resulting analysis focuses on the unit of independence: a new wording of a question is useful robustness evidence, but is not automatically a new patient or a new image. Our historical results retain their original scoring context instead of being presented as contemporary model performance.","task":{"input":"A radiology image and an English question","output":"An image-grounded short answer","unit":"One image-question pair","setting":"Original test: 300 selected free-form questions plus 151 corresponding paraphrases"},"dataOrigin":"MedPix teaching cases; questions and answers produced and reviewed by clinical participants.","facts":[{"label":"Images","value":"315","detail":"Single-image task, spanning head, chest and abdomen.","sourceIds":["vqarad-paper"],"locator":"Data Records; Methods"},{"label":"Released QA records","value":"2,248","detail":"Free-form and rephrased elements reported in Data Records.","sourceIds":["vqarad-paper"],"locator":"Data Records"},{"label":"All question forms","value":"3,515","detail":"Includes 1,267 framed questions alongside 1,515 free-form and 733 rephrased.","sourceIds":["vqarad-paper"],"locator":"Technical Validation"},{"label":"Original test","value":"451 questions","detail":"300 free-form + 151 paraphrases; a question-based selection.","sourceIds":["vqarad-paper"],"locator":"Evaluation of VQA Systems"}],"metric":{"name":"Original manual simple accuracy","description":"Predictions receive full, partial or no credit under manual review; open and closed question groups are reported separately.","formula":"Simple accuracy = awarded credit / number of questions in the selected group","direction":"Higher is better","comparability":"Do not pool this manual partial-credit score with later exact-match or LLM-judge evaluations.","sourceIds":["vqarad-paper","vqarad-results"]},"workflow":[{"label":"Choose a single image","detail":"The paper uses one image per MedPix case.","sourceIds":["vqarad-paper"]},{"label":"Ask an image-grounded question","detail":"Question form and answer type are retained.","sourceIds":["vqarad-paper"]},{"label":"Apply the declared split","detail":"Related question forms must remain visible in the manifest.","sourceIds":["vqarad-paper"]},{"label":"Score with the stated answer policy","detail":"The original paper uses manual interpretation including partial credit.","sourceIds":["vqarad-paper"]}],"slices":[{"label":"Free-form","value":1515,"unit":"question forms","detail":"Original paper category.","sourceIds":["vqarad-paper"]},{"label":"Rephrased","value":733,"unit":"question forms","detail":"Related wording, not independent new images.","sourceIds":["vqarad-paper"]},{"label":"Framed","value":1267,"unit":"question forms","detail":"Template form counted in the headline total.","sourceIds":["vqarad-paper"]}],"sliceTitle":"Why the headline count exceeds the released record count","sliceNote":"The three paper categories sum to 3,515. Free-form plus rephrased sums to the 2,248 reported release records.","results":[{"id":"original-closed","title":"Original closed-ended free-form results","metric":"Simple accuracy","unit":"%","lower":0,"upper":100,"scope":"2018 manual evaluation of the closed-ended free-form subset within the original test; excludes paraphrase questions.","sourceIds":["vqarad-results","vqarad-paper"],"locator":"Table 2 caption, Simple Acc row (publisher URL /tables/3)","rows":[{"label":"MCB_CLEF + RAD","value":57.8,"display":"57.8%","detail":"Trained with ImageCLEF and VQA-RAD."},{"label":"SAN_CLEF + RAD","value":58.9,"display":"58.9%","detail":"Combined training sources."},{"label":"MCB_RAD","value":60.6,"display":"60.6%","detail":"VQA-RAD training."},{"label":"SAN_RAD","value":57.2,"display":"57.2%","detail":"VQA-RAD training."}],"note":"Paper-reported simple accuracy, not mean category accuracy or modern exact match. No uncertainty shown in this table."}],"analysis":[{"heading":"A record count is not a patient count","evidence":"Several questions refer to the same image.","interpretation":"Our inference: resampling question rows alone can overstate the diversity of independent visual evidence.","sourceIds":["vqarad-paper"]},{"heading":"Paraphrases enable a focused robustness question","evidence":"The original test includes linked free-form and paraphrased questions.","interpretation":"Our inference: consistency across a pair can be evaluated separately from correctness; consistently wrong answers should remain failures.","sourceIds":["vqarad-paper"]},{"heading":"The scoring rule can move the result","evidence":"Original manual scoring accepts some partial or differently specified answers.","interpretation":"Our inference: document what constitutes equivalence before comparing a generative model with historical results.","sourceIds":["vqarad-paper","vqarad-results"]}],"limitations":[{"title":"Selected single images","detail":"Performance does not establish reading of full image series or longitudinal studies.","sourceIds":["vqarad-paper"]},{"title":"Question-level test construction","detail":"A test question should not be assumed to imply an unseen image. Audit image identities in the chosen release.","sourceIds":["vqarad-paper"]},{"title":"Different downstream repackagings","detail":"Use the archived manifest and state transformations; headline counts alone do not identify a later dataset variant.","sourceIds":["vqarad-data"]}],"access":{"status":"Public official archive","license":"CC0 1.0 Universal, confirmed via OSF license API","restrictions":"This publication contains metadata and analysis only; consult the archive for assets.","url":"https://osf.io/89kps/","sourceIds":["vqarad-data"]},"sourceIds":["vqarad-paper","vqarad-results","vqarad-data"]},{"slug":"slake","name":"SLAKE","shortName":"SLAKE","version":"2021 paper; cleaned SLAKE 1.0 tracked separately","creators":"Bo Liu, Li-Ming Zhan, Li Xu and Xiao-Ming Wu","paperDate":"2021","headline":"A visual answer and a knowledge answer test different things.","summary":"SLAKE combines image questions with questions that require structured medical knowledge. That distinction makes it useful for inspecting what a multimodal model is actually being asked to contribute. We separate the image split, language setting and external-knowledge condition, then compare a small set of original vision-only baselines. We also preserve the authors’ warning that their cleaned release differs from the paper. A fair replication begins with that version choice. Our tools organize the published task families; they do not claim to measure a new model or infer clinical competence from a question-answer score.","task":{"input":"Radiology image, English or Chinese question, and optionally the provided knowledge graph","output":"A short answer in the configured answer vocabulary","unit":"One image-question pair","setting":"Image-disjoint training, validation and test splits in the original paper"},"dataOrigin":"Selected medical images, physician-supervised questions and bilingual knowledge-graph relations.","facts":[{"label":"Original images","value":"642","detail":"450 training, 96 validation and 96 test images.","sourceIds":["slake-paper"],"locator":"§2.4"},{"label":"Original bilingual QA","value":"14,028","detail":"9,849 training; 2,109 validation; 2,070 test.","sourceIds":["slake-paper"],"locator":"Table 2"},{"label":"Knowledge relations","value":"5,232","detail":"2,603 English and 2,629 Chinese triplets.","sourceIds":["slake-paper"],"locator":"§2.2–2.4"},{"label":"Release caveat","value":"SLAKE 1.0 differs","detail":"The official page explicitly describes another cleaning pass.","sourceIds":["slake-site"],"locator":"Download notice"}],"metric":{"name":"Answer accuracy by task and language","description":"The original classification evaluation separates vision-only from knowledge-based questions.","formula":"Accuracy = correctly classified answers / evaluated questions in the specified slice","direction":"Higher is better","comparability":"Fix release, language, question type, image split and access to the knowledge graph.","sourceIds":["slake-paper","slake-site"]},"workflow":[{"label":"Split by image","detail":"All questions attached to an image follow its partition.","sourceIds":["slake-paper"]},{"label":"Choose language","detail":"English and Chinese are distinct reported conditions.","sourceIds":["slake-paper"]},{"label":"Route the question","detail":"Separate vision-only from knowledge-based tasks.","sourceIds":["slake-paper"]},{"label":"Predict an answer class","detail":"Original test answers are constrained to appear in training.","sourceIds":["slake-paper"]}],"slices":[{"label":"Training","value":450,"unit":"images","detail":"Original image partition.","sourceIds":["slake-paper"]},{"label":"Validation","value":96,"unit":"images","detail":"Original image partition.","sourceIds":["slake-paper"]},{"label":"Test","value":96,"unit":"images","detail":"Original image partition.","sourceIds":["slake-paper"]}],"sliceTitle":"Original image split","sliceNote":"Use the stated counts. The paper contains inconsistent ratio wording; 450/96/96 is unambiguous and sums to 642.","results":[{"id":"vision-only","title":"Original English vision-only baselines","metric":"Accuracy","unit":"%","lower":0,"upper":100,"scope":"2021 paper, English vision-only questions, original split.","sourceIds":["slake-paper"],"locator":"Table 4","rows":[{"label":"VGG + SAN: overall","value":72.73,"display":"72.73%","detail":"All English vision-only questions."},{"label":"VGG + SAN: open-ended","value":70.34,"display":"70.34%","detail":"Open-answer slice."},{"label":"VGG + SAN: closed-ended","value":76.13,"display":"76.13%","detail":"Closed-answer slice."}],"note":"These are different slices of one model, not three competing systems. Cleaned-release scores require a separate comparison."}],"analysis":[{"heading":"Knowledge access is part of the system","evidence":"SLAKE has both vision-only and knowledge-based questions.","interpretation":"Our inference: an overall gain cannot identify improved visual grounding unless the relevant slice is inspected.","sourceIds":["slake-paper"]},{"heading":"The split supports one kind of generalization","evidence":"Images are partitioned while test answer classes occur in training.","interpretation":"Our inference: unseen-image evaluation and unseen-answer evaluation are distinct; this protocol supports the former condition.","sourceIds":["slake-paper"]},{"heading":"Cleaning changes the evaluated population","evidence":"The maintainers explicitly distinguish SLAKE 1.0 from paper data.","interpretation":"Our inference: archive counts after language filtering instead of copying the bilingual paper total into an English-only report.","sourceIds":["slake-site","slake-data"]}],"limitations":[{"title":"Finite answer vocabulary","detail":"The original classification protocol is narrower than unrestricted clinical report generation.","sourceIds":["slake-paper"]},{"title":"Mixed question requirements","detail":"Knowledge-only shortcuts and visual grounding need separate investigations; no single aggregate establishes both.","sourceIds":["slake-paper"]},{"title":"Version-sensitive counts","detail":"Paper, cleaned release and language-filtered subsets have different denominators.","sourceIds":["slake-site"]}],"access":{"status":"Public author-linked release","license":"CC BY 4.0 in author dataset-card metadata","restrictions":"Attribute the creators and specify the cleaned release when using its data.","url":"https://huggingface.co/datasets/BoKelvin/SLAKE","sourceIds":["slake-data","slake-site"]},"sourceIds":["slake-paper","slake-site","slake-data"]}],"explorer":{"kind":"coverage","title":"Inspect the visual-question task","intro":"Compare published question forms, answer formats, evidence requirements and languages. These are qualitative task annotations, not new model scores.","caution":"This is our independent analytical map of published tasks. It does not execute an evaluation, predict a model’s performance or establish clinical benefit.","sourceIds":["vqarad-paper","vqarad-results","vqarad-data","slake-paper","slake-site","slake-data"],"rows":[{"label":"VQA-RAD: free-form question","category":"Question form","input":"Image and naturally phrased question","output":"Short visual answer","metric":"Declared manual or later scoring rule","constraint":"Question novelty does not establish image novelty.","benchmarkSlug":"vqa-rad","sourceIds":["vqarad-paper"]},{"label":"VQA-RAD: paraphrase pair","category":"Question form","input":"Same image with related question wording","output":"Two answers","metric":"Correctness plus separately designed pair consistency","constraint":"Consistency is our proposed analysis, not an official extra score.","benchmarkSlug":"vqa-rad","sourceIds":["vqarad-paper"]},{"label":"VQA-RAD: closed-ended","category":"Answer format","input":"Image and closed question","output":"Restricted answer","metric":"Original simple accuracy","constraint":"Keep manual partial-credit protocol explicit.","benchmarkSlug":"vqa-rad","sourceIds":["vqarad-paper","vqarad-results"]},{"label":"VQA-RAD: open-ended","category":"Answer format","input":"Image and open question","output":"Free-text answer","metric":"Original simple accuracy","constraint":"Different wording can require semantic adjudication.","benchmarkSlug":"vqa-rad","sourceIds":["vqarad-paper"]},{"label":"SLAKE: vision-only","category":"Required evidence","input":"Image and visual question","output":"Answer class","metric":"Accuracy","constraint":"Keep the image split and language fixed.","benchmarkSlug":"slake","sourceIds":["slake-paper"]},{"label":"SLAKE: knowledge-based","category":"Required evidence","input":"Image/question and permitted knowledge graph","output":"Answer class","metric":"Accuracy","constraint":"Knowledge access is part of the system configuration.","benchmarkSlug":"slake","sourceIds":["slake-paper"]},{"label":"SLAKE: English","category":"Language","input":"English questions in a named release","output":"Answer class","metric":"Accuracy","constraint":"Do not use bilingual total as English denominator.","benchmarkSlug":"slake","sourceIds":["slake-paper","slake-site"]},{"label":"SLAKE: Chinese","category":"Language","input":"Chinese questions in a named release","output":"Answer class","metric":"Accuracy","constraint":"Language rows reuse images; they are not new patient cohorts.","benchmarkSlug":"slake","sourceIds":["slake-paper"]}],"parameters":[]},"references":[{"id":"vqarad-paper","title":"A dataset of clinically generated visual questions and answers about radiology images","organization":"Lau et al. / Scientific Data","url":"https://www.nature.com/articles/sdata2018251","note":"Original construction, released records, question variants, test selection and human scoring.","locator":"Data Records; Technical Validation; Methods","version":"2018"},{"id":"vqarad-results","title":"VQA-RAD Table 2: closed-ended free-form question results","organization":"Lau et al. / Scientific Data","url":"https://www.nature.com/articles/sdata2018251/tables/3","note":"Selected original simple-accuracy measurements; these use the paper’s manual scoring.","locator":"Table 2 caption, Simple Acc row (publisher URL /tables/3)","version":"2018"},{"id":"vqarad-data","title":"Visual Question Answering in Radiology","organization":"Lau / US National Library of Medicine","url":"https://osf.io/89kps/","note":"Official dataset archive. Its OSF API license relationship resolves to CC0 1.0 Universal.","locator":"OSF node 89kps; license 563c1cf88c5e4a3877f9e96c","version":"Accessed 2026-09-28"},{"id":"slake-paper","title":"SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering","organization":"Liu et al. / ISBI 2021","url":"https://arxiv.org/html/2102.09542","note":"Original paper counts, image split, question families and reported VGG/SAN results.","locator":"§2.4; Tables 2, 4–5","version":"2021 paper"},{"id":"slake-site","title":"SLAKE official project page","organization":"Bo Liu and Xiao-Ming Wu / Med-VQA","url":"https://www.med-vqa.com/slake/","note":"Explicitly warns that the cleaned SLAKE 1.0 release differs from the paper.","locator":"Download notice","version":"Accessed 2026-09-28"},{"id":"slake-data","title":"SLAKE author-maintained dataset card","organization":"BoKelvin / Bo Liu","url":"https://huggingface.co/datasets/BoKelvin/SLAKE","note":"Author-linked release declares CC BY 4.0 and documents validation filename change.","locator":"README metadata and Modification","version":"Accessed 2026-09-28"}]}