Independent analysis of VQA-RAD and SLAKE: original dataset counts, question variants, image splits, visual grounding, knowledge access and paper-reported results.
Distinguish images, question records and related question forms.
02 /
Task evidence
Separate image understanding from external-knowledge access.
03 /
Comparison policy
Keep release, language, split and answer judging visible.
Our analytical question
A medical VQA result depends on which images, question forms and answer rules were evaluated. We examine VQA-RAD’s clinician-generated questions and SLAKE’s visual and knowledge-based tasks, keeping paper counts separate from cleaned releases. Our original coverage map and guides explain what those choices support. Selected historical results are credited to the original papers; no new model runs are claimed.
This is our independent analytical map of published tasks. It does not execute an evaluation, predict a model’s performance or establish clinical benefit. [1][2][3][4][5][6]
A task-specific audit of images, answer shortcuts and scoring rules in two medical VQA benchmarks.
5 min read
Questions, answered
Read the result in context.
Specific tasks. Stated conditions. Inspect every source.
Why does VQA-RAD have both 3,515 and 2,248 counts?+
The paper’s headline includes free-form, rephrased and framed question forms. Its Data Records section describes 2,248 elements corresponding to the free-form and rephrased categories.
Does a held-out VQA-RAD question imply a new image?+
No such assumption follows from a question-level split. Audit image identifiers and linked question forms in the chosen release.
Does SLAKE 1.0 exactly match the paper?+
The authors’ project page explicitly warns that the cleaned release may differ. State the release and count records after applying the language filter.
Do these sites reproduce the official leaderboards?+
No. We provide independent analysis and selected historical paper-reported measurements, with their original conditions and limitations.
Working tool / saved on this device
Prepare a benchmark comparison brief
Interactive worksheet
Use this secondary checklist to document a run or literature comparison after inspecting the named benchmark conditions. Completion records documentation, not performance.
Evidence you can inspect. Benchmark dossiers distinguish published facts from our interpretation, with source versions and access notes attached.