Medical visual question answering

The image, the question, the split.

Independent analysis of VQA-RAD and SLAKE: original dataset counts, question variants, image splits, visual grounding, knowledge access and paper-reported results.

Visual question answering / dataset anatomy

Keep images and question forms separate

VQA-RAD

315 images [1]

Original question forms

Free-form
1,515 [1]
Rephrased
733 [1]
Framed
1,267 [1]

3,515 question forms in the listed groups.

SLAKE

642 original images [4]

Original image partitions

Training
450 [4]
Validation
96 [4]
Test
96 [4]

642 images in the listed groups.

The two bars have different units. VQA-RAD lists question forms; SLAKE lists image partitions. Read the split and release caveats →

01 /

Counting units

Distinguish images, question records and related question forms.

02 /

Task evidence

Separate image understanding from external-knowledge access.

03 /

Comparison policy

Keep release, language, split and answer judging visible.

Our analytical question

A medical VQA result depends on which images, question forms and answer rules were evaluated. We examine VQA-RAD’s clinician-generated questions and SLAKE’s visual and knowledge-based tasks, keeping paper counts separate from cleaned releases. Our original coverage map and guides explain what those choices support. Selected historical results are credited to the original papers; no new model runs are claimed.

Read the original evidence closely

The benchmark, unpacked.

Primary source library ↗
Dossier01

Original 2018 paper and archive

VQA-RAD ↗

Count images, question records and question variants separately.

UnitOne image-question pairMeasureOriginal manual simple accuracy
Dossier02

2021 paper; cleaned SLAKE 1.0 tracked separately

SLAKE ↗

A visual answer and a knowledge answer test different things.

UnitOne image-question pairMeasureAnswer accuracy by task and language

An original analytical tool

Inspect the visual-question task

Evidence explorer

Compare published question forms, answer formats, evidence requirements and languages. These are qualitative task annotations, not new model scores.

8 of 8 evidence entries shown

Question form

VQA-RAD: free-form question

Read dossier ↗
Input
Image and naturally phrased question
Output
Short visual answer
Measure
Declared manual or later scoring rule
Interpretation boundary

Question novelty does not establish image novelty.

[1]
Question form

VQA-RAD: paraphrase pair

Read dossier ↗
Input
Same image with related question wording
Output
Two answers
Measure
Correctness plus separately designed pair consistency
Interpretation boundary

Consistency is our proposed analysis, not an official extra score.

[1]
Answer format

VQA-RAD: closed-ended

Read dossier ↗
Input
Image and closed question
Output
Restricted answer
Measure
Original simple accuracy
Interpretation boundary

Keep manual partial-credit protocol explicit.

[1][2]
Answer format

VQA-RAD: open-ended

Read dossier ↗
Input
Image and open question
Output
Free-text answer
Measure
Original simple accuracy
Interpretation boundary

Different wording can require semantic adjudication.

[1]
Required evidence

SLAKE: vision-only

Read dossier ↗
Input
Image and visual question
Output
Answer class
Measure
Accuracy
Interpretation boundary

Keep the image split and language fixed.

[4]
Required evidence

SLAKE: knowledge-based

Read dossier ↗
Input
Image/question and permitted knowledge graph
Output
Answer class
Measure
Accuracy
Interpretation boundary

Knowledge access is part of the system configuration.

[4]
Language

SLAKE: English

Read dossier ↗
Input
English questions in a named release
Output
Answer class
Measure
Accuracy
Interpretation boundary

Do not use bilingual total as English denominator.

[4][5]
Language

SLAKE: Chinese

Read dossier ↗
Input
Chinese questions in a named release
Output
Answer class
Measure
Accuracy
Interpretation boundary

Language rows reuse images; they are not new patient cohorts.

[4]

This is our independent analytical map of published tasks. It does not execute an evaluation, predict a model’s performance or establish clinical benefit. [1][2][3][4][5][6]

Original analysis / methods and interpretation

What the score leaves unsaid.

All analyses →

Questions, answered

Read the result in context.

Specific tasks. Stated conditions.
Inspect every source.

Why does VQA-RAD have both 3,515 and 2,248 counts?

The paper’s headline includes free-form, rephrased and framed question forms. Its Data Records section describes 2,248 elements corresponding to the free-form and rephrased categories.

Does a held-out VQA-RAD question imply a new image?

No such assumption follows from a question-level split. Audit image identifiers and linked question forms in the chosen release.

Does SLAKE 1.0 exactly match the paper?

The authors’ project page explicitly warns that the cleaned release may differ. State the release and count records after applying the language filter.

Do these sites reproduce the official leaderboards?

No. We provide independent analysis and selected historical paper-reported measurements, with their original conditions and limitations.

Working tool / saved on this device

Prepare a benchmark comparison brief

Interactive worksheet

Use this secondary checklist to document a run or literature comparison after inspecting the named benchmark conditions. Completion records documentation, not performance.

Name the release and units

Evidence you can inspect. Benchmark dossiers distinguish published facts from our interpretation, with source versions and access notes attached.

Download the evidence ↗