Benchmark dossiers / independent analysis

The benchmark behind the number.

Inspect what each benchmark asks, how its score is produced, and which conclusions its design can support. These are original analyses by Arcophos of the benchmark authors’ work. Any reproduced measurements are labeled with their published source and version.

Dossier01

Original 2018 paper and archive

VQA-RAD ↗

Count images, question records and question variants separately.

UnitOne image-question pairMeasureOriginal manual simple accuracy
Dossier02

2021 paper; cleaned SLAKE 1.0 tracked separately

SLAKE ↗

A visual answer and a knowledge answer test different things.

UnitOne image-question pairMeasureAnswer accuracy by task and language
What we contribute. Source reconciliation, task and metric interpretation, and tools that make assumptions inspectable. We do not claim to have created these benchmarks or run the reported models.