RESEARCH NOTE · PAPER YEAR 2022
Measuring CLEVRness
High accuracy on a visual question-answering benchmark does not tell the whole story. This first-author work with Henryk Michalewski and Mateusz Malinowski uses black-box evaluation to investigate the reasoning capabilities of visual models.
The ICLR 2022 paper explores a dual-agent setup: a reasoning agent and a scene-manipulating agent. Their interaction helps expose model biases without relying on direct access to model gradients or probabilities.
Read “Measuring CLEVRness: Black-box Testing of Visual Reasoning Models”