Researchers introduce Visual Answerability Diagnosis with Rationales (VAD-R), a new benchmark designed to evaluate vision-language models' ability to abstain from unanswerable questions without shortcut cues. The study reveals that while hidden states can distinguish answerability, current models fail to translate this into explicit decisions.
- VAD-R uses a two-stage pipeline of shortcut filtering and quality verification with step-by-step rationales.
- State-of-the-art open-source VLMs show limited spontaneous abstention with average recall rates of only 11.4%.
- The authors propose Rep2Act, a method to align latent answerability awareness with explicit abstention actions.
- Rep2Act improves action accuracy on VAD-R from roughly 56-59% to over 86-88% for Qwen2.5-VL models.
- On the out-of-distribution TUBench, Rep2Act achieves an average F1 score of 53.3%, surpassing GPT-4 Turbo and GPT-4o.
This approach bridges the gap between representation and action, enabling VLMs to effectively recognize when they should not answer a question.