Researchers identify Relevant Visual Information Shift (RVIS) during decoding as the primary cause for the failure of existing visual token pruning methods in Multimodal Large Language Models (MLLMs). To address this, they propose Decoding-stage Shift-aware Token Pruning (DSTP), a training-free framework that aligns visual tokens with shifting reasoning requirements.

  • DSTP is a training-free add-on that corrects RVIS during the decoding stage.
  • It significantly reduces performance degradation of pruning methods in complex reasoning tasks.
  • The method yields consistent performance gains across visual understanding benchmarks.
  • DSTP demonstrates generalizability and efficiency with minimal computational overhead across diverse state-of-the-art architectures.

This approach allows existing pruning techniques to maintain reliability when handling the vast number of visual tokens required for complex visual reasoning.