Microsoft Research has introduced MindTopo, a new benchmark designed to evaluate whether multimodal large language models possess topological intuition regarding connectivity, enclosure, order, separation, and knots.

  • The benchmark organizes tasks into five categories: continuity, separation, order, enclosure, and knots.
  • It measures performance at two cognitive levels: reasoning on static scenes and planning through interactive sequences of actions.
  • All scenes are generated from controlled simulators to provide exact ground truth and adjustable difficulty.
  • Current models perform significantly better on static recognition tasks than on interactive planning tasks.
  • Failures in planning often stem from losing track of structural relationships as scenes change or proposing actions that violate physical constraints.

The findings highlight a substantial gap between recognizing topology in a static image and maintaining an understanding of it while acting, suggesting an opportunity to advance AI systems for robotics and interactive environments.