Microsoft Research has introduced MindTopo, a new benchmark designed to evaluate whether multimodal large language models possess topological intuition regarding connectivity, enclosure, order, separation, and knots.
- The benchmark organizes tasks into five categories: continuity, separation, order, enclosure, and knots.
- It measures performance at two cognitive levels: reasoning on static scenes and planning through interactive sequences of actions.
- All scenes are generated from controlled simulators to provide exact ground truth and adjustable difficulty.
- Current models perform significantly better on static recognition tasks than on interactive planning tasks.
- Failures in planning often stem from losing track of structural relationships as scenes change or proposing actions that violate physical constraints.
The findings highlight a substantial gap between recognizing topology in a static image and maintaining an understanding of it while acting, suggesting an opportunity to advance AI systems for robotics and interactive environments.