A position paper argues that modern AI systems should shift from reactive maintenance flywheels to proactive, test-driven development to improve generalization. Current reactive pipelines rely on observing user errors to patch models, which often ignores broader system objectives and fails to preempt future edge cases.
The authors advocate for creating a "test space" to map feedback data to task objectives, thereby evolving the flywheel from reactive to proactive. They mathematically prove that this proactive approach achieves better long-term scaling with fewer iterations compared to reactive methods.
This shift addresses the statistical difficulty of collecting remaining errors in open-world use cases and aims to build more generalizable AI systems.