VHD-Play generates agentic RL environments from solved mechanisms
VHD-Play is a pipeline that generates diverse agentic reinforcement learning environments by sampling and solving mathematical models before rendering their decision processes as stateful tools. This approach ensures that executable dynamics and trajectory-scoring references are inherited directly from the solved model, addressing the misalignment issues common in existing generation pipelines.