Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
The authors introduce VHD-Play, a pipeline that reverses standard generation by sampling and solving a mathematical model before a corpus-grounded setter renders its decision process as stateful tools. Executable dynamics and trajectory-scoring references are inherited from the same solved model. The pipeline produces 3,300 diverse agentic environments at a cost of a few cents each. Training Qwen3.6-35B-A3B on three families raises its mean agentic score from 0.204 to 0.815 in a five-family diagnostic, with generalization to external function calling, travel planning, and e-commerce benchmarks where it even outperforms Qwen3.7-Max.
The pipeline produces 3,300 diverse agentic environments at a cost of a few cents each.
Training Qwen3.6-35B-A3B on three families raises its mean agentic score from 0.204 to 0.815 in a five-family diagnostic.
On E-Commerce Bench, the trained checkpoint completes every run without bankruptcy and exceeds Qwen3.7-Max.