CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — agents 18 upvotes

Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms

QUESTION — How can diverse agentic RL environments with dependable outcome signals and low extension cost be systematically generated?

The authors introduce VHD-Play, a pipeline that reverses standard generation by sampling and solving a mathematical model before a corpus-grounded setter renders its decision process as stateful tools. Executable dynamics and trajectory-scoring references are inherited from the same solved model. The pipeline produces 3,300 diverse agentic environments at a cost of a few cents each. Training Qwen3.6-35B-A3B on three families raises its mean agentic score from 0.204 to 0.815 in a five-family diagnostic, with generalization to external function calling, travel planning, and e-commerce benchmarks where it even outperforms Qwen3.7-Max.

The pipeline produces 3,300 diverse agentic environments at a cost of a few cents each.

Training Qwen3.6-35B-A3B on three families raises its mean agentic score from 0.204 to 0.815 in a five-family diagnostic.

On E-Commerce Bench, the trained checkpoint completes every run without bankruptcy and exceeds Qwen3.7-Max.

taesiri · 23 Sept 2026 read the original ↗
↑