CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
The authors present CodeMidas, an agentic pipeline that transforms implemented functionality in existing open-source codebases into executable RL environments using only source code as input. The pipeline allocates agentic compute to formulate behavioral specifications, construct execution-grounded tests, and validate candidate tasks through execution checks. The resulting dataset contains 5,545 training tasks from 3,185 codebases across 23 programming languages and 15 domains. Training MiMo-V2.5 on these tasks improves performance across benchmarks including DeepSWE, ProgramBench, and Terminal-Bench v2.1.
CodeMidas yields 5,545 training tasks from 3,185 open-source codebases spanning 23 programming languages and 15 technical domains.
Training MiMo-V2.5 on these tasks improves performance on DeepSWE by +11.7%.
Performance also increases on ProgramBench by +17% and Terminal-Bench v2.1 by +8.5%.