Coding Agents for Generalized Task and Motion Planning Problems
This paper investigates whether coding agents can automate generalized task and motion planning (TAMP) by synthesizing programs within a fixed budget using task descriptions and simulator access. Evaluating Claude Code (Opus 5) and Codex (GPT-5.6 Sol and GPT-6 Astra) across 28 simulated environments from KinDER and PDDLStream over 98,000 evaluation episodes, the authors find that coding agent configurations outperform hand-engineered planners and one-shot generation baselines. Specifically, the agents achieve mean success rates ranging from 56% to 95% versus 47% for standard planners, while requiring an order of magnitude less computation per instance.
Across all program synthesis methods, we evaluate 980 generated programs on 100 held-out instances each, 98,000 evaluation episodes in total.
All three agent configurations outperform hand-engineered planners, one-shot generation, and an LLM-based generalized planning baseline in mean success (56% to 95% versus 47% for the planners, on the 16 environments where a planner is available).
As object counts grow, the agents' programs maintain higher success than the planner, using an order of magnitude less computation per instance on average.