LEGO-Anything: Coding Agents for 3D Scene Reconstruction
The paper presents LEGO-Anything, an Image-to-Code framework where a coding agent iteratively writes and executes Blender code to reconstruct explicit 3D scene programs. To evaluate performance, the authors introduce LEGO-Bench, a simulator-grounded benchmark comprising 208 images from 104 indoor and outdoor scenes. GPT-6-astra achieves the strongest results with 53.4% indoor and 39.6% outdoor scores. Addressing recurring agent flaws like poor initialization and unreliable self-evaluation, the authors introduce LEGO-Plugin, a training-free harness plugin yielding relative gains of up to 62.7% in overall score across six evaluated models.
LEGO-Bench is a simulator-grounded benchmark with 208 images from 104 diverse indoor and outdoor scenes.
GPT-6-astra achieves the strongest overall results, with 53.4% indoor and 39.6% outdoor scores.
LEGO-Plugin improves all six evaluated models, with relative gains of up to 62.7% in overall score.