ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
This paper introduces ProgramDistill, a benchmark evaluating coding agents on features discovered through interaction with fully functional reference applications. Using an automated mine-craft-patch pipeline, the authors discover 1,975 replay-verified behaviors across 26 applications and construct 4,063 tasks without human intervention. Across nine frontier coding agents, GPT-6 Astra and Claude Opus 5 achieve 49.2% and 28.8% success on cumulative workflows in full-application reconstruction, while partial-application success declines as restoration depth increases.
The mine-craft-patch pipeline discovers 1,975 replay-verified behaviors across 26 applications and constructs 4,063 tasks automatically.
GPT-6 Astra and Claude Opus 5 achieve 49.2% and 28.8% success on cumulative workflows in full-application reconstruction.
In partial-application reconstruction, success falls from 100% to 64.0% and from 96% to 32% as restoration depth increases from 1 to 8.