CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — agents 61 upvotes

ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

QUESTION — How can coding agents be evaluated on complex web development features discovered through interaction with fully functional reference applications?

This paper introduces ProgramDistill, a benchmark evaluating coding agents on features discovered through interaction with fully functional reference applications. Using an automated mine-craft-patch pipeline, the authors discover 1,975 replay-verified behaviors across 26 applications and construct 4,063 tasks without human intervention. Across nine frontier coding agents, GPT-6 Astra and Claude Opus 5 achieve 49.2% and 28.8% success on cumulative workflows in full-application reconstruction, while partial-application success declines as restoration depth increases.

The mine-craft-patch pipeline discovers 1,975 replay-verified behaviors across 26 applications and constructs 4,063 tasks automatically.

GPT-6 Astra and Claude Opus 5 achieve 49.2% and 28.8% success on cumulative workflows in full-application reconstruction.

In partial-application reconstruction, success falls from 100% to 64.0% and from 96% to 32% as restoration depth increases from 1 to 8.

beanie00 · 16 Sept 2026 read the original ↗
↑