Can Agents Design Libraries for Agents?
The paper introduces LibraryDesignBench, a two-phase benchmark with 242 expert-validated programming problems to evaluate how agents design code libraries from specifications. Findings show that agent designers successfully reproduce human-written abstractions in 11 out of 15 tasks. However, downstream agents underuse these libraries and write extra code mainly because agent-written libraries are rigid or hard to use. The study demonstrates that providing agent-first guidance and using subagents for testing improves downstream scores and yields simpler programs.
The benchmark spans 242 expert-validated programming problems across 15 library-design tasks in four languages.
On eleven of the fifteen tasks, agent designers reproduce the abstractions of the human-written production library.