CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — evaluation 10 upvotes

WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents

QUESTION — How can we measure and improve the experimental understanding of AI agents regarding how component changes affect outcomes?

The study introduces WhatWorkedBench to measure experimental understanding, defined as the accuracy of predictions about component changes after budgeted experimentation. Agents inspect code, select measurements, and submit a response surface table predicting scores for every configuration. Exhaustive CPU execution provides reference effects across 36 tasks from 30 data sources and 8 workflow types, containing 1,248 configuration records. Fitting a Gaussian process (GP) to agent observations raises effect recovery from 0.632 to 0.698 in the original Flash cohort and from 0.621 to 0.720 in an additional cohort. Encoding code equivalences further raises GP recovery from 0.248 to 0.462 on specific workflows.

Exhaustive CPU execution supplies reference effects across 36 tasks from 30 data sources and 8 workflow types, with 1248 configuration records.

Fitting a Gaussian process (GP) to the same agent observations raises effect recovery from 0.632 to 0.698 in the original Flash cohort and from 0.621 to 0.720 in an additional cohort.

Encoding code equivalences raises GP recovery from 0.248 to 0.462 on six workflows with six binary options at 20 new measurements.

ethanning · 23 Sept 2026 read the original ↗
↑