CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — evaluation 4 upvotes

Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation

QUESTION — How can we improve accuracy and reduce costs when performing disaggregated evaluation of AI systems?

The paper addresses disaggregated AI evaluation by treating evaluation sets as finite populations to achieve accurate domain estimates without exhaustive testing. The authors propose prediction-powered smoothing (PP-S), a Bayesian model fit to each domain's prediction-powered estimate, along with an extension (PP-TS) that borrows strength across taxonomies. For validation, they derive a new design-based cross-validation score to select among estimators. Evaluated on curated benchmarks and deployed agent traffic, these estimators improve point and interval estimation with near-nominal coverage while matching the selection efficacy of independent validation samples.

The proposed estimators improve on the direct estimators in point and interval estimation, with near-nominal coverage.

Derives a new, approximately unbiased design-based cross-validation score for choosing among direct and smoothed estimators.

At the same sampling budget, the proposed score selects as well as an independent validation sample does and estimates the selected estimator's error far more accurately.

kokomarina · 17 Sept 2026 read the original ↗
↑