CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — agents 9 upvotes

Program-Verified Self-Evolution for Vision-Language Models

QUESTION — How can erroneous labels produced by majority voting or model judges be eliminated during the self-evolution of vision-language models?

The authors propose Verifiable QA Generation for Self-Evolving Models (VQS) to fix high error rates in self-evolving vision-language models caused by majority voting or model judges. Instead of voting on answers directly, VQS parses images into structured records (such as scene graphs or chart tables) and uses fixed programs to generate questions and compute answers. The model acts only as a localized visual checker for individual facts. Evaluated across ten benchmarks, VQS significantly boosts accuracy for Qwen3-VL models over multiple training rounds.

24% of majority-vote labels and 18% of model-judge labels produced during self-evolution are wrong.

Human raters find 94% of VQS answers correct, against 76% for majority voting.

Across ten benchmarks, VQS improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales.

ahmedheakl · 27 Sept 2026 read the original ↗
↑