Program-Verified Self-Evolution for Vision-Language Models
The authors propose Verifiable QA Generation for Self-Evolving Models (VQS) to fix high error rates in self-evolving vision-language models caused by majority voting or model judges. Instead of voting on answers directly, VQS parses images into structured records (such as scene graphs or chart tables) and uses fixed programs to generate questions and compute answers. The model acts only as a localized visual checker for individual facts. Evaluated across ten benchmarks, VQS significantly boosts accuracy for Qwen3-VL models over multiple training rounds.
24% of majority-vote labels and 18% of model-judge labels produced during self-evolution are wrong.
Human raters find 94% of VQS answers correct, against 76% for majority voting.
Across ten benchmarks, VQS improves Qwen3-VL by up to 3.18 points at the 2B, 4B, and 8B scales.