CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 101 upvotes

RULER: Instance-aware Rubric Rewards for SVG Generation

QUESTION — How can reinforcement learning policies be optimized for open-ended SVG code generation without relying on ground truth data or human preference labels?

The paper introduces RULER (Instance-aware Rubric Rewards for Reinforcement Learning), a rubric-based scoring approach to optimize Scalable Vector Graphics (SVG) code generation through reinforcement learning. RULER converts each instruction into an instance-aware rubric of six items spanning semantic, visual, and stylistic axes. A judge Vision-Language Model scores rendered rollouts item-by-item, and the weighted satisfactions are optimized via Group Relative Policy Optimization. Results on MMSVG-Illustration and MMSVG-Icon show significant improvements in rubric scores, outperforming dedicated SVG specialists and matching the substantially larger DeepSeek-V3.

Prompting a vision-language judge with a multi-axis rubric correlates with human judgments far better than scalar metrics.

RULER requires neither paired SVG ground truth nor human preference labels.

On MMSVG-Illustration and MMSVG-Icon, RULER lifts the rubric score from 0.432/0.395 to 0.693/0.683.

hangyuran · 21 Sept 2026 read the original ↗
↑