CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 29 upvotes

A Zeroth-Order Paradigm for LLM Preference Alignment

QUESTION — How can large language models be aligned using a zeroth-order method based on comparison oracles rather than directly optimizing a differentiable preference loss?

The authors introduce Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method that leverages comparison oracles to extract directional information from preference pairs without directly optimizing a differentiable preference loss. The approach includes an offline scheme with convergence guarantees and an online variant that utilizes unlabeled policy generations for reverse-KL control. Experiments across multiple model families demonstrate improvements over existing direct alignment methods, including metrics like length-controlled win rates.

ComPO is a zeroth-order alignment method based on comparison oracles to extract directional information.

Online ComPO uses unlabeled policy generations for reverse-KL control relative to a reference policy.

Experiments on Mistral, Llama, Gemma-2, Qwen3, and Gemma-3 models demonstrate improvements over existing direct alignment methods.

PeterLauLukCh · 16 Sept 2026 read the original ↗
↑