CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 2 upvotes

Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition

QUESTION — How can we simultaneously improve reasoning accuracy and inference efficiency in language models by combining independently learned capabilities?

The authors introduce Lightning Weave, a post-training framework that extracts and composes independently learned capabilities into a single student through on-policy distillation. By combining aligned log-ratio shifts at shared student token states and using Tilted-Target DOPD, the framework creates stable learning targets without requiring multiple live anchor models during training. Experiments across mathematics and code show substantial efficiency and accuracy gains, such as on Qwen3.5-4B raising HMMT 2025 accuracy from 59.2% to 64.0% with 10.7% fewer response tokens, and LiveCodeBench v5 from 41.7% to 54.2% with 9.6% fewer response tokens.

On Qwen3.5-4B, it raises HMMT 2025 accuracy from 59.2% to 64.0% with 10.7% fewer response tokens, and LiveCodeBench v5 accuracy from 41.7% to 54.2% with 9.6% fewer response tokens.

gbcfchc · 13 Sept 2026 read the original ↗
↑