CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 21 upvotes

Post-Training Leaves Behavioral Shadows on Unrelated Decisions

QUESTION — How can model capabilities be transferred through task-unrelated text without target-task examples, teacher logits, or teacher parameters?

The authors introduce Active Taskless Distillation (ATD), a method that exploits the behavioral shadow of post-training to transfer capabilities using only single-word prompt pairs where a shared public ancestor is nearly indifferent. In coding experiments using Qwen2.5-1.5B, 5,664 prompts yielded a 5.34 pp gain on HumanEval+ over an exact nuisance-matched control. The approach demonstrates broad capability transfer to scientific knowledge and commonsense reasoning without requiring teacher logits, target-task examples, or parameter access.

In the primary coding experiment with Qwen2.5-1.5B, 5,664 nses yield a 5.34 pp gain on HumanEval+ over an exact nuisance-matched control.

Lines · 24 Sept 2026 read the original ↗
↑