Learning to Learn from Context: Synthetic Training from Perturbed Public Documents
QUESTION — How can high-quality training data be synthesized from public documents to improve the context-dependent reasoning of LLMs without human annotators?
This work addresses the weakness of LLMs in learning from complex context by developing an automated pipeline that perturbs public documents to generate context-dependent reasoning traces and answers. The pipeline rewrites sources, generates dependent questions and rubrics, and filters samples. On CL-bench, SFT raises a Qwen3.6-35B-A3B student from 13.7% to 22.8%, and a subsequent rubric-reward RL stage reaches 24.6%, rivaling frontier models with over a trillion parameters.
SFT raises a Qwen3.6-35B-A3B student from 13.7% to 22.8%, and a subsequent rubric-reward RL stage reaches 24.6%, on CL-bench comparable with a frontier model of over a trillion parameters, Qwen3.8-2.4T (23.9%).
YangXiao-nlp · 27 Sept 2026
read the original ↗