Rufus-Air: An Open LLM Post-Training Recipe
The authors present Rufus-Air, an open post-training recipe structured as an eight-stage serial pipeline built on GLM-4.5-Air-Base (106B-A12B), progressing from SFT and specialized RL stages to RLHF. The approach orders stages by capability progression and reward reliability, leveraging public data and components without requiring new human annotations or an in-house distillation teacher. Evaluations demonstrate that Rufus-Air outperforms the official GLM-4.5-Air post-trained release and remains competitive with similarly sized open models.
Diverse, high-quality SFT establishes a strong capability floor.
Difficulty filtering keeps RL prompts within a productive learning range.
Reward reliability provides a practical principle for ordering stages.