CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 19 upvotes

Rufus-Air: An Open LLM Post-Training Recipe

QUESTION — How can an open, reproducible post-training recipe be engineered for a 106B-A12B base model using an eight-stage serial pipeline?

The authors present Rufus-Air, an open post-training recipe structured as an eight-stage serial pipeline built on GLM-4.5-Air-Base (106B-A12B), progressing from SFT and specialized RL stages to RLHF. The approach orders stages by capability progression and reward reliability, leveraging public data and components without requiring new human annotations or an in-house distillation teacher. Evaluations demonstrate that Rufus-Air outperforms the official GLM-4.5-Air post-trained release and remains competitive with similarly sized open models.

Diverse, high-quality SFT establishes a strong capability floor.

Difficulty filtering keeps RL prompts within a productive learning range.

Reward reliability provides a practical principle for ordering stages.

lawhy · 24 Sept 2026 read the original ↗
↑