Xiaomi scales reinforcement learning runs for MiMo-V2.6
Xiaomi's MiMo team is running large-scale reinforcement learning training for the MiMo-V2.6 model, utilizing approximately 2 billion tokens per step. The process scales compute across environments, harnesses, and automated grading systems.
4 independent accounts
11 posts
1 articles
47,811 interactions
@AndrewCurran_
@Hesamation
@giffmana
@op7418
@srush_nlp
@teortaxesTex
@_LuoFuli
https://mimo.xiaomi.com/rl mimo.xiaomi.com