Xiaomi MiMo tests reinforcement learning scale with MiMo-V2.6
Xiaomi's MiMo team is running reinforcement learning training for the MiMo-V2.6 model. The run scales compute to approximately 2 billion tokens per step using 1568 prompts and 16 rollouts.
5 independent accounts
12 posts
1 articles
47,911 interactions
@AndrewCurran_
@Hesamation
@OpenRouter
@giffmana
@op7418
@srush_nlp
@teortaxesTex
@_LuoFuli
https://mimo.xiaomi.com/rl mimo.xiaomi.com