Xiaomi MiMo teste l'échelle d'apprentissage par renforcement avec MiMo-V2.6
Xiaomi MiMo tests reinforcement learning scale with MiMo-V2.6
L'équipe MiMo de Xiaomi mène actuellement un entraînement par renforcement pour le modèle MiMo-V2.6, en portant la puissance de calcul à environ 2 milliards de tokens par étape avec 1568 invites.
5 comptes indépendants
12 messages
1 articles
47 911 interactions
@AndrewCurran_
@Hesamation
@OpenRouter
@giffmana
@op7418
@srush_nlp
@teortaxesTex
@_LuoFuli
https://mimo.xiaomi.com/rl mimo.xiaomi.com