All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
The authors introduce TextMuSS-10M, a large synthetic dataset spanning 10 scripts and 229 languages, alongside ScriptMoE, a script-aware Mixture-of-Experts architecture. ScriptMoE shares a single visual encoder and replaces the dense decoder with a sparse MoE block using an image-level router and a shared expert. Experiments on TextMuSS-Bench show that ScriptMoE achieves an 82.06% accuracy, outperforming the strongest baseline by 1.31%. On the CC-OCR task, replacing the recognizer in PP-OCRv5 with ScriptMoE lifts the F1 score from 65.71% to 80.89%, slightly surpassing the best VLM at a fraction of the parameter count.
ScriptMoE achieves the highest accuracy of 82.06%, outperforming the strongest STR baseline by 1.31%.
On the CC-OCR end-to-end multilingual task, replacing only the recognizer in PP-OCRv5 with ScriptMoE lifts F1 score from 65.71% to 80.89%, slightly surpassing the best VLM (80.73%) at a fraction of the parameter count.