CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 55 upvotes

All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

QUESTION — How can we build an all-in-one multilingual scene text recognizer that is lighter and more accurate than both VLMs and per-language experts?

The authors introduce TextMuSS-10M, a large synthetic dataset spanning 10 scripts and 229 languages, alongside ScriptMoE, a script-aware Mixture-of-Experts architecture. ScriptMoE shares a single visual encoder and replaces the dense decoder with a sparse MoE block using an image-level router and a shared expert. Experiments on TextMuSS-Bench show that ScriptMoE achieves an 82.06% accuracy, outperforming the strongest baseline by 1.31%. On the CC-OCR task, replacing the recognizer in PP-OCRv5 with ScriptMoE lifts the F1 score from 65.71% to 80.89%, slightly surpassing the best VLM at a fraction of the parameter count.

ScriptMoE achieves the highest accuracy of 82.06%, outperforming the strongest STR baseline by 1.31%.

On the CC-OCR end-to-end multilingual task, replacing only the recognizer in PP-OCRv5 with ScriptMoE lifts F1 score from 65.71% to 80.89%, slightly surpassing the best VLM (80.73%) at a fraction of the parameter count.

Yesianrohn · 21 Sept 2026 read the original ↗
↑