Building a Production Greek-English Speech Recognizer
This engineering report details the development of Sophea, a production bilingual Greek-English automatic speech recognition system evaluated against nine production gates. The authors found that no single training-data composition could pass all nine gates simultaneously. They developed a six-stage data pipeline that reduced the discarded share of scored Greek audio from 98.7 percent to 10.6 percent. Additionally, a three-model ROVER ensemble increased gate coverage to 9 of 9 and reduced overlapping-speech WER from 53.35 percent to 37.87 percent, yielding a 29 percent relative improvement.
An audio-quality filter reduced the discarded share of scored Greek audio from 98.7 percent to 10.6 percent.
A three-model ROVER ensemble reduced overlapping-speech WER from 53.35 percent to 37.87 percent, a 29 percent relative improvement.
It reaches 4.26 percent average WER across eight public English test sets and 25.88 percent WER on live Greek noisy-environment traffic.