Recognition is harder than synthesis for a language like ours. Synthesis needs a few clean voices; recognition needs many speakers, many recording conditions and the dialect variation that real speech carries.
Our current word error rate is too high for transcription of meetings, appeals or public records — the exact tasks people would point it at on day one. Publishing it now would mean publishing a model that mishears the language, and that damages trust in everything else we release.
So the work is data work again: error analysis by speaker group and recording condition, then targeted recording where the model performs worst. Not a bigger model — better coverage.
When the error rate is release-worthy, speech-to-text joins the same API as text-to-speech. Until then this page is the honest status.