Karakalpak speech recognition
A Karakalpak speech-to-text model. The word error rate is not yet at release quality, so we are holding the launch and fixing the data instead of shipping something that mishears people.
The challenge
Recognition is harder than synthesis for a low-resource language: it needs many speakers, many recording conditions and many dialect variants. Our current error rate is too high for the uses people would immediately put it to — transcription of meetings, appeals and public records.
Our approach
Diagnose, do not patch
Error analysis by speaker, dialect and recording condition instead of chasing a single benchmark number.
Grow the corpus where it fails
Targeted recording of the speaker groups and conditions the model handles worst.
Release when it is honest
It joins the karakalpakvoice.uz API only when the error rate is low enough for real transcription work.
We would rather publish late than publish a model that mishears the language. This page will be updated when the error rate is release-worthy.