Until now, a Karakalpak voice was something every project had to build for itself — and almost nobody did. Accessibility tools, public announcement systems, kiosks and education software all defaulted to another language, or to no voice at all.
The model is available as an API at karakalpakvoice.uz. It covers standard text-to-speech and voice cloning: a specific voice can be reproduced from a short reference recording, which matters for institutions that want one consistent voice across their services.
The hard part was not the model. It was the data. Karakalpak speech recordings of usable quality and licensing barely existed in public form, so the first months of the project were recording, cleaning and verification work with native speakers rather than training runs.
Speech-to-text is the natural next step and it is in research. We are not publishing it until the error rate is low enough for real transcription work — the reasoning is in a separate note.