Automatic Speech Recognition: Factors Reducing Accuracy When Working with a Regional Variant of the Russian Language (using Material from Dagestan)

Main Article Content

Abstract

We evaluate the quality of automatic speech recognition (ASR) for Standard Russian and its Dagestani regional variety – a contact idiom formed in the multilingual environment of Dagestan. We tested ten off-the-shelf ASR configurations in a zero-shot setting (without fine-tuning) representing three architectural approaches (Conformer+RNN-T, Conformer+CTC, and Encoder-Decoder Transformer) on datasets of Standard Russian (GOLOS, RuDevices; approx. 20K and 90K utterances respectively) and Dagestani Russian (DaGRuS; approx. 40K utterances). All models demonstrate a word error rate (WER) degradation of 33%–50% when evaluated on Dagestani Russian, highlighting a severe domain shift problem. Error structure analysis reveals three primary degradation mechanisms: “Russification” (systematic replacement of Dagestani forms with phonetically similar Standard Russian words); deletion of discourse markers and out-of-vocabulary (OOV) terms such as ethnonyms and toponyms; and generative model hallucinations on incomprehensible audio segments. We found that end-to-end architectures suffer less degradation than cascaded ones (33%–35% vs. 40%–42%), yet no model achieves acceptable performance (minimum WER ≈ 43%). Finally, we formulate recommendations for designing ASR systems robust to the regional variability of the Russian language.

Article Details

How to Cite
Katsnelson, A. I., and O. N. Lyashevskaya. “Automatic Speech Recognition: Factors Reducing Accuracy When Working With a Regional Variant of the Russian Language (using Material from Dagestan) ”. Russian Digital Libraries Journal, vol. 29, no. 6, Oct. 2026, pp. 2058-79, doi:10.26907/1562-5419-2026-29-6-2058-2079.

References

1. Dobrushina N.R. Mnogoyazychie v Dagestane, ili zachem cheloveku tri yazyka // Sotsiologicheskii zhurnal. 2007. No. 1. P. 103–127.
2. Markl N. Language variation and algorithmic bias: Understanding algorithmic biases in British English automatic speech recognition // Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency. 2022. P. 521–534.
3. Feng S., Kudina O., Halpern B.M., Scharenborg O. Quantifying bias in automatic speech recognition // arXiv:2103.15122. https://doi.org/10.48550/arXiv.2103
4. Conneau A., Baevski A., Collobert R., Mohamed A., Auli M. Unsupervised cross-lingual representation learning for speech recognition // Proceedings of Interspeech 2021. 2021. P. 2426–2430.
5. Karpov N., Denisenko A., Minkin F. Golos: Russian dataset for speech research // Proceedings of Interspeech 2021. 2021. P. 1419–1423.
6. Dobrushina N., Daniel M., von Waldenfels R., Maisak T., Panova A. Corpus of Russian spoken in Daghestan. 2018. URL: http://www.parasolcorpus.org/dagrus/ (accessed: 24.02.2026).
7. Hinsvark A. et al. Accented speech recognition: A survey // arXiv:2104.10747. 2021. https://doi.org/10.48550/arXiv.2104.10747
8. Ramesh G.V. et al. Federated representation learning for automatic speech recognition // Proc. 3rd Symposium on Security and Privacy in Speech Communication. 2023. P. 29–33. https://doi.org/10.21437/SPSC.2023-5
9. Mann H.B., Whitney D.R. On a test of whether one of two random variables is stochastically larger than the other // The Annals of Mathematical Statistics. 1947. Vol. 18, No. 1. P. 50–60.
10. Efron B. Bootstrap methods: Another look at the jackknife // The Annals of Statistics. 1979. Vol. 7, No. 1. P. 1–26.
11. Korobov M. Morphological analyzer and generator for Russian and Ukrainian languages // Analysis of Images, Social Networks and Texts. Communications in Computer and Information Science. Vol. 542. Cham: Springer, 2015. P. 320–332.
12. Naccarato C., Panova A., Stoynova N. Word-order variation in a contact setting: A corpus-based investigation of Russian spoken in Daghestan // Language Variation and Change. 2021. Vol. 33, No. 3. P. 387–411.
13. Guerreiro N.M., Voita E., Martins A.F.T. Looking for a needle in a haystack: A comprehensive study of hallucinations in neural machine translation // Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.