Automatic morphological tagging for Ossetic based on data from the Corpus of Oral Texts

Main Article Content

Anna Sergeevna Shatskikh
Alexey Andreevich Sorokin

Abstract

In this work we present the first morphologically annotated corpus for Iron Ossetic that conforms to the Universal Dependencies schema. The corpus includes 5454 manually annotated sentences from the Iron Ossetic Corpus of Oral Texts, containing 74032 tokens. We use this corpus to train a BERT-based morphological analyzer. The analyzer achieves tag accuracy of 95.60%. Furthermore, the paper presents the results of experiments aimed at improving classification quality. It is shown that neither filtering the model’s output using a context‑free analyzer nor multi‑task approach lead to a significant improvement. Finally, the paper presents the results of testing the analyzer on out‑of‑domain data and shows that the model achieves tag accuracy of 91.45% on them.

Article Details

How to Cite
Shatskikh, A. S., and A. A. Sorokin. “Automatic Morphological Tagging for Ossetic Based on Data from the Corpus of Oral Texts”. Russian Digital Libraries Journal, vol. 29, no. 6, Oct. 2026, pp. 2215-39, doi:10.26907/1562-5419-2026-29-6-2215-2239.

References

1. Ethnologue. URL: http://www.ethnologue.com/show_language.asp?code=oss
2. Осетинский национальный корпус. URL: http://corpus.ossetic-studies.org
3. Устный корпус осетинского. URL: https://www.ossetic-studies.org/ru/texts/iron
4. Arkhangelskiy T., Belyaev O., Vydrin A. The Creation of Large-Scale Annotated Corpora of Minority Languages using UniParser and the EANC platform. // Proceedings of COLING 2012: Posters, Mumbai, India. The COLING 2012 Organizing Committee: 2012 (P. 83–92).
5. Nivre J., de Marneffe M., Ginter F., Hajic J., Manning C., Pyysalo S., Schuster S., Tyers F., Zeman D. Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection. // Proceedings of the Twelfth Language Resources and Evaluation Conference, Marseille, France. European Language Resources Association: 2020 (P. 4034–4043).
6. Abaev V. A grammatical sketch of Ossetic. The Hague: Mouton, 1964.
7. Беляев О.И. Коррелятивная конструкция в осетинском языке (диссертация на соискание степени кандидата филологических наук). 2014.
8. Belyaev O. Evolution of Case in Ossetic. // Iran and the Caucasus. 2010. 14(2). С. 287–322.
9. Serdobolskaya N., Tuzhik O. Differential Object Marking in Modern Ossetic: Referential Properties and Animacy // Tomsk Journal of Linguistics and Anthropology. 2025. 48(2). P. 91–107.
10. Kondratyuk D., Straka M. 75 Languages, 1 Model: Parsing Universal Dependencies Universally. // Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China. Association for Computational Linguistics: 2019 (P. 2779–2795).
11. Straka M., Straková J., Hajic, J. UDPipe at SIGMORPHON 2019: Contextualized Embeddings, Regularization with Morphological Categories, Corpora Merging. // Proceedings of the 16th Workshop on Computational Research in Phonetics, Phonology, and Morphology, Florence, Italy. Association for Computational Linguistics: 2019 (P. 95–103).
12. Devlin J., Chang M., Lee K., Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. // Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, Minnesota. Association for Computational Linguistics: 2020 (P. 4171–4186).
13. google-bert/bert-base-multilingual-cased. URL: https://huggingface.co/google-bert/bert-base-multilingual-cased
14. Kuratov, Y., Arkhipov, M. Adaptation of Deep Bidirectional Multilingual Transformers for Russian Language. 2019.
15. Hulden M. Foma: a Finite-State Compiler and Library // Proceedings of the Demonstrations Session at EACL 2009, Athens, Greece. Association for Computational Linguistics: 2009 (P. 29—32)
16. Inoue G., Shindo H., Matsumoto Y. Joint Prediction of Morphosyntactic Categories for Fine-Grained Arabic Part-of-Speech Tagging Exploiting Tag Dictionary Information. // Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), Vancouver, Canada. Association for Computational Linguistics: 2017 (P. 421–431).
17. Ossetic-COT. URL: https://github.com/ania3000/Ossetic-COT
18. ossetic-encoders/ossbert-morph. URL: https://huggingface.co/ossetic-encoders/ossbert-morph