Architecture of the Linguistic Platform Turklang
Main Article Content
Abstract
The article describes the architecture of the TurkLang multifunctional linguistic platform, currently being developed at the Institute of Applied Semiotics of the Academy of Sciences of the Republic of Tatarstan. The system is based on a three-level knowledge base: linguistic universals, descriptions of units of specific languages, and empirical speech data. The knowledge base is implemented as a graph, where the primary storage unit is the morpheme – the minimal meaningful unit of language (akin to an atomic data element). This design choice is driven by the element-combinatorial nature of Turkic languages: a word form (the concrete realization of a word in text) is a linear chain consisting of a root and sequentially attached affixes – modifiers, each expressing a single grammatical category (case, number, tense, etc.) in a strictly fixed order. The paper examines levels of abstraction, types of relationships between entities, and mechanisms for integration with external sources (an electronic encyclopedia, lexicographic resources, and GIS data). Special attention is paid to the application of international standards for syntactic (Universal Dependencies), semantic (Universal Meaning Representation), and pragmatic (DiAML) annotation. A comparison with existing platforms (Sketch Engine, LingvoDoc, Apertium) is provided. To date, the platform's knowledge base covers more than 30 languages and dialects, contains over 237,000 root morphemes and about 40,000 compatibility rules, while the corpus component comprises over 180 million of word occurrences.
Article Details
References
2. Koskenniemi K. Two-Level Morphology: A General Computational Model for Word-Form Recognition and Production. Helsinki: University of Helsinki, 1983. 160 p.
3. Apresyan Y.D. Izbrannye trudy. Tom II. Integral'noe opisanie yazyka i sistemnaya leksikografiya. Moscow: Shkola “Yazyki russkoi kul’tury”, 1995. 769 p.
4. Krauwer S. The basic language resource kit (BLARK) as the first milestone for the language resources roadmap // Proc. International workshop on speech and computer SPECOM-2003, Moscow, Russia, Oct. 2003. P. 8–15.
5. Kilgarriff A., Rychlý P., Smrž P., Tugwell D. The Sketch Engine // Proceedings of the Eleventh EURALEX International Congress, Lorient, France, Jul. 2004. P. 105–116.
6. Normanskaya Y.V., Borisenko O.D., Beloborodov I.B., Avetisyan A.I. The Software System LingvoDoc and the Possibilities It Offers for Documentation and Analysis of Ob-Ugric Languages // Doklady Mathematics. 2022. Vol. 105, No. 3. P. 187–206. https://doi.org/10.31857/S2686954322030055
7. Khanna, T., Washington, J.N., Tyers, F.M et al. Recent advances in Apertium, a free/open-source rule-based machine translation platform for low-resource languages // Machine Translation. 2021. Vol. 35. P. 475–502. https://doi.org/10.1007/s10590-021-09260-6
8. Bakay Ö., Ergelen Ö., Sarmış E., Yıldırım S., et al. Turkish WordNet KeNet // Proceedings of the 11th Global Wordnet Conference, Virtual, Jan. 2021. P. 166–174. https://doi.org/10.18653/v1/2021.gwc-1.19
9. Agostini A., Usmanov T., Khamdamov U., Abdurakhmonova N., et al. UZWORDNET: A Lexical- Semantic Database for the Uzbek Language // Proceedings of the 11th Global Wordnet Conference, Virtual, Jan. 2021. P. 8–19. https://doi.org/10.18653/v1/2021.gwc-1.2
10. Marsan B., Kara N., Ozcelik M., Arican B. N., et al. Building the Turkish FrameNet // Proceedings of the 11th Global Wordnet Conference, Virtual, Jan. 2021. P. 118–125. https://doi.org/10.18653/v1/2021.gwc-1.14
11. Rezhake M., Kuerban A. Design of the Uyghur FrameNet Desktop // Lecture Notes on Software Engineering. 2015. Vol. 3. No. 1. P. 53–56. https://doi.org/10.7763/LNSE.2015.V3.165
12. Navigli R., Ponzetto S.P. BabelNet: The Automatic Construction, Evaluation and Application of a Wide-Coverage Multilingual Semantic Network // Artificial Intelligence. 2012. Vol. 193. P. 217–250. https://doi.org/10.1016/j.artint.2012.07.001
13. Gangemi A., Alam M., Asprino L., Presutti V., et al. Framester: A Wide Coverage Linguistic Linked Data Hub. // Knowledge Engineering and Knowledge Management. EKAW 2016. Lecture Notes in Computer Science. 2016. Vol. 10024. P. 239–254. https://doi.org/10.1007/978-3-319-49004-5 16
14. Nivre J., de Marneffe M-C., Ginter F., Hajič J., et al. Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection // Proceedings of the Twelfth Language Resources and Evaluation Conference (LREC 2020), Marseille, France, May. 2020. P. 4034–4043.

This work is licensed under a Creative Commons Attribution 4.0 International License.
Presenting an article for publication in the Russian Digital Libraries Journal (RDLJ), the authors automatically give consent to grant a limited license to use the materials of the Kazan (Volga) Federal University (KFU) (of course, only if the article is accepted for publication). This means that KFU has the right to publish an article in the next issue of the journal (on the website or in printed form), as well as to reprint this article in the archives of RDLJ CDs or to include in a particular information system or database, produced by KFU.
All copyrighted materials are placed in RDLJ with the consent of the authors. In the event that any of the authors have objected to its publication of materials on this site, the material can be removed, subject to notification to the Editor in writing.
Documents published in RDLJ are protected by copyright and all rights are reserved by the authors. Authors independently monitor compliance with their rights to reproduce or translate their papers published in the journal. If the material is published in RDLJ, reprinted with permission by another publisher or translated into another language, a reference to the original publication.
By submitting an article for publication in RDLJ, authors should take into account that the publication on the Internet, on the one hand, provide unique opportunities for access to their content, but on the other hand, are a new form of information exchange in the global information society where authors and publishers is not always provided with protection against unauthorized copying or other use of materials protected by copyright.
RDLJ is copyrighted. When using materials from the log must indicate the URL: index.phtml page = elbib / rus / journal?. Any change, addition or editing of the author's text are not allowed. Copying individual fragments of articles from the journal is allowed for distribute, remix, adapt, and build upon article, even commercially, as long as they credit that article for the original creation.
Request for the right to reproduce or use any of the materials published in RDLJ should be addressed to the Editor-in-Chief A.M. Elizarov at the following address: amelizarov@gmail.com.
The publishers of RDLJ is not responsible for the view, set out in the published opinion articles.
We suggest the authors of articles downloaded from this page, sign it and send it to the journal publisher's address by e-mail scan copyright agreements on the transfer of non-exclusive rights to use the work.