• Main Navigation
  • Main Content
  • Sidebar

Russian Digital Libraries Journal

  • Home
  • About
    • About the Journal
    • Aims and Scopes
    • Themes
    • Editor-in-Chief
    • Editorial Team
    • Submissions
    • Open Access Statement
    • Privacy Statement
    • Contact
  • Current
  • Archives
  • Register
  • Login
  • Search
Published since 1998
ISSN 1562-5419
16+
Language
  • Русский
  • English

Search

Advanced filters

Search Results

Analysis of the Effectiveness of Subword Tokenizers in a Low-Resource Linguistic Environment: Implementation Experience for the Tajik Language

Mullosharaf Kurbonovich Arabov, Svetlana Sergeevna Khaybullina
546-564
Abstract:

This paper examines modern approaches to subword tokenization of texts as applied to the low-resource Tajik language, which is characterized by a complex morphological structure and a high degree of word-form variability. In the course of the study, a large-scale heterogeneous corpus was compiled and preprocessed, comprising 99 books and 134,497 textual articles of various genres and topics, with a total volume exceeding 33 million tokens. The corpus was cleaned of noise, normalized, and used as a basis for training and subsequent testing of subword models.


Based on this corpus, five tokenization models implementing the BPE, WordPiece, and Unigram algorithms were trained and analyzed using the Hugging Face Tokenizers and SentencePiece libraries. Comparative evaluation was conducted using a set of key metrics, including the proportion of out-of-vocabulary (OOV) words, the degree of text representation compression, tokenization speed, as well as characteristics of n-gram distribution, which make it possible to assess the ability of the models to capture the morphological and structural organization of the language. The experimental results made it possible to identify the strengths and weaknesses of different approaches to subword segmentation and to determine the most effective tokenization strategies under conditions of the morphological complexity of the Tajik language. The findings obtained can be used in the development of language models and applied NLP tools for Tajik and other low-resource languages, contributing to the expansion of their presence in the digital environment.

Keywords: Tajik language, subword tokenization, low-resource languages, BPE, WordPiece, Unigram, Hugging Face Tokenizers, SentencePiece, corpus linguistics, natural language processing (NLP).

Recommender system of text analytics of legal documents

Денис Сергеевич Зуев, Марат Фаритович Насрутдинов, Айрат Фаридович Хасьянов
435-449
Abstract:

The paper discusses the use of machine learning mechanisms, natural language analysis and intellectual search in the field of jurisprudence. The main expected results are the methodology for applying text-based analytics and semantic natural language processing (NLP) algorithms in knowledge management cases in different types of legal practice. The obtained results can be applied in the field of education and knowledge management in a wider context, since the study lies at the union of jurisprudence, mathematical and computer linguistics.

We describe a prototype of a multi-agent system of intellectual analysis of legal texts that is capable of identifying general dependencies on the existing database of legal documents, providing legal cases with similar topics, recommending the most likely outcomes of judicial review.
Keywords: data analytics and data mining, data intensive domains, digital libraries, clustering, classification of judicial acts, recommender system, micro-service architecture.

Large Language Models for Word-In-Context

Denis Vladislavovich Kokosinskii
1133-1154
Abstract:

Task-specific models have long dominated natural language processing tasks. However, general-purpose large language models (LLMs) have recently begun to successfully compete with highly specialized solutions across various NLP domains. In this paper, we investigate the applicability of LLMs to the task of estimating the semantic proximity of a word's meanings in a pair of usages, known as Word-in-Context (WiC). Drawing on the multilingual CoMeDi benchmark, we propose novel approaches to building automated WiC-systems based on LLMs. We conduct a systematic comparison of five different configurations in terms of quality and computational costs. In particular, we propose a configuration where LLM predictions are adjusted using a training set, without the need to fine-tune the LLM itself. Results on test sets across seven languages show that our proposed approaches enable LLMs to outperform all existing specialized systems, establishing a new state-of-the-art (SOTA) on the CoMeDi benchmark. Nevertheless, the achieved high quality comes with a significant increase in computational costs: LLM-based systems require several orders of magnitude more computations compared to compact specialized models (such as XL-DURel). This work represents a step towards understanding the trade-off between accuracy and resource efficiency when using modern LLMs in lexical semantics tasks.

Keywords: Word-in-Context, Large Language Model, Natural Language Processing.

Linguistic Ontology for Natural Science and Technology OENT: Structure, Concepts, Relations

Б.В. Добров, Н.В. Лукашевич
Abstract: We present the large-scale linguistic ontology OENT aimed for the using in NLP for the solving of information retrieval tasks such as categorization and query expansion in the domain of natural science and technology. OENT is based on three approaches - traditions of librarian information retrieval thesauri, formal ontologies, WordNet-like resources.Also we describe several additional rules for concepts, text entries and relations. OENT includes 50 thousands concepts, more than 150 thousands text entries, 200 thousands direct relations and more than two millions inherited relations.
Keywords: онтология, лингвистическая онтология, Онтология по естественным наукам и технологиям ОЕНТ, структурные особенности ОЕНТ.
1 - 4 of 4 items
Information
  • For Readers
  • For Authors
  • For Librarians
Make a Submission
Current Issue
  • Atom logo
  • RSS2 logo
  • RSS1 logo

Russian Digital Libraries Journal

ISSN 1562-5419

Information

  • About the Journal
  • Aims and Scopes
  • Themes
  • Author Guidelines
  • Submissions
  • Privacy Statement
  • Contact
  • eLIBRARY.RU
  • dblp computer science bibliography

Send a manuscript

Authors need to register with the journal prior to submitting or, if already registered, can simply log in and begin the five-step process.

Make a Submission
About this Publishing System

© 2015-2026 Kazan Federal University; Institute of the Information Society