Published: 05.10.2026
Full Issue
Part 1. Computer Resources for Automatic Language Processing
Automatic Speech Recognition: Factors Reducing Accuracy When Working with a Regional Variant of the Russian Language (using Material from Dagestan)
We evaluate the quality of automatic speech recognition (ASR) for Standard Russian and its Dagestani regional variety – a contact idiom formed in the multilingual environment of Dagestan. We tested ten off-the-shelf ASR configurations in a zero-shot setting (without fine-tuning) representing three architectural approaches (Conformer+RNN-T, Conformer+CTC, and Encoder-Decoder Transformer) on datasets of Standard Russian (GOLOS, RuDevices; approx. 20K and 90K utterances respectively) and Dagestani Russian (DaGRuS; approx. 40K utterances). All models demonstrate a word error rate (WER) degradation of 33%–50% when evaluated on Dagestani Russian, highlighting a severe domain shift problem. Error structure analysis reveals three primary degradation mechanisms: “Russification” (systematic replacement of Dagestani forms with phonetically similar Standard Russian words); deletion of discourse markers and out-of-vocabulary (OOV) terms such as ethnonyms and toponyms; and generative model hallucinations on incomprehensible audio segments. We found that end-to-end architectures suffer less degradation than cascaded ones (33%–35% vs. 40%–42%), yet no model achieves acceptable performance (minimum WER ≈ 43%). Finally, we formulate recommendations for designing ASR systems robust to the regional variability of the Russian language.
Pragmatic Boundaries of Greetings and Goodbyes: Evaluating LLMs on the Russian Multimedia Politeness Corpus
This study investigates the ability of large language models to detect greeting and goodbye sequences in Russian-language dialogues and determine their precise pragmatic boundaries. YandexGPT 5.1 Pro, GigaChat 2 Max, and GPT-5 are evaluated on two tasks: classifying dialogues based on the presence of these sequences and extracting individual utterances, taking their intentional boundaries into account. The experiment is conducted on 216 dialogues from the Russian Multimedia Politeness Corpus, thereby accounting for the broad communicative and social context of interaction. GPT-5 achieves the strongest dialogue-level classification performance, whereas YandexGPT and GigaChat attain near-perfect precision through conservative prediction strategies. Performance declines markedly for all models when exact boundaries must be extracted. Error analysis shows that the models systematically over-include contextual nonverbal actions and miss non-conventional markers, including gestures, address forms, discourse markers, and implicatures indicating future contact. The article presents an updated description of the corpus and its web interface, as well as facilities for constructing diagnostic subcorpora and benchmarks. The findings underscore the need to model discourse structure and the contextual variation of everyday communication more precisely.
On Improving RAG Quality for Contradiction Detection in Court Rulings
The paper extends and continues the work on a pilot system for detecting contradictions in court rulings. The tool aims to identify contradictions between court rulings on administrative offense cases and the legislative norms governing them. The previously presented pipeline, which includes rule-based document preprocessing, search over a structured repository of legal norms, and LLM fine-tuning using a LoRA adapter, demonstrated fairly high metrics [1].
However, this implementation had a number of limitations, primarily related to the RAG component. Unlike LLM, the RAG module was not trained on legal texts, which led to a substantial loss of relevant fragments for comparison at the retrieval stage.
The present work therefore focuses on a detailed quality analysis and improving the RAG component of the system. A dataset for evaluating premise retrieval recall was compiled, comprising 30 documents, along with automated tests containing all relevant sentence pairs (more than one hundred pairs in total). Using this data, the recall@n was calculated for the original (baseline) version of the model.
To improve the recall of the retrieved data, the following experiments were conducted: (1) fine-tuning the original embeddings on the system's training dataset; (2) adding a reranking model, both with and without fine-tuning; (3) combining variations of base/fine-tuned embeddings/reranker across several values of the top_k parameter (the number of retrieved matches); and (4) replacing the sentence embedding model with alternative options.
These experiments improved RAG recall from 72% to 90%. However, increasing the number of the retrieved fragments reduced the system's overall precision metrics, indicating a need for further training.
Other notable optimizations include reorganizing the structure of the legal codes database, expanding the test set, and improving the algorithm for filtering out irrelevant sentences from the rulings.
Communicative Bias in ASR Recognition of Atypical Russian Speech
The article addresses the problem of communicative bias in automatic speech recognition systems when processing atypical speech in Russian. The relevance of the study is driven by the active integration of artificial intelligence technologies into the social sphere and public administration in Russia, which presupposes ensuring their accessibility for various categories of citizens, including people with speech disorders. The authors introduce the concept of communicative bias of ASR models as a property of a system that, under conditions of uncertainty (atypical speech, noise, pauses), makes decisions about filling semantic and acoustic gaps based on its own training data rather than the speaker's actual communicative intention. The empirical basis of the study is the RuAphasiaBank corpus, which includes speech recordings from patients with various types of aphasia of varying severity, as well as neurotypical informants. For comparative analysis, six ASR models operating with Russian were selected: two open-source (GigaAM, T-One) and four commercial (Yandex SpeechKit, Any2Text, Charla, Speech2Text). Recognition quality was assessed using WER and CER metrics, as well as a qualitative analysis of the ratio of insertions, deletions, and substitutions at the word and character levels. The results show that the models implement different strategies when encountering atypical speech: “panic shortening” (GigaAM, Yandex SpeechKit), “lexical preservation” at the expense of accuracy (T-One), and a “stable” strategy of balanced errors (Any2Text, Speech2Text, Charla). Linguistic analysis reveals that each model produces a specific distorted communicative style, ranging from colloquial dialogic speech to a style associated with cognitive impairments or the use of youth slang. Based on the obtained data, potential risks to the social well-being and legal security of aphasia patients when using voice assistants in public institutions are modeled.
FROM AN ANNOTATED CORPUS TO AN ONTOLOGY: A FORMAL MODEL OF MULTI-LEVEL SEMANTIC ANNOTATION OF FAKE NEWS IN THE KAZAKH–RUSSIAN MEDIA SPACE (KazFakeCorpus / KazFakeOnto)
The paper addresses the formal representation of knowledge about disinformation in a low-resource bilingual media environment. Automatic fake news detection is traditionally formulated as a binary REAL/FAKE classification task, which does not capture the internal structure of an unreliable message: the type of fake content, the disinformation technique employed, the author’s communicative intent, and the characteristics of the source and evidence base.
Two interrelated resources are presented. The first is KazFakeCorpus, a bilingual corpus of 4,276 texts in Kazakh and Russian, balanced by class within each language; the REAL class was drawn from official materials of the Gov.kz portal and the FAKE class from their controlled transformations. The texts were annotated in Label Studio by two independent experts following a multi-level semantic scheme; Krippendorff’s Alpha for the principal levels ranged from 0.79 to 0.88. The second resource is KazFakeOnto, an ontology that recasts the annotation scheme from a flat relational representation into a formal OWL 2 DL model: 38 classes, 26 object properties, 22 data properties, 112 individuals and 1,155 asserted triples. Eleven defined classes are not asserted manually but inferred by a reasoner from property values; a property chain links a news item to its annotator through an annotation object. Consistency was verified with HermiT, and the query layer was evaluated on fifteen SPARQL and thirteen DL queries.
The ontological layer is shown to provide a reproducible specification of the annotation scheme, the automatic inference of semantic categories, and a unified formal environment for formulating corpus-linguistic questions and analysing machine-learning model errors. The limitations of the approach and directions for scaling the ontology to the full corpus and to naturally occurring disinformation are discussed.
RusHallu-RAG: a Comprehensive Approach to Hallucination Detection in Russian-Language RAG Systems
This paper introduces RusHallu-RAG, the first comprehensive benchmark for detecting hallucinations in Russian-language RAG systems. The dataset of 1,000 query-answer pairs covers general knowledge and scientific domains with controlled context relevance balancing. A fine-grained taxonomy of six hallucination types is proposed and used for annotation at both response and span levels with human experts and LLMs.
For training a hallucination detector, a balanced dataset of 3,000 examples is collected using synthetic generation methods. A detector based on YandexGPT-5-Lite-8B is trained via SFT. We evaluate 15 open-weight and proprietary models alongside the trained detector. Results show that model size does not guarantee performance: medium-sized models (20–33B) sometimes outperform larger ones. The trained detector achieves an average accuracy of 63.7%, ranking second after proprietary Gemini-2.5-Pro, with a gain of 47.2 p.p. over the baseline. The gap between proprietary and open-weight models is reduced from 20–35 to 8.4 p.p., confirming the effectiveness of the approach. The benchmark, code, and trained model are publicly released.
Automatic morphological tagging for Ossetic based on data from the Corpus of Oral Texts
In this work we present the first morphologically annotated corpus for Iron Ossetic that conforms to the Universal Dependencies schema. The corpus includes 5454 manually annotated sentences from the Iron Ossetic Corpus of Oral Texts, containing 74032 tokens. We use this corpus to train a BERT-based morphological analyzer. The analyzer achieves tag accuracy of 95.60%. Furthermore, the paper presents the results of experiments aimed at improving classification quality. It is shown that neither filtering the model’s output using a context‑free analyzer nor multi‑task approach lead to a significant improvement. Finally, the paper presents the results of testing the analyzer on out‑of‑domain data and shows that the model achieves tag accuracy of 91.45% on them.
Verb Forms in Emotion Recognition Tasks: Analysis of Data Characteristics and Classification Results
Emotion recognition in text represents an important task in natural language processing. This study investigates the influence of verb morphological form on emotion classification performance by encoder models and identifies factors that may determine morphological sensitivity in language models. Three models (ruBERT-tiny2, ruBERT-base, ruRoBERTa-large), fine-tuned on a combined corpus of four Russian-language datasets (22,852 examples), were employed in the experiment. Each model classified 5,676 morphological forms of 473 verbs from the Russian Emotion Lexicon. Results revealed systematic morphological influence: differences in classification performance between forms, as measured by F1-score, reached 20–25% across all models. To account for the observed patterns, training data characteristics were analysed. Frequency of morphological forms in the corpus was shown to partially correlate with classification performance, though it does not appear to be the sole determining factor. Analysis of emotion distribution across morphological forms revealed that in the training corpus, imperative forms occur predominantly in contexts associated with Joy. However, Sadness-labelled verbs in imperative mood are systematically classified by encoders as Anger. This discrepancy between training data characteristics and classification outcomes suggests that imperatives combined with sadness semantics are interpreted by models as expressions of aggression, indicating an interaction between grammatical form and verb semantics in determining emotion class.
Part 2. Original articles
PromptAlign: Automatic Iterative Alignment of a Large Language Model to Expert Annotations without Fine-Tuning
Large language models are increasingly used in social and psychological sciences; however, their predictions are systematically biased relative to expert annotations (alignment shift), and attempts to manually refine instructions or to supply examples lead to a “blanket-pulling” effect – degradation of minority classes.
We propose PromptAlign, an automatic prompt-optimization pipeline that requires neither access to model weights nor human involvement in iterative correction, because the taxonomy and the feature space are specified by an expert a priori. The pipeline extracts hybrid linguistic and semantic features using the large language model itself, builds interpretable L1-regularized logistic regressions to localize disagreements with the expert, performs a meta-analysis of errors, and compiles new prompts while controlling the class balance through a class-imbalance index (Blanket Pulling Index, BPI). An experiment on a corpus of 541 Russian-language texts with three psychological classes showed an increase in macro-F1 on the held-out test set from 0.655 to 0.749 and in the minimum per-class F1 from 0.449 to 0.656 (p>0.99 by a paired bootstrap test). The method significantly outperforms both the zero-shot baseline and few-shot prompting. A component-contribution analysis confirms the critical role of semantic interpretation of features and of periodic prompt rebuilding. The final prompt contains five rules that reflect latent expert criteria; unlike the “black boxes” produced by automatic optimizers, these rules are available for inspection and correction by a specialist.
Capabilities of the Developed License Support and Maintenance Information System
The purpose of creating the License Support and Maintenance Information System is to automate the management, acquisition, maintenance and use of licensed software products. The results of the system development over the last two years are presented. A request approval mechanism has been implemented, and various requests have been created on its basis: user requests for a new license, for adding new software products to the catalog of supported software, auditor requests for purchasing additional licenses. Work has been completed to fill the database, conveniently present information on licenses and various statistical information, as well as integrate with other services within the JINR Digital Ecosystem.
Requirements and Technical Solutions for a Scientific Data Management System Architecture
This paper addresses the problem of defining requirements for the architecture of a scientific data management system and selecting technical solutions for implementing its components. Scientific data are divided into three categories – research datasets, scientific and technical information (STI), and scientometric information (SMI) – which differ in structure, volume, update patterns, access profile, and metadata requirements, complicating their effective management within a single general-purpose store. Based on a data typology and use cases for the key roles (researcher, analyst, data engineer, administrator), requirements for the components are formulated. A comparative analysis of alternative solutions is performed for each component, and a modular architecture is proposed that separates sources of truth from derived representations. A consistent technology stack compatible with on-premises deployment is determined, along with cross-cutting mechanisms of identification, authorization, and resource management.
Architecture of the Linguistic Platform Turklang
The article describes the architecture of the TurkLang multifunctional linguistic platform, currently being developed at the Institute of Applied Semiotics of the Academy of Sciences of the Republic of Tatarstan. The system is based on a three-level knowledge base: linguistic universals, descriptions of units of specific languages, and empirical speech data. The knowledge base is implemented as a graph, where the primary storage unit is the morpheme – the minimal meaningful unit of language (akin to an atomic data element). This design choice is driven by the element-combinatorial nature of Turkic languages: a word form (the concrete realization of a word in text) is a linear chain consisting of a root and sequentially attached affixes – modifiers, each expressing a single grammatical category (case, number, tense, etc.) in a strictly fixed order. The paper examines levels of abstraction, types of relationships between entities, and mechanisms for integration with external sources (an electronic encyclopedia, lexicographic resources, and GIS data). Special attention is paid to the application of international standards for syntactic (Universal Dependencies), semantic (Universal Meaning Representation), and pragmatic (DiAML) annotation. A comparison with existing platforms (Sketch Engine, LingvoDoc, Apertium) is provided. To date, the platform's knowledge base covers more than 30 languages and dialects, contains over 237,000 root morphemes and about 40,000 compatibility rules, while the corpus component comprises over 180 million of word occurrences.
Predicting Changes in Personality Behavior Under the Influence of an Interlocutor
The study of interactions between users of information systems is an urgent task in humanitarian and technical research. The paper presents the development of a regression model that predicts changes in the severity of the content components of a text under the influence of other people. The data source was short conversations (message triplets) between two users on social networks. The trained model is designed to be used as part of a cognitive architecture that reflects a model of a subject's behavior in a digital environment.
Organization and Management Metadata of Scientific Articles in Digital Mathematical Libraries
This paper presents the main approaches to creating metadata management systems for digital mathematics libraries. It describes approaches to editing metadata adopted in various digital mathematics projects. Methods for representing metadata in digital libraries are presented. A metadata editor for digital mathematics collections is proposed. Approaches to adding, editing, and deleting article metadata in a digital library are described, and a method for editing metadata in a digital library is proposed.
The Tatar Language Semron Database as a Digital Linguistic Resource for Explainable Artificial Intelligence
The article considers the Tatar language semron database as a Digital linguistic resource intended for representing lexical and grammatical units of the Tatar language and for their use in explainable artificial intelligence. The basic unit of the database is the semron, understood as a potential meaning-forming unit described through its form, class, syntactic, semantic and pragmatic characteristics, valencies, corpus evidence, links to a knowledge graph and expert verification status. The article proposes a model for organizing the semron database, including the frame-based semron record, the distinction between the base frame of a semron and the specialized frame of a seiron, and the lifecycle of a database record. It is shown that such a database cannot be reduced to a dictionary, corpus, morphological analyzer or knowledge graph, since it performs an integrative function by linking the formal description of a linguistic unit with the conditions of its contextual semantic activation. The practical significance of the work lies in developing a model of a digital linguistic resource suitable for expert curation, machine processing, corpus-based validation, linkage with a knowledge graph and subsequent use in the Semiotic Generative Recognizing Model.
Automatic Correction of Individual Educational Trajectory in Adaptive Learning Systems Based on a Composite Approach
In the context of the expanding use of adaptive courses in higher education, the problem of early identification of underperforming students and automatic adjustment of their individual educational trajectory becomes increasingly relevant. Existing methods for identifying underperforming students do not provide an in‑depth analysis of the individualized learning approach employed by the student. The generalized adaptive learning algorithms currently in use do not directly address the task of identifying underperforming students, which may lead to their late detection – by which time it is no longer possible to adjust the students’ approach to learning to ensure successful course completion. The identified problem requires an immediate solution.
The aim of the study was to develop algorithms for identifying underperforming students in adaptive learning systems using a composite approach – based on analysing the individualized learning approach and enabling automatic adjustment of the individual educational trajectory to ensure successful course completion.
The developed algorithms detect student underperformance in mastering an academic discipline by predicting the outcomes of individualized learning using a domain model. In addition, algorithms for the automatic adjustment of the individual educational trajectory have been proposed. Unlike existing approaches, these algorithms generate an individualized educational strategy to ensure successful completion of the course. These algorithms can be integrated into existing adaptive learning systems to increase the proportion of students who successfully complete adaptive courses.
Multi-agency System for Digital Intelligence in Geological Knowledge
A new approach to building geographically distributed information systems based on a multi-agent architecture is proposed. This approach eliminates the rigid unification of protocols and resource cataloging in favor of dynamic integration of data sources. A key element of the system is a verification agent, which verifies the validity of scientific research results through formal control and the use of multiple LLMs. The developed prototype, implemented on the GeologyScience.ru platform, eliminates the shortcomings of the existing AI assistant related to hallucinations and unreliability of information and creates a foundation for accumulating verified knowledge.
An Approach to Assessing the Reliability of Information in Popular Science Discourse Based on the Analysis of Argumentation and Scientific Fact-Checking
This paper examines the challenges of assessing the reliability of information extracted from popular science texts. Popular science is known to be characterized by a free style of presentation, loose references to primary sources and expert opinions, and the presence of controversial statements and conclusions. Clarifying these components of the author's reasoning can significantly increase the credibility among critical readers or, conversely, refute the authors' assertions.
This paper proposes a new approach to assessing the reliability of information contained in popular science publications. This approach utilizes D. Walton's argumentation theory to identify potentially unreliable claims and logical fallacies. Searching scientific databases and calculating trust metrics for scientific publications based on their metadata and scientific ratings will enable the verification of controversial claims and an assessment of their reliability. The use of large language models will enable the generation of explanations for decisions made.
To support the development and validation of this approach, two data collection methods are proposed. The first uses existing corpora of Russian-language popular science texts and commentaries with annotated argumentation. The annotation includes, among other things, polemical argumentation schemes that challenge the author's thesis or argument. The second method utilizes LLM for generating synthetic theses and popular science articles, based on actual scientific publications.
Information and Analytical System for Assessing Cognitive Functions of Aviation Personnel Based on Fuzzy Logic
The article presents an information and analytical system for assessing eight cognitive functions of aviation personnel based on fuzzy logic. The system receives 14 input variables, which are the results of subjects passing 13 validated psychodiagnostic methods. The system operates based on the Mamdani fuzzy inference algorithm, scalable to the input space dimension (1, 2 or 3 metrics) and using adaptive bases of production rules (3, 9 or 27). The generated quantitative assessments of eight cognitive functions in the range [0; 1] are fed to a three-level classifier, which implements their qualitative interpretation by assigning them to one of three zones: risk [0;0.4], acceptable level (0.4;0.7] or high level (0.7;1]. The practical significance of the developed information-analytical system lies in its role as a tool for automated decision support for professional selection specialists. Implementing the system will enable a shift from a binary «fit/unfit» assessment to a differentiated approach to managing personnel risks that takes into account the structure of a candidate's cognitive profile.
Mirera Digital Educational Platform
This paper presents the Mirera digital educational platform designed for comprehensive support of the educational process. The platform integrates learning management tools, automated assessment, learning analytics, attendance monitoring, and AI-based support for both instructors and students. The paper describes the key functionalities of the system, including automated assessment across various disciplines, adaptive learning pathways, analysis of students' digital footprints, integrated communication tools, and AI-powered assistants. Particular attention is paid to the automation of routine educational processes and data-driven pedagogical decision-making. The results of experimental deployment demonstrate that the platform reduces instructors' workload, increases student engagement, and improves the efficiency of educational process management. The findings confirm the potential of integrated digital educational platforms as an effective tool for the digital transformation of higher education.
Methodology, Results, and Comparative Analysis of the Use of Mask-RCNN and YOLO Neural Network Models for the Task of Automatically Selecting Mode Tracks on Ionograms of Oblique Ionospheric Radiosonde
Automatic processing of ionospheric radiosonde ionograms is necessary both for operational monitoring of the ionosphere and for analyzing archived radiosonde data arrays. The central place in this task is occupied by automatic detection and classification of tracks – traces of reflected radio signal modes on ionograms. This paper presents the methodology, results, and comparative analysis of the use of neural network models (neural network architectures) for Mask RCNN and YOLO instance segmentation. The most successful models were the YOLO family, which provide both high accuracy and computational efficiency. At the same time, the exact pixel–by-pixel matching of the masks is not critical – the key requirement is the correct localization of the useful signal. The use of these models to identify tracks on ionograms has demonstrated moderate but sufficient efficiency: the main structures of the tracks are reliably identified. The most successful models were the YOLO family, which provide both high accuracy and computational efficiency. In the future, the set of models used will be expanded, primarily due to visual transformers.
Processing of Scanned Mathematical PDF Documents in the Semantic Formula Search Service
The article addresses the problem of extracting semantic relations between mathematical formulas and concepts from unstructured PDF documents, including scanned ones. A complete processing pipeline is proposed that combines optical character recognition, paragraph structure recovery, formula and variable extraction, and their linking to concepts using syntactic patterns of the Russian language. Formal models are developed for two types of relations: "variable-concept" and "variable-main formula". Experimental evaluation on a collection of 150 mathematical articles shows a significant improvement in linking accuracy, with an increase of 18 percentage points in F-measure compared to the previous version. Architectural improvements of the web service prototype are described, providing asynchronous processing and scalability.
Method for Semantic Analysis of Software Requirements Changes in Technical Documentation Versions
A formalized method for the automated detection of software requirement changes during the evolution of technical documentation versions is described. The input data consist of sets of requirements extracted and classified into functional and non-functional categories, with further granularity regarding non-functional requirement types (e.g., performance, security, availability, etc.) for each documentation version. The task of matching requirements between versions is reduced to constructing a partial mapping between the requirement sets and assigning one of the following statuses to each requirement: Unchanged, Modified, New, or Deleted.
A composite similarity function for requirement pairs is proposed, integrating semantic similarity based on vector representations (embeddings), lexical similarity, requirement class consistency, and the structural context of the document. The results of applying the proposed method to an open dataset of requirement pairs are presented. A prototype system for requirement change analysis in technical documentation has been developed.
Methodology and First Results of Ionogram Synthesis by Physics-Informed Neural Networks
Abstract. The problem of synthesizing ionograms of vertical ionosphere radiosonde using physically informed neural networks is solved. In the proposed approach, the neural network is used not for direct image generation, but for approximating the high-altitude profile of the electron concentration. Network training is performed taking into account a stationary one-dimensional balance equation, including an ionization source, recombination losses, and effective transfer of charged particles. The loss function additionally includes boundary conditions, a restriction on the smoothness of the profile, and diagnostic parameters of the F2 layer. The resulting electron concentration profile is transformed into a height distribution of the plasma frequency. Therefore, without taking into account the Earth's magnetic field and collisions of electrons with other particles, the reflection points, the group refractive index and the dependence of the effective height on the sounding frequency are calculated. Based on this dependence, a synthetic image of an ionogram with background noise and vertical interference structures is formed. The results confirm the possibility of constructing a single physically consistent chain from the height profile of the electron concentration to the image of the ionogram. The developed approach can be used to generate controlled synthetic data, study the effect of ionospheric parameters on the shape of ionogram trails, and subsequently develop algorithms for automatic processing of radiosonde results.
Cognitive Modeling of User Interaction in Multi‑Role Web Applications with 3D Functionality
The article investigates cognitive load in users of a multi-role 3D platform. An experiment (71 participants: 34 designers, 37 customers; 639 observations) using objective metrics and subjective scales revealed different load profiles: designers are slowed down by data entry forms, while customers are hindered by 3D content. Based on 34 problem areas, seven interface design principles and a task classification by cognitive complexity level, accounting for role-specific factors, were formulated. Retesting (21 participants) confirmed their effectiveness: task completion time decreased by 21–28%, errors by 33–44%, frustration by 25–35%, and the use of 3D features increased from 40.5% to 90.9%. The developed three-component model (taxonomic, diagnostic, and design levels) is recommended for designing similar multi-role platforms with 3D functionality.