• Main Navigation
  • Main Content
  • Sidebar

Russian Digital Libraries Journal

  • Home
  • About
    • About the Journal
    • Aims and Scopes
    • Themes
    • Editor-in-Chief
    • Editorial Team
    • Submissions
    • Open Access Statement
    • Privacy Statement
    • Contact
  • Current
  • Archives
  • Register
  • Login
  • Search
Published since 1998
ISSN 1562-5419
16+
Language
  • Русский
  • English

Search

Advanced filters

Search Results

The SAO RAS archive system. Maintenance and upgrading

О.П. Желенкова, В.В. Витковский, Т.А. Пляскина
Abstract: The observatory archive system includes a digital data storage and a search information system (SIS) with a dynamic web-based interface and http data access. To date, the system includes 16 digital collections of observational data (local archives), obtained on different instruments, operating or working on the telescopes. The earliest data refer to the end of 1994. There are currently actively replenished 6 local archives. The data storage includes a temporary storage area, located on a file server of the 6-m telescope (BTA), and the area of permanent storage. Permanent storage area includes CD / DVD-discs, a hard disk of the dedicated archive server and also a large capacity flash disk. For data protection during emergency situations or I/O defects of disks, we provide two full copies of the CD / DVD disks and two copies of the archive data on the hard disk of the archive server. One copy (A0) repeats optical discs, the other (A1), with the slightly modified directory structure, actually used by the SIS. D igital media devices and read-write data drives can not be attributed to long-term storage devices. For long-term storage of digital data it is necessary to provide rewriting of information every 5-10 years on a new type of media. Archive copies A0 and A1 are also supported for this procedure of rewriting. Archival data (A1) is repeated on a flash disk with the addition of a dump of tables and programs. There is the system backup for restoring after an emergency on the server. To ensure the modernization of the SIS, we support the two schemes of the database - test and operational. All our developments take place in the test database schema. When modifying the scheme after its checking the SIS switched to an updated version of the database. The original copy of the A0 and the availability of test database scheme allow to modernize the SIS, even at the level of the tables. Currently, the SIS implemented on DBMS PostgreSQL 8.3.7.
Keywords: цифровые коллекции экспериментальных данных, веб-доступ к архивам наблюдений, виртуальная обсерватория, предметно-ориентированные базы данных.

Adaptive RAG-Architecture for Intelligent Search in the Corpus of Educational Institutions Documents

Anna Dmitrievna Budrevich, Mikhail Mikhailovich Abramskiy, Iskander Airatovich Valishin
1338-1360
Abstract:

The problem of improving the quality of intelligent search in a corpus of educational institution documents, including curricula, course syllabus, and regulatory documents, was solved. Classic architectures of the Retrieval-Augmented Generation (RAG) approach, based on a single plaintext search module, demonstrate low accuracy when used with documents containing tables, logical relationships between entities, and strict regulatory document wording. An adaptive RAG architecture consisting of four layers is proposed, each taking into account the specifics of data storage in such documents. The results show that considering document structure and adaptive query routing significantly improve the actual accuracy of responses. The proposed architecture can be used in the design of intelligent assistants for administrative and educational services at higher education institutions.

Keywords: RAG, RAG architecture, educational institution documents, intelligent search, artificial intelligence, document processing, semantic models.

Fuzzy-Logic Adaptation of Sliding Window Parameters in Data Preparation for Large Language Models

Maxim Vladimirovich Bobyr, Natalya Anatolyevna Milostnaya, Svetlana Yurievna Belskaia
1318-1337
Abstract:

The article proposes a fuzzy regulator for calculating sliding window parameters to prepare training data for Large Language Models. The traditional approach sets the stride and context length parameters as fixed constants, uniform for the entire text, and does not account for the linguistic characteristics of individual fragments, such as dense scientific text and monotonous, repetitive text. The proposed method utilizes two automatically computed features of a fragment: lexical diversity and the average BPE token length. Based on the Mamdani algorithm with a base of 9 fuzzy logic rules and defuzzification using the center of gravity method, the fuzzy regulator adaptively calculates the stride and context length values for each fragment. The proposed approach has a cognitive interpretation, as it mimics the mechanism of adaptive human attention during reading, for example, complex fragments are processed more attentively with a small stride size.

Keywords: fuzzy inference, Mamdani algorithm, sliding window, LLM (Large Language Model), type-token ratio, BPE tokenization, cognitive modeling, adaptive text processing.

Strong and Weak Relations in the Academic Web

Andrey Anatolievich Pechnikov
526-542
Abstract: The web graph is the most popular model of real Web fragments used in Web science. The study of communities in the web graph contributes to a better understanding of the organization of the fragment of the Web and the processes occurring in it. It is proposed to allocate a communication graph in a web graph containing only those vertices (and arcs between them) that have counter arcs, and in it to investigate the problem of splitting into communities. By analogy with social studies, connections realized through edges in a communication graph are proposed to be called "strong" and all others "weak". Thematic communities with meaningful interpretations are built on strong connections. At the same time, weak links facilitate communication between sites that do not have common features in the field of activity, geography, subordination, etc., and basically preserve the coherence of the fragments of the Web even in the absence of strong links. Experiments conducted for a fragment of the scientific and educational Web of Russia show the possibility of meaningful interpretation of the results and the prospects of such an approach.
Keywords: web graph, communication graph, community in graph, strength of the linkages.

Artificial Intelligence in Several Fragments

Yuri Evgenievich Polak
442-485
Abstract:

This paper is a mosaic of vivid fragments describing the industrial aspects of artificial intelligence (AI). These are sketches of the overall picture, which will likely never be completed, as each day brings information about new achievements, ideas, and threats. Discussions cover issues of civilian AI in short-term workstations, the development of algorithms for intelligent games, the threats and dangers posed by AI, AI ethics, and standards and international norms for artificial intelligence. Each fragment is a review of the latest (mid-January 2026) Russian and international sources, including quotes, translations, screenshots, and links to original documents.


This text remains an immense "fragment" on the benefits of AI applications, which was presented with the greatest speed. Perhaps this will be the beginning of a separate, never-ending study.

Keywords: artificial intelligence, Dartmouth Seminar, predecessors of AI, development of intellectual game algorithms, threat and danger, this AI, regulation of artificial intelligence.

Database and astronomy - a practical approach

О.С. Бартунов, С.В. Карпов
Abstract: Modern astronomy experiences a rapid increase of data being acquired in obser-vations or in numerical simulations. Efficient storage and management of these data is a problem as important as the scientific analysis itself. In the article we analyze the reasons behind such an "information explosion" and outline the problems the Virtual Observatory faces due to it – primarily, the development of standards and technologies for accessing the data remotely and in an automated way. We also review the main requirements for modern scientific information such as repeatability of results, versioning, data provenance etc. Obvious solution for long-term storage of such a data is Database Management Systems (DBMS). We discuss how well different kinds of astronomical information – catalogues, spectra, images, time series, simulation results – are compatible to relational model of most popular DBMSs and formulate requirements for specia-lized systems which are optimal for storage and analysis of scientific data.
Keywords: виртуальная обсерватория, научные данные, системы хранения научной информации, СУБД.

Semantic Fragment of a Research e-Infrastructure: necessary information objects, tools and services

С.И. Паринов
Abstract: Basically semantic linkage technique is used to specify in computer readable form known facts or relationships that definitely exist between information objects like people, organizations, research results, etc. Recently developers also started using it to visualize scientists’ opinions or scientific hypotheses (e.g. inference/deduction, impact/usage, theoretical hierarchy, etc.) about relationships between research objects. Based on CERIF Link entity and the Semantic Layer and assuming that scientists typically re-use research objects by making relation-ships between them we propose a sketch of a research e-infrastructure semantic fragment, which allow scientists unlimited re-use of research information systems (RIS) content. After some development a semantic linkage technique provides scientists with tools and services for semantic linking of any pair of research objects, which metadata are available within content of any RIS. This application also allows scientists a decentralized development of semantic vocabularies that guarantee a covering by this technique any new types of relationships. In the paper we discuss information objects, tools and services which are necessary for proper functioning of proposed fragment of the research e-infrastructure. We also discuss a “quality control” topic which in this context is very important.
Keywords: ESRI ArcGIS for INSPIRE, Estonia, Spatial Data Infrastructure, automatic data update mechanisms.

Vit Quantization: CPU-Centric Analysis of the Trade-Off between Size and Speed

Amir Ramisovich Nigmatullin, Rustam Arifovich Lukmanov, Ahmad Taha
262-286
Abstract:

Using Vision Transformer (ViT) models in real medical practice – for example, in hospitals or diagnostic centers – is often difficult because doctors' work computers usually do not have powerful graphics processors (GPUs), and computing resources are limited. This work investigates a complete practical pipeline for model inference, aimed at reducing computational costs without significant loss of predictive performance.


The proposed approach combines several optimization techniques. First, knowledge distillation (KD) is used, where a compact student model learns to mimic the behavior of a larger, more accurate teacher model. Second, Exponential Moving Average (EMA) of the model weights is applied to stabilize training and improve generalization. Third, post-training INT8 quantization (PTQ) is explored to reduce model size and accelerate inference. Additionally, a simplified quantization-aware training variant (QAT-lite) is considered, where the effects of quantization are partially incorporated during fine-tuning.


Experiments are conducted on the ISIC dataset, which contains dermoscopic images of skin lesions. Model performance is evaluated using standard classification metrics, including accuracy, macro-averaged F1 score, and area under the ROC curve (ROC-AUC). CPU performance is also analyzed, including inference latency, throughput, memory consumption, and the final model size.


The results show that post-training INT8 quantization preserves performance close to the FP32 baseline while substantially reducing memory and computational requirements. In contrast, QAT-lite does not consistently provide reproducible improvements over PTQ.

Keywords: Vision Transformer, knowledge distillation, EMA, post-training quantization, quantization-aware training.

Automatic Extraction of Argumentative Relations from Scientific Communication Texts

Yury Alekseevich Zagorulko, Elena Anatolievna Sidorova, Irina Ravilevna Akhmadeeva
1070-1084
Abstract:

The complexity of the problem of extracting argumentative structures is associated with such problems as selecting argumentative segments, predicting long-range connections between non-contact segments, and training on data labeled with a low degree of inter-annotator consistency. In this paper, we consider an approach to extracting argumentative relations from fairly large texts related to scientific communication. A comparative analysis was performed of fine-tuning methods using a pre-trained Longformer-type language model that takes into account long contexts and two methods that take into account annotator discrepancies in argument labeling by using the so-called soft labels obtained by uniformly smoothing labels and averaging expert assessments. The experiments were conducted on four datasets containing positive and negative examples of statement pairs (premise, conclusion) and differing in segmentation methods and average text size. The best results were obtained using the model with averaging expert assessments. At the same time, it is noted that the model using smoothed labels also increases the accuracy of classifiers, but worsens the recall.

Keywords: argument mining, argumentative relation extraction, scientific communication, segmentation problem, soft label, label smoothing, language model.

A Recommendation System for Finding Semantically Similar Fragments of Program Code

Vitaly Ivanovich Zorin, Evgeny Konstantinovich Lipachev
751-781
Abstract:

Recommendation systems in the scientific information space serve as essential tools for search and navigation when working with scientific documents. Software code is currently considered as an object of scientific knowledge and, as a result, an important task is to create software lifecycle support systems, in particular, to find similar software solutions, detect code borrowings, analyze and evaluate code quality.


This paper proposes a content-based recommender system that provides users with a personalized list of code fragments that are functionally equivalent to the input query code presented in one of the programming languages from the established set.


The basic algorithm of the system is based on the representation of the program code in the form of an abstract syntax tree followed by the construction of a vector space of program codes. The semantic similarity of program codes is determined by the distance between code vectors in a multidimensional space.


The personalization of recommendations is achieved through a filtering module that ranks the retrieved fragments taking into account the user's profile. The factors under consideration are the language preferences of the user and his areas of scientific interests, extracted through integration with ORCID.


To ensure the system's operation, a specialized dataset was created based on the CodeNet corpus. The problem of automated language detection from a snippet of the presented code in one of the 19 languages included in the current rating list of programming languages has also been solved.

Keywords: abstract syntax tree, code embedding, content-based filtering, cross-language clone, cross language code search, code similarity, recommender system.

Comparison of Approaches to the Problem of Automatic Generation of Official Reply Letters using LLM

Ivan Evgenievich Nikolaev, Andrey Vitalievich Melnikov, Kirill Evgenievich Alekseev, Alexander Sergeevich Belonogov, Mikhail Aleksandrovich Rusanov
1174-1188
Abstract:

One of the key automation challenges for government agencies is the preparation of official response letters. This article presents an empirical comparison of two approaches to the automatic generation of official response letters: one based on templates defining the letter structure and one based on relevant letter examples selected using RAG. The original GovLetter dataset, which includes real-life business correspondence from government agencies in the Khanty-Mansi Autonomous Okrug – Yugra, was used as a basis. Generation was performed using a locally deployed open-source language model. The quality of the results was assessed according to 12 criteria using the Schema-Guided Reasoning (SGR) framework and the LLM-as-a-Judge methodology. The experimental results demonstrate that the example-based approach outperforms the template-based approach in most metrics, particularly in the accuracy of arguments, formal tone, and correct formatting. The obtained results confirm the potential of solutions based on document history for the effective automation of official response letter preparation.

Keywords: text generation, official correspondence, large language models, RAG, quality assessment, SGR.

Development of an Adaptive System for Generating Game Quests and Dialogues Based on Large Language Models

Vsevolod Tarasovich Trofimchuk, Vlada Vladimirovna Kugurakova
953-993
Abstract:

This article addresses the problem of creating dynamic narrative systems for video games with real-time interactivity. It presents the development and testing of a GPT integration component for dialogue generation, which revealed a critical limitation of cloud-based solutions – a 30-second latency unacceptable for gameplay. A hybrid architecture of an adaptive system is proposed, combining LLMs with reinforcement learning mechanisms. Particular attention is given to solving the problems of game world consistency and managing long-term context of NPC interactions through a RAG approach. The transition to the Edge AI paradigm with the application of quantization methods to achieve a target latency of 200–500 ms is substantiated. Metrics for evaluating personalization and dynamic content adaptation have been developed.

Keywords: video games, large language models, LLM, dialogue generation, quest generation, adaptive quests, procedural content generation, agent behavior, game AI, machine learning in games.

CALCULATION OF ROD ELEMENTS WITH CRACKS BASED ON A COMBINA-TION OF ROD THEORY AND ELASTICITY THEORY

Murat Nurievich Serazutdinov
330-350
Abstract:

Mathematical models for calculating the stress-strain state of rods with cracks under tension-compression and bending deformations are presented. A combination of the relations of the theory of elasticity and the theory of rods is used. The main provisions of the proposed modeling method are based on dividing the rod into fragments and finding deformations and stresses for each of the selected fragments according to the theory of rods or the theory of elasticity. Calculation algorithms are described, which are relatively simple to implement. Numerical data for solving problems are provided to illustrate the reliability and accuracy of calculations based on the models described in the article.

Keywords: rod, beam, crack, stress-strain state, Mohr's formula, theory of elasticity, theory of rods.

Taking into Account the Structure of the Document in the Method of Automatic Annotation of Mathematical Concepts in Educational Texts

Konstantin Sergeevich Nikolaev
558-577
Abstract:

The enrichment of educational texts with semantic content (in particular, adding hyperlinks to the pages of the service that displays detailed information about concepts in the text) helps to increase the efficiency of students' assimilation of the material. The existing methods of semantic markup of educational texts do not take into account the structural features of such documents, which leads to excessive recognition of concepts. This article describes the development of the method of automatic annotation of mathematical concepts in educational mathematical texts by adding functionality to account for the structure of an educational document. The main purpose of the method is to process educational materials of the distance education course "Technology for solving planimetric problems". Following a single template when creating course pages allows you to apply an analysis of the web page markup and keywords used by the course creators. The main task in this process is to determine the type of table cell containing text fragments of educational materials. In accordance with the recommendations of the course creators, definitions should be highlighted in the cells containing the task statement, as well as in those blocks where the input data of the task is indicated. The type of table cells is determined by analyzing their attributes and searching for keywords in their contents. This limitation of recognizable text fragments will improve the student's perception of the course pages and improve the quality of learning.

Keywords: semantic analysis, mathematical ontology, didactic relations, mathematical education, document markup.

Leveraging Semantic Markups for Incorporating External Resources Data to the Content of a Web Page

Evgeny L’vovich Kitaev, Rimma Yuryevna Skornyakova
494-513
Abstract: The semantic markups of the World Wide Web have accumulated a large amount of data and their number continues to grow. However, the potential of these data is, in our opinion, not fully utilized. The semantic markups contents are widely used by search systems, partly by social networks, but the usual approach to using that data by application developers is based on converting data to RDF standard and executing SPARQL queries, which requires good knowledge of this language and programming skills. In this paper, we propose to leverage the semantic markups available on the Web to automatically incorporate their contents to the content of other web pages. We also present a software tool for implementing such incorporation that does not require a web page developer to have knowledge of any programming languages ​​other than HTML and CSS. The developed tool does not require installation, the work is performed by JavaScript plugins. Currently, the tool supports semantic data contained in the popular types of semantic markups “microdata” and JSON-LD, in the tags of HTML documents and the properties of Word and PDF documents.
Keywords: semantic web, semantic technologies, semantic markup, microdata, JSON-LD, web development, web technologies.

Construction of a “Neuron Activity–Class” Mapping in a Convolutional Spiking Neural Network with STDP Training

Alexander Sergeevich Toschev
1399-1417
Abstract:

This paper investigates the problem of constructing a mapping between the spiking activity of a convolutional spiking neural network and the class of an input image. This problem arises after unsupervised training: the convolutional layer forms an event-based representation of the image, but the network itself does not contain a ready-made mechanism for assigning a class label. The aim of the study is to evaluate how informative the spike counters of the convolutional layer are after training with the STDP rule (Spike-Timing-Dependent Plasticity, a biologically plausible learning rule for impulse-based, or spiking, neural networks) and to assess the performance of a linear readout classifier as the number of training presentations increases.


The experiments were conducted on the MNIST dataset, a standard dataset of handwritten digit images. A single-layer architecture was used: one convolutional layer with 32 feature maps, a 5 × 5 kernel, and an output dimensionality of 32 × 24 × 24, corresponding to 18432 LIF neurons, where LIF denotes the leaky integrate-and-fire neuron model. The input images were encoded using a deterministic Poisson encoder; the Poisson encoding multiplier was set to 0.006. The presentation duration, that is, the time during which an object was shown to the network, was 100 time steps. Training was performed from scratch using the STDP rule, without resuming from a previously saved state. For evaluation, a protocol was used in which spike counters were collected on 10000 training images and 10000 test images.


In the main experiment, with 15000 STDP training presentations, an accuracy of 0.8883 was achieved. The baseline label-map approach, close to the method of Diehl and Cook, produced an accuracy of about 0.48 under the same configuration. The coverage of the calibrated label map was 219 out of 18,432 neurons, while the average activity was 4,146.9416 spikes per image. The neighboring scale point with 10,000 training presentations yielded an accuracy of 0.8876. The results show that the chosen method for reading out activity has a substantial effect on the final classification quality: a linear readout classifier based on the full vector of network spike counters extracts distributed information that is lost when labels are assigned directly to individual neurons.

Keywords: spiking neural network, convolutional spiking neural network, readout, spike counters, image classification .

MetaHuman Synthetic Dataset for Optimizing 3D Model Skinning

Rim Radikovich Gazizov, Makar Dmitrievich Belov
244-279
Abstract:

In this study, we present a method for creating a synthetic dataset using the MetaHuman framework to optimize the skinning of 3D models. The research focuses on improving the quality of skeletal deformation (skinning) by leveraging a diverse array of high-fidelity virtual human models. Using MetaHuman, we generated an extensive dataset comprising dozens of virtual characters with varied anthropometric fea-tures and precisely defined skinning weight parameters. This data was used to train an algorithm that optimizes the distribution of skinning weights between bones and the character mesh.


The proposed approach automates the weight rigging process, significantly reducing manual effort for riggers and increasing the accuracy of deformations during animation. Experimental results show that leveraging synthetic data reduces skinning errors and produces smoother character movements compared to traditional methods. The outcomes have direct applications in the video game, animation, virtual reality, and simulation industries, where rapid and high-quality rigging of numerous characters is required. The method can be integrated into existing graphics engines and development pipelines (such as Unreal Engine or Unity) as a plugin or tool, facilitating the adoption of this technology in practical projects.

Keywords: synthetic dataset, Metahuman, neural networks, 3D model skinning, computer animation, machine learning.

Using FSM-Based Strategies for Deriving Tests with Guaranteed Fault Coverage for Input/Output Automata

Igor Borisovich Burdonov, Nina Vladimirovna Yevtushenko, Alexander Sergeevich Kossachev
18-34
Abstract:

In this paper, we study the possibility of using Finite State Machine (FSM-) based methods for deriving finite test suites with guaranteed fault coverage for Input / Output automata. A method for deriving an FSM for a given automaton is proposed and it is shown that finite test suites derived for such an FSM are complete for two fault models based on Input/Output automata if they are applied within the framework of proper timeouts.

Keywords: Input/Output automaton, Finite State machine, fault model, complete test suite.

Technology trends handling of big data and tools storage of multiformat data and analytics

Марат Рамилевич Биктимиров, Александр Михайлович Елизаров, Андрей Юрьевич Щербаков
390-407
Abstract: This article analyzes the development trends of processing Big Data tools and multi-format data storage and analysis. This analysis was carried out as part of our program of basic research of the Department of Mathematical Sciences, Russian Academy of Sciences “Algebraic and Combinatorial Methods of Mathematical Cybernetics and information systems of the new generation", as well as RFBR grant number 14-07-00783 “Way to store and process a large volume of scientific and reference data modern hardware platforms”.
Keywords: Big Data, storage systems, analysis, information, software, grid computing, cloud computing.

A Digital Platform for Integration and Analysis of Geophysical Monitoring Data from the Baikal Natural Zone

Andrey Pavlovich Grigoryuk, Lyudmila Petrovna Braginskaya, Igor Konstantinovich Seminskiy, Konstantin Zhanovich Seminskiy, Valeriy Viktorovich Kovalevskiy
303-316
Abstract:

This paper presents a digital platform for complex monitoring data of dangerous geodynamic, engineering-geological and hydrogeological processes occurring in the region of intensive nature management of the central ecological zone of the Baikal natural territory (CEZ BNT). The platform is intended for integration and analysis of data coming from several polygons located within the CEZ BPT in order to assess the state of the geological environment and forecasting of hazardous processes manifestation.


The platform is built on a client-server architecture. Storage, processing and analysis of data is carried out on the server, which users can access via the Internet using a web browser. Several data filtering methods (linear frequency, Savitsky-Goley and others), various methods of spectral and wavelet analysis, multifractal and entropy analysis, and spatial data analysis are currently available. The digital platform has been tested on real data.

Keywords: geophysical monitoring, digital platform, precursors, seismic forecast, earthquakes.

An Algorithmic Framework for Accurately Extracting Main Content from News Websites

Hamza Salem, Alexander Sergeevich Toschev
931-942
Abstract:

A new precise MCE algorithm for extracting the main content from news websites is presented. The proposed algorithm uses analysis of the Document Object Model (DOM) structure and content density metrics to identify and extract the informational core of a web page. The implemented approach combines three key features: the maximum number of direct child elements containing text, the maximum textual content without child elements containing text, and the closest position to the average node depth. The algorithm demonstrated superior performance compared to existing solutions such as Boilerpipe and Readability, achieving 99.96% precision, 99.69% recall, and 99.80% F1-score on a comprehensive dataset of 500 diverse web pages. Its language-independent design makes the algorithm particularly effective for extracting multilingual content, including languages with complex structures such as Arabic.

Keywords: NLP, Data Extraction, Language-Independent Algorithm, RAG (Retrieval-Augmented Generation).

Evolution of Visualization Methods for Research Publication Collections

Alexander Ivanovich Legalov, Igor Alexandrovich Legalov, Ivan Vasilievich Matkovsky
788-807
Abstract: It is proposed to add a static system of types to the dataflow functional model of parallel computing and the dataflow functional parallel programming language developed on its basis. The use of static typing increases the possibility of transforming dataflow functional parallel programs into programs running on modern parallel computing systems. Language constructions are proposed. Their syntax and semantics are described. It is noted that the need to use the single assignment principle in the formation of data storages of a particular type. The features of instrumental support of the proposed approach are considered.
Keywords: visualization of document collections, text analysis, text and metadata visualization algorithms, LDA, NMF, word2vec.

Analysis of Software System Optimization using the Example of Free Automated Library and Information Systems

Oleg Ivanovich Vasyliev, Valentin Yurevich Medvedev
151-163
Abstract:

This article is devoted to the study of the possibilities of optimizing the operability and improving the efficiency of complex multifunctional software systems using the example of free automated library and information systems (hereinafter - ALIS).


By 2023, the world has accumulated valuable experience in the creation and operation of integrated ALIS of various scales and purposes, but the issues of improving their design solutions remain relevant. First of all, this concerns the need to optimize the structure of the source code in order to increase its readability and maintainability, reduce the execution time of individual functional modules, and reduce the amount of RAM used.


As part of the study, a comparative analysis of the source codes of several existing open source databases implemented in various programming languages was carried out. The main approaches to the design of the code structure were studied, the most frequently used algorithms and patterns were identified. To assess the degree of optimization of the source code, a set of indicators was developed, including an assessment of the structure, readability, modularity and other characteristics. On this basis, individual code fragments were compared before and after the use of well-known refactoring techniques.


As a result of the work carried out, it was possible to identify the most common errors and shortcomings in the structuring of the source codes of the ALIS, to determine the main directions of their optimization. Data has been obtained on the possible reduction of testing and technical support costs by improving the quality of source codes.

Keywords: software code correction, software system optimization, refactoring, multilingual system, software system quality assessment, automated library and information systems, software development process.

Common Digital Space of Scientific Knowlege Content Fragments Modeling

Svetlana Aleksandrovna Vlasova, Nikolay Evgenievich Kalenov, Alexander Nikolaevich Sotnikov
353-368
Abstract:

The article reflects the new results of research related to the formation of the Common Digital Space of Scientific Knowledge (CDSSK). This work has been carried out since 2019 in a number of academic organizations, including the Interdepartmental Supercomputer Center of the Russian Academy of Sciences (now the Department of Supercomputer Systems and Parallel Computing at the National Research Center "Kurchatov Institute"). As part of these studies, the structure of the CDSSK ontology, a language for its description, and a number of unified software tools have been developed to ensure the formation of the ontology of individual subspaces and the input of various types and kinds of object attributes and named relationships into the CDSSK. Currently, the formation of the CDSSK content is being modeled using the example of a universal and a number of thematic subspaces. The results of this modeling are presented below. The attributes and relationships of the "Administrative Units" class objects belonging to the "Geography" subspace, the "Organizations and Their Subdivisions" class, and the "Classification Systems" class belonging to the universal subspace are presented. The ability to navigate through the loaded real resources is demonstrated.

Keywords: scientific knowledge space, ontology, named relationships, data loading, Russian administrative units, city coats of arms.

Further Development of Studies of Pressure Fields in the Arctic Region of Russia

Natalia Pavlovna Tuchkova, Konstantin Pavlovich Belyaev, Gury Mickailovich Mickailov, Alexey Nikolaevich Salnikov
1217-1232
Abstract:

The results of studies of atmospheric pressure in the Arctic region of Russia in the period from 1948 to 2008 are presented. The analysis of the climatic seasonal variation of the atmospheric pressure fields is carried out. As the main research method, the probabilistic and statistical analysis of the time series of the pressure field 60 years long at fixed points in the region of the Arctic zone of Russia was used. In total, about 90,000 daily (in six-hour increments) pressure values were examined. On the basis of these data, a climatic seasonal variation was constructed as an averaging of the values of a given time series at each point in space and for a fixed date. The characteristics of the seasonal course, its amplitude and phase have been studied. These characteristics were analyzed and their geophysical interpretation was carried out. In particular, the minimum and maximum values ​​of the series were determined for the entire region and the time series of these characteristics were constructed. It is shown that the deviation is asymmetric, this is an unobvious research result. For the maximum and minimum, the best approximations were constructed, and these approximations were tested by known methods of statistical analysis, including maximum likelihood, least squares and goodness of fit methods (tests), in particular, the χ2-criterion. The conducted research has applications both purely physical (allows to explain the nature, genesis and distribution of large-scale atmospheric formations in a climatic year) and prognostic (allows understanding and tracking trends in climate, as well as quantitatively assessing the scale and variability of large-scale atmospheric processes). Numerical calculations were performed on the Lomonosov-2 supercomputer of the Lomonosov Moscow State University.

Keywords: time series analysis, climatic seasonal cycle, maximum and minimum pressure values within a climatic year.
1 - 25 of 68 items 1 2 3 > >> 
Information
  • For Readers
  • For Authors
  • For Librarians
Make a Submission
Current Issue
  • Atom logo
  • RSS2 logo
  • RSS1 logo

Russian Digital Libraries Journal

ISSN 1562-5419

Information

  • About the Journal
  • Aims and Scopes
  • Themes
  • Author Guidelines
  • Submissions
  • Privacy Statement
  • Contact
  • eLIBRARY.RU
  • dblp computer science bibliography

Send a manuscript

Authors need to register with the journal prior to submitting or, if already registered, can simply log in and begin the five-step process.

Make a Submission
About this Publishing System

© 2015-2026 Kazan Federal University; Institute of the Information Society