• Main Navigation
  • Main Content
  • Sidebar

Russian Digital Libraries Journal

  • Home
  • About
    • About the Journal
    • Aims and Scopes
    • Themes
    • Editor-in-Chief
    • Editorial Team
    • Submissions
    • Open Access Statement
    • Privacy Statement
    • Contact
  • Current
  • Archives
  • Register
  • Login
  • Search
Published since 1998
ISSN 1562-5419
16+
Language
  • Русский
  • English

Search

Advanced filters

Search Results

Post-Correction of Weak Transcriptions by Large Language Models in the Iterative Process of Handwritten Text Recognition

Valerii Pavlovich Zykov, Leonid Moiseevich Mestetskiy
1385-1414
Abstract:

This paper addresses the problem of accelerating the construction of accurate editorial annotations for handwritten archival texts within an incremental training cycle based on weak transcription. Unlike our previously published results, the present work focuses on integrating automatic post-correction of weak transcriptions using large language models (LLMs). We propose and implement a protocol for applying LLMs at the line level in a few-shot setup with carefully designed prompts and strict output format control (preservation of pre-reform orthography, protection of proper names and numerals, prohibition of structural changes to lines). Experiments are conducted on the corpus of diaries by A.V. Sukhovo-Kobylin. As the base recognition model, we use the line-level variant of the Vertical Attention Network (VAN). Results show that LLM post-correction–exemplified by the ChatGPT-4o service–substantially improves the readability of weak transcriptions and significantly reduces the word error rate (in our experiments by about −12 percentage points), without degrading the character error rate. Another service tested, DeepSeek-R1, demonstrated less stable behavior. We discuss practical prompt engineering, limitations (context length limits, risk of “hallucinations”), and provide recommendations for the safe integration of LLM post-correction into an iterative annotation pipeline to reduce expert annotators’ workload and speed up the digitization of historical archives.

Keywords: handwritten text recognition, weak markup, Vertical Attention Network (VAN), large language models (LLM), post-correction, iterative retraining.

Comparison of Approaches to the Problem of Automatic Generation of Official Reply Letters using LLM

Ivan Evgenievich Nikolaev, Andrey Vitalievich Melnikov, Kirill Evgenievich Alekseev, Alexander Sergeevich Belonogov, Mikhail Aleksandrovich Rusanov
1174-1188
Abstract:

One of the key automation challenges for government agencies is the preparation of official response letters. This article presents an empirical comparison of two approaches to the automatic generation of official response letters: one based on templates defining the letter structure and one based on relevant letter examples selected using RAG. The original GovLetter dataset, which includes real-life business correspondence from government agencies in the Khanty-Mansi Autonomous Okrug – Yugra, was used as a basis. Generation was performed using a locally deployed open-source language model. The quality of the results was assessed according to 12 criteria using the Schema-Guided Reasoning (SGR) framework and the LLM-as-a-Judge methodology. The experimental results demonstrate that the example-based approach outperforms the template-based approach in most metrics, particularly in the accuracy of arguments, formal tone, and correct formatting. The obtained results confirm the potential of solutions based on document history for the effective automation of official response letter preparation.

Keywords: text generation, official correspondence, large language models, RAG, quality assessment, SGR.

The Problem of Building Synthetic Psychological Data: Experience in Modeling Reactions to Frustration

Anfisa Anvarovna Chuganskaya, Danil Alekseevich Kireev, Ivan Valentinovich Smirnov, Oleg Georgievich Grigoriev
1235-1252
Abstract:

The issue of generating synthetic data for psychological research remains relevant and complex. The problems of confidentiality, reliability, validity, and accuracy of conclusions are unevenly represented in different areas of psychology and are actually related to the use of synthetic data in related sciences such as medicine, sociology, history, political science, and economics. The study of various psychological phenomena in large social groups is associated with the analysis of complex and difficult-to-formalize constructs. Synthetic data refers to artificially generated data based on algorithms and modeling.


  1. Rosenzweig's classification of types of reactions to frustration was chosen as the basis for this study. When analyzing online discourse, there is a problem of the low number of certain types. This is especially true for the class of impunitive reactions. The paper analyzes the possibility of creating text corpora using synthetic data of reactions to frustration generated by large language models.

During the experiments, the experts created prompts and generated examples of impulsive reactions using four large language models, with 10 examples of each type of reaction. They then evaluated the contextual validity and quality of the generated responses.


The results obtained allowed them to identify the weaknesses in generating texts with complex psychological phenomena for training neural network models.

Keywords: frustration, large language models (LLM), synthetic data, artificial intelligence, prompt, online discussion, classification of the Rosenzweig.

Errors of Artificial Intelligence in Solving Combinatorial Problems

Elena Vladimirovna Krutenko, Boris Yakovlevich Steinberg
428-441
Abstract:

Several combinatorial exercises have been considered, which artificial intelligence solves with errors. The representatives of artificial intelligence examined are ChatGPT and DeepSeek systems. Questions (prompts) to these systems are provided, and the obtained answers are analyzed. Hypotheses are proposed regarding the reasons for the errors made by artificial intelligence when solving the tasks under consideration. It is suggested that similar errors may occur when using artificial intelligence for software development and other applications. Topics for further research are proposed, which may be of interest for determining the conditions for the continued use of artificial intelligence.

Keywords: neural network, artificial intelligence, errors, combinatorics, technical specifications.

Abstractive Summarization for Trade News Analysis Based on a New Domain-Specific Dataset

Daria Andreevna Lyutova, Valentin Andreevich Malykh
1120-1137
Abstract:

We present TradeNewsSum—a corpus for abstractive summarization of international trade news—covering Russian- and English-language publications from domain-specific sources. All summaries are manually prepared following unified guidelines. We conducted experiments with fine-tuning transformer and seq2seq models and performed automatic evaluation using the LLM-as-a-judge scheme. LLaMA 3.1 in instruction-prompting mode achieved the best results, showing high scores across metrics, including factual completeness.

Keywords: abstractive summarization, multilingual corpus, international trade news, sanctions, trade regimes, TradeNewsSum, transformers, large language models, LLM-as-a-judge, NER-based entity evaluation.
1 - 5 of 5 items
Information
  • For Readers
  • For Authors
  • For Librarians
Make a Submission
Current Issue
  • Atom logo
  • RSS2 logo
  • RSS1 logo

Russian Digital Libraries Journal

ISSN 1562-5419

Information

  • About the Journal
  • Aims and Scopes
  • Themes
  • Author Guidelines
  • Submissions
  • Privacy Statement
  • Contact
  • eLIBRARY.RU
  • dblp computer science bibliography

Send a manuscript

Authors need to register with the journal prior to submitting or, if already registered, can simply log in and begin the five-step process.

Make a Submission
About this Publishing System

© 2015-2026 Kazan Federal University; Institute of the Information Society