Pragmatic Boundaries of Greetings and Goodbyes: Evaluating LLMs on the Russian Multimedia Politeness Corpus
Main Article Content
Abstract
This study investigates the ability of large language models to detect greeting and goodbye sequences in Russian-language dialogues and determine their precise pragmatic boundaries. YandexGPT 5.1 Pro, GigaChat 2 Max, and GPT-5 are evaluated on two tasks: classifying dialogues based on the presence of these sequences and extracting individual utterances, taking their intentional boundaries into account. The experiment is conducted on 216 dialogues from the Russian Multimedia Politeness Corpus, thereby accounting for the broad communicative and social context of interaction. GPT-5 achieves the strongest dialogue-level classification performance, whereas YandexGPT and GigaChat attain near-perfect precision through conservative prediction strategies. Performance declines markedly for all models when exact boundaries must be extracted. Error analysis shows that the models systematically over-include contextual nonverbal actions and miss non-conventional markers, including gestures, address forms, discourse markers, and implicatures indicating future contact. The article presents an updated description of the corpus and its web interface, as well as facilities for constructing diagnostic subcorpora and benchmarks. The findings underscore the need to model discourse structure and the contextual variation of everyday communication more precisely.
Article Details
References
2. Kádár D.Z., Haugh M. Understanding politeness. Cambridge University Press, 2013.
3. Schegloff E.A. Sequence Organization in Interaction: A Primer in Conversation Analysis I. Cambridge University Press, 2007.
4. Ruis L. et al. The goldilocks of pragmatic understanding: Fine-tuning strategy matters for implicature resolution by LLMs // arXiv:2210.14986. 2024.
https://doi.org/10.48550/arXiv.2210.14986
5. Jian M., Narayanaswamy S. Are LLMs good pragmatic speakers? //
arXiv:2411.01562. 2024. https://doi.org/10.48550/arXiv.2411.01562
6. Hu J., Floyd S., Jouravlev O., Fedorenko E., Gibson E. A fine-grained comparison of pragmatic language understanding in humans and language models // Proc. of the 61st Annual Meeting of the Association for Computational Linguistics. 2023. Vol. 1. P. 4194–4213.
7. Lipkin B., Wong L., Grand G., Tenenbaum J.B. Evaluating statistical language models as pragmatic reasoners. // arXiv:2305.01020. 2023.
https://doi.org/10.48550/arXiv.2305.01020
8. Iida A., Okuoka K., Omori T., Nakashima R., Osawa M. LLM-based evaluation of utterances with implicature understanding: A preliminary study // Proc. of the 13th International Conference on Human-Agent Interaction. 2025. P. 485–487. https://doi.org/10.1145/3765766.3765841
9. Yerukola A., Vaduguru S., Fried D., Sap M. Is the pope catholic? Yes, the pope is catholic. Generative evaluation of non-literal intent resolution in LLMs // Proc. of the 62nd Annual Meeting of the Association for Computational Linguistics. 2024. Vol. 2. P. 265–275.
10. Liu R., Sumers T. R., Dasgupta I., Griffiths T.L. How do large language models navigate conflicts between honesty and helpfulness? // Proc. of the 41st International Conference on Machine Learning. 2024. P. 31844–31865.
11. Chen Z., Yang R., Zhao Z., Cai D., He X. Dialogue act recognition via CRF-attentive structured network // Proc. of the 41st International ACM SIGIR Conference on Research & Development in Information Retrieval. 2018. P. 225–234.
12. Qamar A., Tong J., Huang R. Do LLMs understand dialogues? A case study on dialogue acts. // Proc. of the 63rd Annual Meeting of the Association for Computational Linguistics. 2025. Vol. 1. P. 26219–26237.
13. Schegloff E., Sacks H. Opening up closings // Semiotica. 1973. Vol. 8, No. 4. P. 289–327.
14. Klokova K., Krongauz M., Shulginov V., Yudina T. Towards a russian multimedia politeness corpus. // Proc. of the International Conference “Dialogue 2023”. 2023. P. 233–244.
15. Terkourafi M. An argument for a frame-based approach to politeness: Evidence from the use of the imperative in Cypriot Greek // Pragmatics and beyond. New series. 2005. Vol. 139. P. 99–116.
16. Laver J. Communicative Functions of Phatic Communion // Organisation of Behaviour in Face-to-face Interaction. 1975. P. 215–238.
17. Laver J. Linguistic Routines and Politeness in Greeting and Parting // Conversational routines. 1981. P. 289–305.
18. Brown T.B. et al. Language models are few-shot learners // arXiv:2005.14165. 2020. https://doi.org/10.48550/arXiv.2005.14165
19. Winkler W. E. String comparator metrics and enhanced decision rules in the fellegi-sunter model of record linkage // Proc. of the Section on Survey Research. 1990. P. 354–359.
20. Cahyono S.C. Comparison of document similarity measurements in scientific writing using jaro-winkler distance method and paragraph vector method // IOP Conference Series: Materials Science and Engineering. 2019. Vol. 662, No. 5. https://doi.org/10.1088/1757-899X/662/5/052016

This work is licensed under a Creative Commons Attribution 4.0 International License.
Presenting an article for publication in the Russian Digital Libraries Journal (RDLJ), the authors automatically give consent to grant a limited license to use the materials of the Kazan (Volga) Federal University (KFU) (of course, only if the article is accepted for publication). This means that KFU has the right to publish an article in the next issue of the journal (on the website or in printed form), as well as to reprint this article in the archives of RDLJ CDs or to include in a particular information system or database, produced by KFU.
All copyrighted materials are placed in RDLJ with the consent of the authors. In the event that any of the authors have objected to its publication of materials on this site, the material can be removed, subject to notification to the Editor in writing.
Documents published in RDLJ are protected by copyright and all rights are reserved by the authors. Authors independently monitor compliance with their rights to reproduce or translate their papers published in the journal. If the material is published in RDLJ, reprinted with permission by another publisher or translated into another language, a reference to the original publication.
By submitting an article for publication in RDLJ, authors should take into account that the publication on the Internet, on the one hand, provide unique opportunities for access to their content, but on the other hand, are a new form of information exchange in the global information society where authors and publishers is not always provided with protection against unauthorized copying or other use of materials protected by copyright.
RDLJ is copyrighted. When using materials from the log must indicate the URL: index.phtml page = elbib / rus / journal?. Any change, addition or editing of the author's text are not allowed. Copying individual fragments of articles from the journal is allowed for distribute, remix, adapt, and build upon article, even commercially, as long as they credit that article for the original creation.
Request for the right to reproduce or use any of the materials published in RDLJ should be addressed to the Editor-in-Chief A.M. Elizarov at the following address: amelizarov@gmail.com.
The publishers of RDLJ is not responsible for the view, set out in the published opinion articles.
We suggest the authors of articles downloaded from this page, sign it and send it to the journal publisher's address by e-mail scan copyright agreements on the transfer of non-exclusive rights to use the work.