On Improving RAG Quality for Contradiction Detection in Court Rulings
Main Article Content
Abstract
The paper extends and continues the work on a pilot system for detecting contradictions in court rulings. The tool aims to identify contradictions between court rulings on administrative offense cases and the legislative norms governing them. The previously presented pipeline, which includes rule-based document preprocessing, search over a structured repository of legal norms, and LLM fine-tuning using a LoRA adapter, demonstrated fairly high metrics [1].
However, this implementation had a number of limitations, primarily related to the RAG component. Unlike LLM, the RAG module was not trained on legal texts, which led to a substantial loss of relevant fragments for comparison at the retrieval stage.
The present work therefore focuses on a detailed quality analysis and improving the RAG component of the system. A dataset for evaluating premise retrieval recall was compiled, comprising 30 documents, along with automated tests containing all relevant sentence pairs (more than one hundred pairs in total). Using this data, the recall@n was calculated for the original (baseline) version of the model.
To improve the recall of the retrieved data, the following experiments were conducted: (1) fine-tuning the original embeddings on the system's training dataset; (2) adding a reranking model, both with and without fine-tuning; (3) combining variations of base/fine-tuned embeddings/reranker across several values of the top_k parameter (the number of retrieved matches); and (4) replacing the sentence embedding model with alternative options.
These experiments improved RAG recall from 72% to 90%. However, increasing the number of the retrieved fragments reduced the system's overall precision metrics, indicating a need for further training.
Other notable optimizations include reorganizing the structure of the legal codes database, expanding the test set, and improving the algorithm for filtering out irrelevant sentences from the rulings.
Article Details
References
2. Harabagiu S., Hickl A., Lacatusu F. Negation, contrast and contradiction in text processing // Proc. of the 21st National Conference on Artificial Intelligence (AAAI-06). 2006. P. 755–762.
3. De Marneffe M.-C., Rafferty A.N., Manning C.D. Finding contradictions in text // Proc. of the 46th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies (ACL-08: HLT). 2008. P. 1039–1047.
4. Mehdad Y., Negri M., Federico M. Towards cross-lingual textual entailment // Proc. of the 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT 2010). 2010. P. 321–324.
5. Bowman S., Angeli G., Potts C., Manning C.D. A large annotated corpus for learning natural language inference // arXiv:1508.05326. 2015. https://doi.org/10.48550/arXiv.1508.05326
6. Dragos V. Detection of contradictions by relation matching and uncertainty assessment // Procedia Computer Science. 2017. Vol. 112. P. 71–80.
7. Pielka M., et al. A linguistic investigation of machine learning based contradiction detection models: an empirical analysis and future perspectives // arXiv:2210.10434. 2022. https://doi.org/10.48550/arXiv.2210.10434
8. Putra I M. S., Siahaan D., Saikhu A. Recognizing textual entailment: a review of resources, approaches, applications, and challenges // ICT Express. 2024. Vol. 10, No. 1. P. 132–155. https://doi.org/10.1016/j.icte.2023.08.012
9. Shajalal M., et al. Textual entailment recognition with semantic features from empirical text representation // arXiv:2210.09723v4. 2022. https://doi.org/10.48550/arXiv.2210.09723
10. Rahimi Z., ShamsFard M. A knowledge-based approach for recognizing textual entailments with a focus on causality and contradiction // Research Square preprint. 2024. https://doi.org/10.21203/rs.3.rs-3826973/v1
11. Asif M., Ahmed Khan T., Song W.-C. Evaluating large language models for optimized intent translation and contradiction detection using KNN in IBN // IEEE Access. 2025. Vol. 13. P. 20316–20327.
12. Wu X., Niu X., Rahman R. Topological analysis of contradictions in text // Proc. of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2022). 2022. P. 2478–2483. https://doi.org/10.1145/3477495.3531881
13. Williams A., Nangia N., Bowman S. A broad-coverage challenge corpus for sentence understanding through inference // Proc. of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT 2018). 2018. P. 1112–1122. https://doi.org/10.18653/v1/N18-1101
14. Camarda A.D., Ianni G. A study on contradiction detection using a neuro-symbolic approach // Proc. of 40th Italian Conference on Computational Logic (CILC 2025). 2025, paper 08.
15. Feng X., Hunter A. Identification of entailment and contradiction relations between natural language sentences: a neurosymbolic approach //
arXiv:2405.01259. 2024. https://doi.org/10.48550/arXiv.2405.01259
16. Walton D., Reed C. Argumentation schemes and defeasible inferences // Working Notes of the ECAI 2002 Workshop on Computational Models of Natural Argument (CMNA 2002). 2002. P. 45–55.
17. Ibraim S.S., Mendonça P.C.C. Book Review: «Argumentation Schemes» // Ensaio Pesquisa em Educação em Ciências. 2013. Vol. 15, No. 3. P. 255–262.
18. Walton D. Legal argumentation and evidence // Penn State University Press, 2010. 392 p.
19. Mantravadi A., et al. LegalWiz: a multi-agent generation framework for contradiction detection in legal documents // arXiv:2510.03418. 2025. https://doi.org/10.48550/arXiv.2510.03418
20. Koreeda Y., Manning C.D. ContractNLI: a dataset for document-level natural language inference for contracts // Findings of the Association for Computational Linguistics: EMNLP 2021. 2021, P. 1907–1919. https://doi.org/10.18653/v1/2021.findings-emnlp.164
21. Nguyen H.-T., et al. Transformer-based approaches for legal text processing: JNLP team – COLIEE 2021 // The Review of Socionetwork Strategies. 2022. Vol. 16, No. 1. P. 135–155. https://doi.org/10.1007/s12626-022-00102-2
22. Nguyen C., et al. CAPTAIN at COLIEE 2023: efficient methods for legal information retrieval and entailment tasks // arXiv:2401.03551. 2024. https://doi.org/10.48550/arXiv.2401.03551
23. Aires J.P., Strube de Lima V.L., Meneguzzi F. Identifying potential conflicts between norms in contracts // 18th International Workshop on Coordination, Organizations, Institutions, and Norms (COIN 2015). 2016. P. 179–188.
24. Surana S., Dembla S., Bihani P. Identifying contradictions in the legal proceedings using natural language models // SN Computer Science. 2022. Vol. 3, No. 3. Art. no. 187. https://doi.org/10.1007/s42979-022-01075-3
25. Ariai F., Mackenzie J., Demartini G. Natural language processing for the legal domain: a survey of tasks, datasets, models, and challenges // ACM Computing Surveys. 2025. Vol. 58, No. 6. P. 1–37. https://doi.org/10.48550/arXiv.2410.21306
26. Hosseini M.J., Petrov A., Fabrikant A., Louis A. A synthetic data approach for domain generalization of NLI models // Proc. of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024). 2024. P. 2212–2226. https://doi.org/10.18653/v1/2024.acl-long.120

This work is licensed under a Creative Commons Attribution 4.0 International License.
Presenting an article for publication in the Russian Digital Libraries Journal (RDLJ), the authors automatically give consent to grant a limited license to use the materials of the Kazan (Volga) Federal University (KFU) (of course, only if the article is accepted for publication). This means that KFU has the right to publish an article in the next issue of the journal (on the website or in printed form), as well as to reprint this article in the archives of RDLJ CDs or to include in a particular information system or database, produced by KFU.
All copyrighted materials are placed in RDLJ with the consent of the authors. In the event that any of the authors have objected to its publication of materials on this site, the material can be removed, subject to notification to the Editor in writing.
Documents published in RDLJ are protected by copyright and all rights are reserved by the authors. Authors independently monitor compliance with their rights to reproduce or translate their papers published in the journal. If the material is published in RDLJ, reprinted with permission by another publisher or translated into another language, a reference to the original publication.
By submitting an article for publication in RDLJ, authors should take into account that the publication on the Internet, on the one hand, provide unique opportunities for access to their content, but on the other hand, are a new form of information exchange in the global information society where authors and publishers is not always provided with protection against unauthorized copying or other use of materials protected by copyright.
RDLJ is copyrighted. When using materials from the log must indicate the URL: index.phtml page = elbib / rus / journal?. Any change, addition or editing of the author's text are not allowed. Copying individual fragments of articles from the journal is allowed for distribute, remix, adapt, and build upon article, even commercially, as long as they credit that article for the original creation.
Request for the right to reproduce or use any of the materials published in RDLJ should be addressed to the Editor-in-Chief A.M. Elizarov at the following address: amelizarov@gmail.com.
The publishers of RDLJ is not responsible for the view, set out in the published opinion articles.
We suggest the authors of articles downloaded from this page, sign it and send it to the journal publisher's address by e-mail scan copyright agreements on the transfer of non-exclusive rights to use the work.