Abstract:
The paper extends and continues the work on a pilot system for detecting contradictions in court rulings. The tool aims to identify contradictions between court rulings on administrative offense cases and the legislative norms governing them. The previously presented pipeline, which includes rule-based document preprocessing, search over a structured repository of legal norms, and LLM fine-tuning using a LoRA adapter, demonstrated fairly high metrics [1].
However, this implementation had a number of limitations, primarily related to the RAG component. Unlike LLM, the RAG module was not trained on legal texts, which led to a substantial loss of relevant fragments for comparison at the retrieval stage.
The present work therefore focuses on a detailed quality analysis and improving the RAG component of the system. A dataset for evaluating premise retrieval recall was compiled, comprising 30 documents, along with automated tests containing all relevant sentence pairs (more than one hundred pairs in total). Using this data, the recall@n was calculated for the original (baseline) version of the model.
To improve the recall of the retrieved data, the following experiments were conducted: (1) fine-tuning the original embeddings on the system's training dataset; (2) adding a reranking model, both with and without fine-tuning; (3) combining variations of base/fine-tuned embeddings/reranker across several values of the top_k parameter (the number of retrieved matches); and (4) replacing the sentence embedding model with alternative options.
These experiments improved RAG recall from 72% to 90%. However, increasing the number of the retrieved fragments reduced the system's overall precision metrics, indicating a need for further training.
Other notable optimizations include reorganizing the structure of the legal codes database, expanding the test set, and improving the algorithm for filtering out irrelevant sentences from the rulings.