PromptAlign: Automatic Iterative Alignment of a Large Language Model to Expert Annotations without Fine-Tuning
Main Article Content
Abstract
Large language models are increasingly used in social and psychological sciences; however, their predictions are systematically biased relative to expert annotations (alignment shift), and attempts to manually refine instructions or to supply examples lead to a “blanket-pulling” effect – degradation of minority classes.
We propose PromptAlign, an automatic prompt-optimization pipeline that requires neither access to model weights nor human involvement in iterative correction, because the taxonomy and the feature space are specified by an expert a priori. The pipeline extracts hybrid linguistic and semantic features using the large language model itself, builds interpretable L1-regularized logistic regressions to localize disagreements with the expert, performs a meta-analysis of errors, and compiles new prompts while controlling the class balance through a class-imbalance index (Blanket Pulling Index, BPI). An experiment on a corpus of 541 Russian-language texts with three psychological classes showed an increase in macro-F1 on the held-out test set from 0.655 to 0.749 and in the minimum per-class F1 from 0.449 to 0.656 (p>0.99 by a paired bootstrap test). The method significantly outperforms both the zero-shot baseline and few-shot prompting. A component-contribution analysis confirms the critical role of semantic interpretation of features and of periodic prompt rebuilding. The final prompt contains five rules that reflect latent expert criteria; unlike the “black boxes” produced by automatic optimizers, these rules are available for inspection and correction by a specialist.
Article Details
References
2. Zhou Y., Muresanu A.I., Han Z., et al. Large Language Models Are Human-Level Prompt Engineers. arXiv preprint arXiv:2211.01910. 2023. https://doi.org/10.48550/arXiv.2211.01910 (дата обращения: 15.05.2025).
3. Yang C., Wang X., Lu Y., et al. Large Language Models as Optimizers. arXiv preprint arXiv:2309.03409. 2024. https://doi.org/10.48550/arXiv.2309.03409 (дата обращения: 15.05.2025).
4. Khattab O., Singhvi A., Maheshwari P., et al. DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines. arXiv preprint arXiv:2310.03714. 2023. https://doi.org/10.48550/arXiv.2310.03714 (дата обращения: 15.05.2025).
5. Brown T.B., Mann B., Ryder N., et al. Language Models are Few-Shot Learners // Advances in Neural Information Processing Systems. 2020. Vol. 33. P. 1877–1901.
6. Reynolds L., McDonell K. Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm // Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems. 2021. P. 1–7. https://doi.org/10.1145/3411763.3451760
7. Christiano P.F., Leike J., Brown T.B., et al. Deep Reinforcement Learning from Human Preferences // Advances in Neural Information Processing Systems. 2017. Vol. 30. P. 4299–4307.
8. Ouyang L., Wu J., Jiang X., et al. Training Language Models to Follow Instructions with Human Feedback // Advances in Neural Information Processing Systems. 2022. Vol. 35. P. 27730–27744.
9. Rafailov R., Sharma A., Mitchell E., et al. Direct Preference Optimization: Your Language Model is Secretly a Reward Model // Advances in Neural Information Processing Systems. 2023. Vol. 36. P. 53728–53741.
10. Wei J., Wang X., Schuurmans D., et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models // Advances in Neural Information Processing Systems. 2022. Vol. 35. P. 24824–24837.
11. Davidson T., Warmsley D., Macy M., Weber I. Automated Hate Speech Detection and the Problem of Offensive Language // Proceedings of the 11th International AAAI Conference on Web and Social Media. 2017. P. 512–515. https://doi.org/10.1609/icwsm.v11i1.14955
12. Zampieri M., Malmasi S., Nakov P., et al. Predicting the Type and Target of Offensive Posts in Social Media // Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2019. P. 1415–1420. https://doi.org/10.18653/v1/N19-1144
13. Kolhatkar V., Wu H., Cavasso L., et al. The SFU Opinion and Comments Corpus: A Corpus for the Analysis of Online News Comments // Corpus Pragmatics. 2020. Vol. 4, No. 2. P. 155–190. https://doi.org/10.1007/s41701-019-00065-w
14. Park G., Schwartz H.A., Eichstaedt J.C., et al. Automatic Personality Assessment through Social Media Language // Journal of Personality and Social Psychology. 2015. Vol. 108, No. 6. P. 934–952. https://doi.org/10.1037/pspp0000020
15. Tibshirani R. Regression Shrinkage and Selection via the Lasso // Journal of the Royal Statistical Society. Series B (Methodological). 1996. Vol. 58, No. 1. P. 267–288. https://doi.org/10.1111/j.2517-6161.1996.tb02080.x
16. Efron B. Bootstrap Methods: Another Look at the Jackknife // The Annals of Statistics. 1979. Vol. 7, No. 1. P. 1–26. https://doi.org/10.1214/aos/1176344552
17. Doshi-Velez F., Kim B. Towards a Rigorous Science of Interpretable Machine Learning. arXiv preprint arXiv:1702.08608. 2017. https://doi.org/10.48550/arXiv.1702.08608 (дата обращения: 15.05.2025).
18. Lipton Z.C. The Mythos of Model Interpretability // Communications of the ACM. 2018. Vol. 61, No. 10. P. 36–43. https://doi.org/10.1145/3233231
19. Van Hee C., Lefever E., Hoste V. We Usually Don’t Like Going to the Dentist: Using Common Sense to Detect Irony on Twitter // Computational Linguistics. 2018. Vol. 44, No. 4. P. 793–832. https://doi.org/10.1162/coli_a_00337

This work is licensed under a Creative Commons Attribution 4.0 International License.
Presenting an article for publication in the Russian Digital Libraries Journal (RDLJ), the authors automatically give consent to grant a limited license to use the materials of the Kazan (Volga) Federal University (KFU) (of course, only if the article is accepted for publication). This means that KFU has the right to publish an article in the next issue of the journal (on the website or in printed form), as well as to reprint this article in the archives of RDLJ CDs or to include in a particular information system or database, produced by KFU.
All copyrighted materials are placed in RDLJ with the consent of the authors. In the event that any of the authors have objected to its publication of materials on this site, the material can be removed, subject to notification to the Editor in writing.
Documents published in RDLJ are protected by copyright and all rights are reserved by the authors. Authors independently monitor compliance with their rights to reproduce or translate their papers published in the journal. If the material is published in RDLJ, reprinted with permission by another publisher or translated into another language, a reference to the original publication.
By submitting an article for publication in RDLJ, authors should take into account that the publication on the Internet, on the one hand, provide unique opportunities for access to their content, but on the other hand, are a new form of information exchange in the global information society where authors and publishers is not always provided with protection against unauthorized copying or other use of materials protected by copyright.
RDLJ is copyrighted. When using materials from the log must indicate the URL: index.phtml page = elbib / rus / journal?. Any change, addition or editing of the author's text are not allowed. Copying individual fragments of articles from the journal is allowed for distribute, remix, adapt, and build upon article, even commercially, as long as they credit that article for the original creation.
Request for the right to reproduce or use any of the materials published in RDLJ should be addressed to the Editor-in-Chief A.M. Elizarov at the following address: amelizarov@gmail.com.
The publishers of RDLJ is not responsible for the view, set out in the published opinion articles.
We suggest the authors of articles downloaded from this page, sign it and send it to the journal publisher's address by e-mail scan copyright agreements on the transfer of non-exclusive rights to use the work.