PromptAlign: Automatic Iterative Alignment of a Large Language Model to Expert Annotations without Fine-Tuning

Main Article Content

Alexander Alekseevich Agathonov
Akhat Rustemovich Khusainov
Nikolai Arkadievich Prokopyev

Abstract

Large language models are increasingly used in social and psychological sciences; however, their predictions are systematically biased relative to expert annotations (alignment shift), and attempts to manually refine instructions or to supply examples lead to a “blanket-pulling” effect – degradation of minority classes.


We propose PromptAlign, an automatic prompt-optimization pipeline that requires neither access to model weights nor human involvement in iterative correction, because the taxonomy and the feature space are specified by an expert a priori. The pipeline extracts hybrid linguistic and semantic features using the large language model itself, builds interpretable L1-regularized logistic regressions to localize disagreements with the expert, performs a meta-analysis of errors, and compiles new prompts while controlling the class balance through a class-imbalance index (Blanket Pulling Index, BPI). An experiment on a corpus of 541 Russian-language texts with three psychological classes showed an increase in macro-F1 on the held-out test set from 0.655 to 0.749 and in the minimum per-class F1 from 0.449 to 0.656 (p>0.99 by a paired bootstrap test). The method significantly outperforms both the zero-shot baseline and few-shot prompting. A component-contribution analysis confirms the critical role of semantic interpretation of features and of periodic prompt rebuilding. The final prompt contains five rules that reflect latent expert criteria; unlike the “black boxes” produced by automatic optimizers, these rules are available for inspection and correction by a specialist.

Article Details

How to Cite
Agathonov, A. A., A. R. Khusainov, and N. A. Prokopyev. “PromptAlign: Automatic Iterative Alignment of a Large Language Model to Expert Annotations Without Fine-Tuning”. Russian Digital Libraries Journal, vol. 29, no. 6, Oct. 2026, pp. 2265-92, doi:10.26907/1562-5419-2026-29-6-2265-2292.
Author Biographies

Alexander Alekseevich Agathonov

Cand. Sc. (Physics and Mathematics), Associate Professor at the Department of Higher Mathematics and Mathematical Modeling, N.I. Lobachevsky Institute of Mathematics and Mechanics, Kazan Federal University. Scientific interests: natural language processing, large language models, machine learning, artificial intelligence in medicine, behavior analysis in digital environment, e-Learning.

Akhat Rustemovich Khusainov

Laboratory Assistant, Kazan Federal University, Kazan, Russia. Scientific interests: natural language processing, machine learning, artificial intelligence.

Nikolai Arkadievich Prokopyev

Cand. Sc. (Technology), Associate Professor at the Department of Information Systems, Institute of Computational Mathematics and Information Technologies, Kazan Federal University. Researcher at the Institute of Applied Semiotics of the Academy of Sciences of the Republic of Tatarstan. Scientific interests: natural language processing, artificial intelligence, e-Learning.

References

1. Agathonov, A.A., Khusainov, A.R., Ustin, P.N., & Popov, L.M. Large language models as a tool for predicting types of personal behavior // Russian Psychological Journal. 2026. № 23(2) (in print).
2. Zhou Y., Muresanu A.I., Han Z., et al. Large Language Models Are Human-Level Prompt Engineers. arXiv preprint arXiv:2211.01910. 2023. https://doi.org/10.48550/arXiv.2211.01910 (дата обращения: 15.05.2025).
3. Yang C., Wang X., Lu Y., et al. Large Language Models as Optimizers. arXiv preprint arXiv:2309.03409. 2024. https://doi.org/10.48550/arXiv.2309.03409 (дата обращения: 15.05.2025).
4. Khattab O., Singhvi A., Maheshwari P., et al. DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines. arXiv preprint arXiv:2310.03714. 2023. https://doi.org/10.48550/arXiv.2310.03714 (дата обращения: 15.05.2025).
5. Brown T.B., Mann B., Ryder N., et al. Language Models are Few-Shot Learners // Advances in Neural Information Processing Systems. 2020. Vol. 33. P. 1877–1901.
6. Reynolds L., McDonell K. Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm // Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems. 2021. P. 1–7. https://doi.org/10.1145/3411763.3451760
7. Christiano P.F., Leike J., Brown T.B., et al. Deep Reinforcement Learning from Human Preferences // Advances in Neural Information Processing Systems. 2017. Vol. 30. P. 4299–4307.
8. Ouyang L., Wu J., Jiang X., et al. Training Language Models to Follow Instructions with Human Feedback // Advances in Neural Information Processing Systems. 2022. Vol. 35. P. 27730–27744.
9. Rafailov R., Sharma A., Mitchell E., et al. Direct Preference Optimization: Your Language Model is Secretly a Reward Model // Advances in Neural Information Processing Systems. 2023. Vol. 36. P. 53728–53741.
10. Wei J., Wang X., Schuurmans D., et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models // Advances in Neural Information Processing Systems. 2022. Vol. 35. P. 24824–24837.
11. Davidson T., Warmsley D., Macy M., Weber I. Automated Hate Speech Detection and the Problem of Offensive Language // Proceedings of the 11th International AAAI Conference on Web and Social Media. 2017. P. 512–515. https://doi.org/10.1609/icwsm.v11i1.14955
12. Zampieri M., Malmasi S., Nakov P., et al. Predicting the Type and Target of Offensive Posts in Social Media // Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2019. P. 1415–1420. https://doi.org/10.18653/v1/N19-1144
13. Kolhatkar V., Wu H., Cavasso L., et al. The SFU Opinion and Comments Corpus: A Corpus for the Analysis of Online News Comments // Corpus Pragmatics. 2020. Vol. 4, No. 2. P. 155–190. https://doi.org/10.1007/s41701-019-00065-w
14. Park G., Schwartz H.A., Eichstaedt J.C., et al. Automatic Personality Assessment through Social Media Language // Journal of Personality and Social Psychology. 2015. Vol. 108, No. 6. P. 934–952. https://doi.org/10.1037/pspp0000020
15. Tibshirani R. Regression Shrinkage and Selection via the Lasso // Journal of the Royal Statistical Society. Series B (Methodological). 1996. Vol. 58, No. 1. P. 267–288. https://doi.org/10.1111/j.2517-6161.1996.tb02080.x
16. Efron B. Bootstrap Methods: Another Look at the Jackknife // The Annals of Statistics. 1979. Vol. 7, No. 1. P. 1–26. https://doi.org/10.1214/aos/1176344552
17. Doshi-Velez F., Kim B. Towards a Rigorous Science of Interpretable Machine Learning. arXiv preprint arXiv:1702.08608. 2017. https://doi.org/10.48550/arXiv.1702.08608 (дата обращения: 15.05.2025).
18. Lipton Z.C. The Mythos of Model Interpretability // Communications of the ACM. 2018. Vol. 61, No. 10. P. 36–43. https://doi.org/10.1145/3233231
19. Van Hee C., Lefever E., Hoste V. We Usually Don’t Like Going to the Dentist: Using Common Sense to Detect Irony on Twitter // Computational Linguistics. 2018. Vol. 44, No. 4. P. 793–832. https://doi.org/10.1162/coli_a_00337