A Hybrid Approach to Autonomous Control of a Robotic Device That Combines Reinforcement Learning and Operator-Based Demonstration Learning
Main Article Content
Abstract
A reproducible hybrid approach to learning autonomous control of a robotic device has been proposed and experimentally verified, implemented using publicly available tools without specialized computing hardware. The approach combines learning from operator demonstrations, reinforcement learning, curriculum learning, and domain randomization into a single sequential framework. Based on a comparative analysis of five prototypes representing key classes of architectural solutions, applicable approaches for resource-constrained environments are identified. Three controlled experiments in Unity ML-Agents demonstrated the superiority of the hybrid combination over isolated approaches in terms of both learning speed and robustness to environmental changes. Limitations associated with catastrophic forgetting are outlined, and prospects for integrating large language models to generate synthetic training data are identified.
Article Details
References
2. Kugurakova V.V. et al. VR-Telecontrol of Multi-Arm Devices: Problems, Hypotheses, Problem Statement // Russian Digital Libraries Journal. 2022. Vol. 25, No. 5. P. 441–471. https://doi.org/10.26907/1562-5419-2022-25-5-441-471
3. Zykova A.E., Zykov E.Yu., Kugurakova V.V. Primenenie tekhnologiy virtualnoy realnosti dlya udalennogo upravleniya robotizirovannymi ustroystvami // Itogovaya nauchno-prakticheskaya konferentsiya KFU. Kazan, 2023. P. 87. EDN GDBPQP.
4. Sergunin I.D., Zykov E.Yu. «RUSAK» – modulʹnoe reshenie sistemy upravleniya udalyonnymi ustrojstvami s integrirovannymi zashchishchyonnymi kanalami svyazi // Sostoyanie i perspektivy razvitiya sovremennoj nauki po napravleniyu «Robototekhnika». 2023. P. 265–271.
5. Kugurakova V.V., Iskhakov R.T. et al. Tochechnoe teleupravlenie dlya selʹskokhozyajstvennoj tekhniki, vypolnyayushchej zadachi na osnove II // International Forum KAZAN DIGITAL WEEK – 2025: conference papers, volume 1. 2025. P. 1964–1971.
6. Programma dlya sinkhronizatsii stereoizobrazheniya so stereopary, razmeshchyonnoj na udalyonnom dvigayushchemsya ustrojstve: Svidetelʹstvo o gos. reg. programmy dlya EVM № 2022684277 / Kugurakova V.V., Khafizov M.R. Gabdullina D.R.; KFU. Date of registration: 13.12.2022.
7. Programma adaptivnogo upravleniya tipovymi marshrutami i rutinnymi dejstviyami modeli robota-manipulyatora v virtualʹnoj srede: Svidetelʹstvo o gos. reg. programmy dlya ЭВМ № 2024689377 / Kugurakova V.V., Iskhakov R.T. et al.; KFU. Date of registration: 05.12.2022.
8. Iskhakov R.T. Razrabotka generatora bolʹshogo kolichestva variatsij dvizhenij robotizirovannogo ustrojstva na osnove bazovykh zapisannykh dejstvij operatora: vypusknaya kvalifikatsionnaya rabota: 09.03.04 «Programmnaya inzheneriya» / R.T. Iskhakov; Kazan (Volga Region) Federal University. Kazan, 2024. 41 p.
9. Guzaerov D.A., Sergunin I.D., Kugurakova V.V. Infrared Video-Based Object Recognition for Agricultural Research Field Uses // BIO Web of Conferences. EDP Sciences, 2024. Vol. 141. No. 01025. https://doi.org/10.1051/bioconf/202414101025
10. Zare M. et al. A survey of imitation learning: Algorithms, recent developments, and challenges // IEEE Transactions on Cybernetics. 2024. Vol. 54, No. 12. P. 7173–7186. https://doi.org/10.48550/arXiv.2309.02473
11. Arulkumaran K. et al. Deep Reinforcement Learning: A Brief Survey // IEEE Signal Processing Magazine. 2017. Vol. 34, No. 6. P. 26–38. https://doi.org/10.1109/MSP.2017.2743240
12. Kober J., Bagnell J.A., Peters J. Reinforcement Learning in Robotics: A Survey // The International Journal of Robotics Research. 2013. Vol. 32, No. 11. P. 1238–1274. https://doi.org/10.1177/0278364913495721
13. Chebotar Y. et al. Actionable models: Unsupervised offline reinforcement learning of robotic skills //arXiv preprint arXiv:2104.07749. 2021.
14. Bengio Y. et al. Curriculum Learning // Proceedings of ICML. 2009. P. 41–48. https://doi.org/10.1145/1553374.1553380
15. Graves A. et al. Automated Curriculum Learning for Neural Networks // Proceedings of ICML. 2017. P. 1311–1320. https://doi.org/10.48550/arXiv.1704.03003
16. Chukwurah N., Adebayo A.S., Ajayi O.O. Sim-to-real transfer in robotics: Addressing the gap between simulation and real-world performance //International Journal of Robotics and Simulation. 2024. Vol. 6, No. 1. P. 89–102. https://doi.org/10.54660/.IJFMR.2024.5.1.33-39
17. Tobin J. et al. Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World // IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). 2017. P. 23–30. https://doi.org/10.1109/IROS.2017.8202133
18. Pinto L. et al. Robust Adversarial Reinforcement Learning // Proceedings of the 34th International Conference on Machine Learning (ICML). 2017. https://doi.org/10.48550/arXiv.1703.02702
19. Abnar S. et al. Exploring the limits of large scale pre-training // arXiv preprint arXiv:2110.02095. 2021.
20. Tesla Optimus // teslarobots.ru. 2025. URL: https://teslarobots.ru/ (accessed at 05.02.26)
21. Brohan A. et al. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control // arXiv preprint arXiv:2307.15818. 2023.
22. Fu Z. et al. Humanplus: Humanoid shadowing and imitation from humans // arXiv preprint arXiv:2406.10454. 2024.
23. Xpeng Shares Achievements in Physical AI: Unveils XPeng VLA 2.0, Next-Gen Iron // xpeng.com. 2025. URL: https://www.xpeng.com/pressroom/news/ 019a56f54fe99a2a0a8d8a0282e402b7 (дата обращения: 04.03.2026).
24. Figure AI Opens New Helix Lab to Accelerate Learning for Its Humanoid Robots // mlq.ai. 2025. URL: https://mlq.ai/news/figure-ai-opens-new-helix-lab-to-accelerate-learning-for-its-humanoid-robots/ (accessed at 04.03.26)
25. Apanasevich I. et al. Green-VLA: Staged Vision-Language-Action Model for Generalist Robots //arXiv preprint arXiv:2602.00919. 2026.
26. Yu A., Palefsky-Smith R., Bedi R. Deep reinforcement learning for simulated autonomous vehicle control // Standford University Course Project Reports: Winter. 2016. P. 1–7. URL: https://cs231n.stanford.edu/reports/2016/pdfs/112_Report.pdf
27. Done C. et al. Reinforcement Learning-Driven Prosthetic Hand Actuation in a Virtual Environment Using Unity ML-Agents // Virtual Worlds. MDPI. 2025. Vol. 4, No. 4. P. 53. https://doi.org/10.3390/virtualworlds4040053
28. Kirkpatrick J. et al. Overcoming Catastrophic Forgetting in Neural Networks // Proceedings of the National Academy of Sciences. 2017. Vol. 114, No. 13. P. 3521–3526. https://doi.org/10.1073/pnas.1611835114
29. Fedus W. et al. Revisiting fundamentals of experience replay //International conference on machine learning. 2020. P. 3061–3071. https://doi.org/10.48550/arXiv.2007.06700
30. Caruana R. Multitask Learning // Machine Learning. 1997. Vol. 28. P. 41–70. https://doi.org/10.1023/A:1007379606734
31. Pateria S. et al. Hierarchical reinforcement learning: A comprehensive survey //ACM Computing Surveys (CSUR). 2021. Vol. 54, No. 5. P. 1–35. https://doi.org/10.1145/3453160

This work is licensed under a Creative Commons Attribution 4.0 International License.
Presenting an article for publication in the Russian Digital Libraries Journal (RDLJ), the authors automatically give consent to grant a limited license to use the materials of the Kazan (Volga) Federal University (KFU) (of course, only if the article is accepted for publication). This means that KFU has the right to publish an article in the next issue of the journal (on the website or in printed form), as well as to reprint this article in the archives of RDLJ CDs or to include in a particular information system or database, produced by KFU.
All copyrighted materials are placed in RDLJ with the consent of the authors. In the event that any of the authors have objected to its publication of materials on this site, the material can be removed, subject to notification to the Editor in writing.
Documents published in RDLJ are protected by copyright and all rights are reserved by the authors. Authors independently monitor compliance with their rights to reproduce or translate their papers published in the journal. If the material is published in RDLJ, reprinted with permission by another publisher or translated into another language, a reference to the original publication.
By submitting an article for publication in RDLJ, authors should take into account that the publication on the Internet, on the one hand, provide unique opportunities for access to their content, but on the other hand, are a new form of information exchange in the global information society where authors and publishers is not always provided with protection against unauthorized copying or other use of materials protected by copyright.
RDLJ is copyrighted. When using materials from the log must indicate the URL: index.phtml page = elbib / rus / journal?. Any change, addition or editing of the author's text are not allowed. Copying individual fragments of articles from the journal is allowed for distribute, remix, adapt, and build upon article, even commercially, as long as they credit that article for the original creation.
Request for the right to reproduce or use any of the materials published in RDLJ should be addressed to the Editor-in-Chief A.M. Elizarov at the following address: amelizarov@gmail.com.
The publishers of RDLJ is not responsible for the view, set out in the published opinion articles.
We suggest the authors of articles downloaded from this page, sign it and send it to the journal publisher's address by e-mail scan copyright agreements on the transfer of non-exclusive rights to use the work.