A Hybrid Approach to Autonomous Control of a Robotic Device That Combines Reinforcement Learning and Operator-Based Demonstration Learning

Main Article Content

Robert Tagirovich Iskhakov
Vlada Vladimirovna Kugurakova

Abstract

A reproducible hybrid approach to learning autonomous control of a robotic device has been proposed and experimentally verified, implemented using publicly available tools without specialized computing hardware. The approach combines learning from operator demonstrations, reinforcement learning, curriculum learning, and domain randomization into a single sequential framework. Based on a comparative analysis of five prototypes representing key classes of architectural solutions, applicable approaches for resource-constrained environments are identified. Three controlled experiments in Unity ML-Agents demonstrated the superiority of the hybrid combination over isolated approaches in terms of both learning speed and robustness to environmental changes. Limitations associated with catastrophic forgetting are outlined, and prospects for integrating large language models to generate synthetic training data are identified.

Article Details

How to Cite
Iskhakov, R. T., and V. V. Kugurakova. “A Hybrid Approach to Autonomous Control of a Robotic Device That Combines Reinforcement Learning and Operator-Based Demonstration Learning”. Russian Digital Libraries Journal, vol. 29, no. 5, Sept. 2026, pp. 1495-20, doi:10.26907/1562-5419-2026-29-5-1495-1520.

References

1. Kugurakova V.V., Khafizov M.R., Kadyrov S.A., Zykov E.Yu. Remote control of a robotic device using virtual reality // Software & Systems. 2022. Vol. 35, No. 3. P. 348–361. https://doi.org/10.15827/0236-235X.139.348-361
2. Kugurakova V.V. et al. VR-Telecontrol of Multi-Arm Devices: Problems, Hypotheses, Problem Statement // Russian Digital Libraries Journal. 2022. Vol. 25, No. 5. P. 441–471. https://doi.org/10.26907/1562-5419-2022-25-5-441-471
3. Zykova A.E., Zykov E.Yu., Kugurakova V.V. Primenenie tekhnologiy virtualnoy realnosti dlya udalennogo upravleniya robotizirovannymi ustroystvami // Itogovaya nauchno-prakticheskaya konferentsiya KFU. Kazan, 2023. P. 87. EDN GDBPQP.
4. Sergunin I.D., Zykov E.Yu. «RUSAK» – modulʹnoe reshenie sistemy upravleniya udalyonnymi ustrojstvami s integrirovannymi zashchishchyonnymi kanalami svyazi // Sostoyanie i perspektivy razvitiya sovremennoj nauki po napravleniyu «Robototekhnika». 2023. P. 265–271.
5. Kugurakova V.V., Iskhakov R.T. et al. Tochechnoe teleupravlenie dlya selʹskokhozyajstvennoj tekhniki, vypolnyayushchej zadachi na osnove II // International Forum KAZAN DIGITAL WEEK – 2025: conference papers, volume 1. 2025. P. 1964–1971.
6. Programma dlya sinkhronizatsii stereoizobrazheniya so stereopary, razmeshchyonnoj na udalyonnom dvigayushchemsya ustrojstve: Svidetelʹstvo o gos. reg. programmy dlya EVM № 2022684277 / Kugurakova V.V., Khafizov M.R. Gabdullina D.R.; KFU. Date of registration: 13.12.2022.
7. Programma adaptivnogo upravleniya tipovymi marshrutami i rutinnymi dejstviyami modeli robota-manipulyatora v virtualʹnoj srede: Svidetelʹstvo o gos. reg. programmy dlya ЭВМ № 2024689377 / Kugurakova V.V., Iskhakov R.T. et al.; KFU. Date of registration: 05.12.2022.
8. Iskhakov R.T. Razrabotka generatora bolʹshogo kolichestva variatsij dvizhenij robotizirovannogo ustrojstva na osnove bazovykh zapisannykh dejstvij operatora: vypusknaya kvalifikatsionnaya rabota: 09.03.04 «Programmnaya inzheneriya» / R.T. Iskhakov; Kazan (Volga Region) Federal University. Kazan, 2024. 41 p.
9. Guzaerov D.A., Sergunin I.D., Kugurakova V.V. Infrared Video-Based Object Recognition for Agricultural Research Field Uses // BIO Web of Conferences. EDP Sciences, 2024. Vol. 141. No. 01025. https://doi.org/10.1051/bioconf/202414101025
10. Zare M. et al. A survey of imitation learning: Algorithms, recent developments, and challenges // IEEE Transactions on Cybernetics. 2024. Vol. 54, No. 12. P. 7173–7186. https://doi.org/10.48550/arXiv.2309.02473
11. Arulkumaran K. et al. Deep Reinforcement Learning: A Brief Survey // IEEE Signal Processing Magazine. 2017. Vol. 34, No. 6. P. 26–38. https://doi.org/10.1109/MSP.2017.2743240
12. Kober J., Bagnell J.A., Peters J. Reinforcement Learning in Robotics: A Survey // The International Journal of Robotics Research. 2013. Vol. 32, No. 11. P. 1238–1274. https://doi.org/10.1177/0278364913495721
13. Chebotar Y. et al. Actionable models: Unsupervised offline reinforcement learning of robotic skills //arXiv preprint arXiv:2104.07749. 2021.
14. Bengio Y. et al. Curriculum Learning // Proceedings of ICML. 2009. P. 41–48. https://doi.org/10.1145/1553374.1553380
15. Graves A. et al. Automated Curriculum Learning for Neural Networks // Proceedings of ICML. 2017. P. 1311–1320. https://doi.org/10.48550/arXiv.1704.03003
16. Chukwurah N., Adebayo A.S., Ajayi O.O. Sim-to-real transfer in robotics: Addressing the gap between simulation and real-world performance //International Journal of Robotics and Simulation. 2024. Vol. 6, No. 1. P. 89–102. https://doi.org/10.54660/.IJFMR.2024.5.1.33-39
17. Tobin J. et al. Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World // IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). 2017. P. 23–30. https://doi.org/10.1109/IROS.2017.8202133
18. Pinto L. et al. Robust Adversarial Reinforcement Learning // Proceedings of the 34th International Conference on Machine Learning (ICML). 2017. https://doi.org/10.48550/arXiv.1703.02702
19. Abnar S. et al. Exploring the limits of large scale pre-training // arXiv preprint arXiv:2110.02095. 2021.
20. Tesla Optimus // teslarobots.ru. 2025. URL: https://teslarobots.ru/ (accessed at 05.02.26)
21. Brohan A. et al. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control // arXiv preprint arXiv:2307.15818. 2023.
22. Fu Z. et al. Humanplus: Humanoid shadowing and imitation from humans // arXiv preprint arXiv:2406.10454. 2024.
23. Xpeng Shares Achievements in Physical AI: Unveils XPeng VLA 2.0, Next-Gen Iron // xpeng.com. 2025. URL: https://www.xpeng.com/pressroom/news/ 019a56f54fe99a2a0a8d8a0282e402b7 (дата обращения: 04.03.2026).
24. Figure AI Opens New Helix Lab to Accelerate Learning for Its Humanoid Robots // mlq.ai. 2025. URL: https://mlq.ai/news/figure-ai-opens-new-helix-lab-to-accelerate-learning-for-its-humanoid-robots/ (accessed at 04.03.26)
25. Apanasevich I. et al. Green-VLA: Staged Vision-Language-Action Model for Generalist Robots //arXiv preprint arXiv:2602.00919. 2026.
26. Yu A., Palefsky-Smith R., Bedi R. Deep reinforcement learning for simulated autonomous vehicle control // Standford University Course Project Reports: Winter. 2016. P. 1–7. URL: https://cs231n.stanford.edu/reports/2016/pdfs/112_Report.pdf
27. Done C. et al. Reinforcement Learning-Driven Prosthetic Hand Actuation in a Virtual Environment Using Unity ML-Agents // Virtual Worlds. MDPI. 2025. Vol. 4, No. 4. P. 53. https://doi.org/10.3390/virtualworlds4040053
28. Kirkpatrick J. et al. Overcoming Catastrophic Forgetting in Neural Networks // Proceedings of the National Academy of Sciences. 2017. Vol. 114, No. 13. P. 3521–3526. https://doi.org/10.1073/pnas.1611835114
29. Fedus W. et al. Revisiting fundamentals of experience replay //International conference on machine learning. 2020. P. 3061–3071. https://doi.org/10.48550/arXiv.2007.06700
30. Caruana R. Multitask Learning // Machine Learning. 1997. Vol. 28. P. 41–70. https://doi.org/10.1023/A:1007379606734
31. Pateria S. et al. Hierarchical reinforcement learning: A comprehensive survey //ACM Computing Surveys (CSUR). 2021. Vol. 54, No. 5. P. 1–35. https://doi.org/10.1145/3453160


Most read articles by the same author(s)

1 2 3 > >>