MEMORY-AUGMENTED REINFORCEMENT LEARNING FOR UAV NAVIGATION USING PPO-LSTM
DOI:
https://doi.org/10.30572/2018/KJE/170327Keywords:
Autonomous drone navigation, obstacle avoidance, PPO, LSTM, Partial observabilityAbstract
The problems of partial observability and sensor shortage pose a significant challenge for autonomous Unmanned Aerial Vehicles (UAVs) as they prove to be challenging for conventional Deep Reinforcement Learning (DRL) methods to undertake well under such conditions. In this paper, a memory-augmented Proximal Policy Optimization (PPO) model extended using a Long Short-Term Memory (LSTM) network is proposed as a solution to such challenges. The observation space is constructed from 2D LiDAR and Inertial Measurement Unit (IMU) data to sense simultaneously external observation and internal state of motion, whereas the action space consists of continuous velocity commands. A shaped reward function is optimized for encouraging safe target approaching, obstacle avoidance, and convergence speed. Experimental outcomes show that the PPO-LSTM described herein achieves smoother paths, more robust reward convergence, and a much lower rate of collision than regular PPO. It also generalizes to new environments with movable obstacles. Qualitatively, the success rate increased from 64.5% to 83.9%, collision frequency reduced by over 70%, and path efficiency increased from 0.60 to 0.85, without suffering from unstable training behavior
Downloads
References
Al-Hsnawy, T. and Al-Ghanimi, A. (2024), “A REVIEW OF CONTROL METHODS FOR QUADROTOR UAV”, Kufa Journal of Engineering, University of Kufa, Vol. 15 No. 4, pp. 98–124, doi: 10.30572/2018/KJE/150408.
Hanover, D., Loquercio, A., Bauersfeld, L., Romero, A., Penicka, R., Song, Y., Cioffi, G., et al. (2023), “Autonomous Drone Racing: A Survey”, doi: 10.1109/TRO.2024.3400838.
Huang, X., Wang, W., Ji, Z. and Cheng, B. (2023), “Representation Enhancement-Based Proximal Policy Optimization for UAV Path Planning and Obstacle Avoidance”, Hindawi, Vol. 2023, p. 15, doi: 10.1155/2023/6654130.
Ibrahim, S., Mostafa, M., Jnadi, A., Salloum, H. and Osinenko, P. (2024), “Comprehensive Overview of Reward Engineering and Shaping in Advancing Reinforcement Learning Applications”, IEEE Access, Institute of Electrical and Electronics Engineers Inc., doi: 10.1109/ACCESS.2024.3504735.
Ismael, H.H., Jasim, M.A., Aidi Sharif, M. and Jasim, F.Z. (2025), “ARTIFICIAL INTELLIGENCE IN ROBOTIC MANIPULATORS: EXPLORING OBJECT DETECTION AND GRASPING INNOVATIONS”, Kufa Journal of Engineering, University of Kufa, Vol. 16 No. 1, pp. 136–159, doi: 10.30572/2018/KJE/160109.
Jiawei, X., Xufang, Z., Zhong, L. and Qingtao, X. (2023), “LSTM-DPPO based deep reinforcement learning controller for path following optimization of unmanned surface vehicle”, Journal of Systems Engineering and Electronics, Beijing Institute of Aerospace Information, Vol. 34 No. 5, pp. 1343–1358, doi: 10.23919/JSEE.2023.000113.
Kalidas, A.P., Joshua, C.J., Md, A.Q., Basheer, S., Mohan, S. and Sakri, S. (2023), “Deep Reinforcement Learning for Vision-Based Navigation of UAVs in Avoiding Stationary and Mobile Obstacles”, Drones, doi: 10.3390/drones7040245.
Liu, J., Yan, Y., Yang, Y. and Li, J. (2024), “An Improved Artificial Potential Field UAV Path Planning Algorithm Guided by RRT Under Environment-Aware Modeling: Theory and Simulation”, IEEE Access, Vol. 12, pp. 12080–12097, doi: 10.1109/ACCESS.2024.3355275.
Mansour, H.S., Valizadeh, M., Abdulaal, A.H. and Amirani, M.C. (2025), “A NOVEL DEEP 2D-CNN MODEL FOR ECG-BASED ARRHYTHMIA DIAGNOSIS WITH SELECTIVE ATTENTION MECHANISM AND CWT INTEGRATION”, Kufa Journal of Engineering, University of Kufa, Vol. 16 No. 2, pp. 423–444, doi: 10.30572/2018/KJE/160225.
Pendyala, A., Atamna, A. and Glasmachers, T. (2024), “Solving a Real-World Optimization Problem Using Proximal Policy Optimization with Curriculum Learning and Reward Engineering”, doi: https://doi.org/10.48550/arXiv.2404.02577.
Rodriguez-Ramos, A., Sampedro, C., Bavle, H., de la Puente, P. and Campoy, P. (2019a), “A Deep Reinforcement Learning Strategy for UAV Autonomous Landing on a Moving Platform”, Journal of Intelligent and Robotic Systems: Theory and Applications, Springer Netherlands, Vol. 93 No. 1–2, pp. 351–366, doi: 10.1007/s10846-018-0891-8.
Rodriguez-Ramos, A., Sampedro, C., Bavle, H., de la Puente, P. and Campoy, P. (2019b), “A Deep Reinforcement Learning Strategy for UAV Autonomous Landing on a Moving Platform”, Journal of Intelligent & Robotic Systems, IEEE, Vol. 93 No. 1–2, pp. 351–366, doi: 10.1007/s10846-018-0891-8.
Schulman, J., Levine, S., Moritz, P., Jordan, M.I. and Abbeel, P. (2015), “Trust Region Policy Optimization”. https://doi.org/10.48550/arXiv.1502.054771.
Schulman, J., Wolski, F., Dhariwal, P., Radford, A. and Openai, O.K. (2017), Proximal Policy Optimization Algorithms. https://doi.org/10.48550/arXiv.1707.06347.
Sheng, Y., Liu, H., Li, J. and Han, Q. (2024), “UAV Autonomous Navigation Based on Deep Reinforcement Learning in Highly Dynamic and High-Density Environments”, Drones, Multidisciplinary Digital Publishing Institute (MDPI), Vol. 8 No. 9, doi: 10.3390/drones8090516.
Wang, X., Baldi, S., Feng, X., Wu, C., Xie, H. and De Schutter, B. (2023), “A Fixed-Wing UAV Formation Algorithm Based on Vector Field Guidance”, IEEE Transactions on Automation Science and Engineering, Vol. 20 No. 1, pp. 179–192, doi: 10.1109/TASE.2022.3144672.
Wisniewski, M., Chatzithanos, P., Guo, W. and Tsourdos, A. (2024), Benchmarking Deep Reinforcement Learning for Navigation in Denied Sensor Environments, doi: https://doi.org/10.48550/arXiv.2410.14616.
Downloads
Published
Issue
Section
Categories
License
Copyright (c) 2026 Maryam Allawi Haddad, Dhayaa Raissan Khudher

This work is licensed under a Creative Commons Attribution 4.0 International License.












