LLM-driven Semantic Guidance for Enhancing Reinforcement Learning-based Autonomous Driving through Reward Engineering
강화학습 기반 자율주행 성능 향상을 위한 보상 설계 기반 LLM 주도 의미적 가이던스
Ahmad Mouri Zadeh Khaki, Kyunghwan Choi*
Key Figure

  • Access the paper
  • 제어로봇시스템학회 (ICROS), 2026 accepted [📃 Full-Text]
    • Abstract
    • This Proximal Policy Optimization (PPO) is widely used for autonomous driving but often faces sparse rewards and unsafe exploration, which are critical issues in safety-sensitive scenarios. We propose LLM-PPO Driver, a framework that improves PPO by using a Large Language Model (LLM) as a source of prior driving knowledge during training. The LLM supports learning through reward shaping and imitation learning, without participating in real-time control. Experiments in the Gym highway-v0 environment show that both methods outperform baseline PPO, with imitation learning achieving the best results. These findings demonstrate that LLM-guided prior knowledge can improve safety, task success, and learning efficiency in autonomous driving.