Constrained Optimization Formulation of Bellman Optimality Equation for Online Reinforcement Learning
온라인 강화학습을 위한 벨만 최적 방정식의 제약 최적화 문제
Hyochan Lee, Kyunghwan Choi*
Key Figure

  • Access the paper
  • 제어로봇시스템학회 (ICROS), 2026 accepted [📃 Full-Text]
    • Abstract
    • This paper proposes a constrained optimization-based reinforcement learning (RL) method to enhance learning stability for discrete-time nonlinear systems. To address the oscillatory convergence and hyperparameter sensitivity of conventional Adaptive Dynamic Programming (ADP), the Bellman Optimality Equation (BOE) is reformulated as a constrained optimization problem. By treating the control policy and critic weights as simultaneous decision variables and enforcing the BOE as an equality constraint, the framework ensures consistent optimality throughout the learning process. Online update laws are derived from the Karush-Kuhn-Tucker (KKT) conditions within a Lagrangian framework to seek the optimal solution in real time. The proposed approach demonstrates improved convergence and robustness over standard RL-based control methods.