This paper proposes a constrained optimization-based reinforcement learning (RL) method to enhance learning stability for discrete-time nonlinear systems. To address the oscillatory convergence and hyperparameter sensitivity of conventional Adaptive Dynamic Programming (ADP), the Bellman Optimality Equation (BOE) is reformulated as a constrained optimization problem. By treating the control policy and critic weights as simultaneous decision variables and enforcing the BOE as an equality constraint, the framework ensures consistent optimality throughout the learning process. Online update laws are derived from the Karush-Kuhn-Tucker (KKT) conditions within a Lagrangian framework to seek the optimal solution in real time. The proposed approach demonstrates improved convergence and robustness over standard RL-based control methods.