Key Moments
Stanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 10: Reachibility Analysis
Want to know something specific about what's covered?
We've already dissected every moment. Ask and we will deliver (with timestamps).
Key Moments
Continuous-time optimal control transitions to differential games where an adversarial 'player two' tries to minimize outcomes, requiring new mathematical tools like the Hamilton-Jacobi-Bellman equation.
Key Insights
The Hamilton-Jacobi-Bellman equation is the continuous-time analog of the Bellman equation for optimal control.
Differential games, involving an adversarial player attempting to minimize an objective, are formulated using a minimax objective.
The Hamilton-Jacobi-Bellman equation for differential games involves a min-max (or max-min) structure due to the adversarial player's influence.
Reachable set analysis, crucial for collision avoidance and goal achievement, can be computed using techniques derived from differential games and the Hamilton-Jacobi-Isaacs equation.
Linear Quadratic Regulator (LQR) problems in continuous time are solved by guessing a quadratic cost-to-go and solving an associated ordinary differential equation for its parameters.
Bridging discrete and continuous time optimal control
The lecture begins by summarizing the infinite horizon extension of stochastic dynamic programming, focusing on the fixed-point Bellman equation for both value and Q-functions. It then transitions to continuous-time optimal control, highlighting the need for a different mathematical framework. The continuous-time dynamic programming equations, known as the Hamilton-Jacobi-Bellman (HJB) equation, are introduced as the counterpart to the discrete-time Bellman equation. These equations are essential for solving problems in continuous time, which can offer cleaner policies and easier implementation in some scenarios. The course will explore LQR and reachability analysis using these continuous-time formulations.
Introducing differential games with an adversarial player
A key development is the introduction of differential games, where the system dynamics are influenced by a second player, referred to as 'player two' or 'nature,' whose objective is to minimize a cost (or maximize a negative reward) that player one (the controller) aims to maximize. This adversarial setting is motivated by the need for robust control policies that can handle worst-case disturbances, such as in collision avoidance for autonomous vehicles or aircraft. The formulation assumes player two acts non-anticipatorily, reacting to player one's past actions and current state, giving player two a slight advantage. This structure leads to a minimax optimization problem.
The Hamilton-Jacobi-Bellman equation for adversarial settings
The formulation of differential games leads to a generalized version of the HJB equation, often referred to as the Hamilton-Jacobi-Isaacs (HJI) equation. This equation incorporates a min-max or max-min structure reflecting the adversarial nature of the problem. Deriving this equation involves approximations of the continuous-time integral cost using Taylor expansions and a critical step where the order of optimization (min-max vs. max-min) is adjusted to consistently give the disturbance player an advantage. While mathematically complex, this equation provides a framework for finding optimal strategies in the presence of an adversary, with applications in ensuring survival or achieving goals under worst-case conditions.
Solving continuous-time LQR using the HJB equation
The lecture demonstrates how the HJB equation can be applied to the Linear Quadratic Regulator (LQR) problem in continuous time. The standard approach involves making an educated guess (ansatz) that the optimal cost-to-go is a quadratic function of the state. This guess is plugged into the HJB equation, leading to an ordinary differential equation for the unknown quadratic matrix. Solving this differential equation yields the parameters for the optimal cost-to-go and, consequently, the optimal control law, which remains a linear feedback of the state. This confirms that dynamic programming principles extend effectively to continuous-time LQR.
Reachability analysis for guaranteed outcomes
Reachability analysis, a crucial application of differential games, aims to compute backward reachable sets. These sets define initial states from which a system is guaranteed to reach a target set by a specific final time, despite adversarial actions. Two main scenarios are discussed: avoidance sets, where the target is an undesired region (e.g., hitting an obstacle), and goal-reaching sets, where the target is a desired region. For avoidance, one must ensure the initial state is outside the backward reachable set. For goal-reaching, the initial state must be within this set. The Hamilton-Jacobi-Isaacs framework is used to compute these sets, providing strong guarantees for planning and control.
Mentioned in This Episode
●Concepts
Common Questions
Value iteration focuses on iteratively updating the value function until it converges to the optimal value function. Policy iteration, on the other hand, alternates between evaluating a policy and then improving it until convergence to the optimal policy.
Topics
Mentioned in this video
The core equation used to define optimal control problems, extended here to infinite horizon and continuous time settings. It relates the value of a state to the immediate reward and the value of future states.
An algorithm for solving infinite horizon MDPs where the value function is iteratively updated until convergence. It involves repeatedly applying the Bellman equation.
An alternative algorithm for solving infinite horizon MDPs that iteratively improves a policy by first evaluating the current policy and then improving it.
The continuous-time analog of the Bellman equation, used for solving optimal control problems in continuous time. It is a partial differential equation.
A field in control theory concerned with determining the set of initial states from which a system can reach a target set, often under adversarial conditions.
Games involving multiple players whose actions are constrained by differential dynamics. They are used to model optimal control problems with adversarial disturbances.
A decision-making strategy used in game theory and decision theory that identifies the optimal move for a player assuming that their opponent also makes the optimal move, aiming to minimize the maximum possible loss or maximize the minimum possible gain.
The equation that arises from solving differential games, analogous to the Hamilton-Jacobi-Bellman equation for standard optimal control problems. It involves a minimax or maximin structure.
The set of states from which a system is guaranteed to reach a specific target set by a final time, often considering adversarial controls.
More from Stanford Online
View all 126 summaries
78 minStanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 9: Stochastic Dyn. Program
75 minAStanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 8: Nonlinearity
76 minStanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 7: Dynamic Programming
81 minStanford AA203 Optimal and Learning-Based Control | Spring 2026 | Lecture 6: Optimal Control
Ask anything from this episode.
Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.
Get Started Free