Key Moments

Stanford AA274A Principles of Robotic Autonomy | Autumn 2019 | Advanced Trajectory Optimization

Stanford OnlineStanford Online
Education6 min read78 min video
Oct 9, 2026|1,321 views|5
Save to Pod
TL;DR

Advanced trajectory optimization uses indirect methods derived from optimal control theory, by first defining necessary optimality conditions and then solving them, to find cost-minimizing paths for robots.

Key Insights

1

Trajectory planning can be formulated as an optimal control problem, aiming to find a control input (U) and a state trajectory (X) that minimize a cost function, subject to system dynamics.

2

Indirect methods for optimal control involve deriving and solving the necessary conditions for optimality, often by constructing a Hamiltonian function similar to the Lagrangian in constrained optimization.

3

The necessary optimality conditions for optimal control problems result in a set of coupled differential equations (state and costate equations) and algebraic equations, forming a two-point boundary value problem (TPBVP).

4

Solving TPBVPs, which have boundary conditions at both the initial and final times, is more complex than initial value problems and requires specialized solvers like BVP solvers in Python or MATLAB.

5

When the final time (TF) is not fixed, time-rescaling techniques are used to transform the problem into a standard TPBVP with a fixed time interval (e.g., 0 to 1), allowing existing solvers to be applied.

6

Direct methods discretize the problem first, converting it into a finite-dimensional nonlinear programming problem, while indirect methods first derive optimality conditions and then discretize.

Formulating trajectory planning as optimal control

The lecture introduces advanced trajectory optimization by framing it as an optimal control problem, a more rigorous approach than differential flatness. The goal is to find not just a feasible path but one that optimizes a specific cost function, such as minimizing time, control effort, or a combination thereof. The problem is defined by system dynamics (differential equations relating state X and input U), input constraints (e.g., maximum force or acceleration), and a cost function comprising an integral term (running cost) and a terminal term (final cost). The primary task is to find a control input history U(t) and the corresponding state trajectory X(t) that minimize the total cost, where the state trajectory is governed by the system's differential equations and the initial state is typically given. This differs from finite-dimensional optimization as it involves optimizing over functions, an infinite-dimensional decision space.

Indirect methods: Optimality conditions first

Indirect methods for optimal control follow a strategy of first deriving the necessary conditions for optimality and then solving them. This approach is analogous to unconstrained optimization where setting the gradient to zero provides necessary conditions. For constrained optimization problems, like those with equality constraints, the concept of a Lagrangian function is introduced. The Lagrangian combines the original cost function with penalty terms for constraint violations, weighted by Lagrange multipliers. For optimal control, the equivalent of the Lagrangian is the Hamiltonian function. The necessary conditions for optimality in optimal control are expressed as a set of coupled differential equations, known as state equations (the original system dynamics), costate equations (derived from the Hamiltonian), and algebraic equations relating inputs to costates. Solving these conditions yields candidate optimal trajectories.

The Hamiltonian and costate equations

The Hamiltonian for an optimal control problem is defined as the running cost plus the dot product of the costate vector (P) and the system's dynamics (A*X*U*T). The costate variables (P) play a role analogous to Lagrange multipliers in constrained optimization. The necessary conditions for optimality require that the time derivative of the state (X dot) equals the partial derivative of the Hamiltonian with respect to the costate (P), and the time derivative of the costate (P dot) equals the negative partial derivative of the Hamiltonian with respect to the state (X). Additionally, the partial derivative of the Hamiltonian with respect to the control input (U) must be zero, which can often be used to solve for the optimal control input U in terms of the costates. This system of equations forms a two-point boundary value problem (TPBVP).

Solving two-point boundary value problems

The set of optimality conditions derived for optimal control problems typically results in a TPBVP. This means that boundary conditions are specified at both the initial time (e.g., initial state) and the final time (e.g., final state, or conditions related to the terminal cost and time). Unlike initial value problems (IVPs) where all conditions are at the start, TPBVPs are significantly more complex to solve numerically. Specialized solvers, such as BVP solvers in Python or MATLAB (like BVP4C), are required. These solvers iteratively adjust the solution to satisfy conditions at both ends of the time interval. The lecture demonstrates the setup for a BVP solver using a simple example, requiring the definition of the differential equations, boundary condition residuals, and an initial guess.

Handling free final time

A common scenario in optimal control is when the final time (TF) is not fixed but is itself a variable to be optimized. When TF is free, the standard TPBVP formulation needs adjustment. A key technique is time rescaling, where the original time variable T is mapped to a new time variable tau in a fixed interval (e.g., 0 to 1) by dividing by TF. This transforms the problem into a standard TPBVP with a fixed time horizon. An additional state variable representing TF is introduced with trivial dynamics (its derivative is zero). The main equation derived from the optimality conditions, which relates the Hamiltonian, terminal cost, and time, is used to derive an additional boundary condition that allows the solver to determine the optimal free final time.

Example: Minimum time for a double integrator

A practical example illustrates the application of indirect methods to a double integrator system. The objective is to drive a particle from an initial state (position=10, velocity=0) to a final state (position=0, velocity=0) in minimum time, while also minimizing control effort (U^2). The problem is converted into a first-order system, and the Hamiltonian is constructed. The necessary optimality conditions yield a system of differential equations for the states (X1, X2) and costates (P1, P2), along with an algebraic equation for the control input U. Since the final time is free, time rescaling is applied, introducing an extra state variable (R) for TF. The resulting system is a TPBVP with five differential equations and five boundary conditions, solvable using a BVP solver. The results show that as the weight on time cost increases, the final time decreases, and numerical solutions closely match analytical ones.

Relationship between direct and indirect methods

The lecture concludes by discussing the relationship between direct and indirect methods for trajectory optimization. Direct methods discretize the continuous-time problem first, transforming it into a finite-dimensional nonlinear programming problem, which is then solved. Indirect methods, as detailed in the lecture, first derive the necessary conditions for optimality in the continuous-time domain and then solve them, often involving discretization at a later stage. The key insight is that as the discretization step size (time step) in direct methods becomes infinitesimally small, the optimality conditions derived from the discretized problem converge to the optimality conditions obtained by indirect methods. Therefore, both approaches are fundamentally linked and should yield consistent results for the same problem, especially with fine discretization. This relationship serves as a consistency check between the two methodologies.

Common Questions

Optimal control in trajectory planning involves finding a control input (like steering or acceleration) over time that minimizes a specific cost function, such as time, control effort, or a combination of both, while satisfying the system's dynamics and constraints.

Topics

Mentioned in this video

More from Stanford Online

View all 148 summaries

Ask anything from this episode.

Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.

Get Started Free