An iterative learning algorithm is presented for continuous-time linear–quadratic optimal control problems where the system is externally symmetric with unknown dynamics. Both finite-horizon and infinite-horizon problems are considered. It is shown that the proposed algorithm is globally convergent to the optimal solution and has some advantages over adaptive dynamic programming, including unbiased performance under noisy measurements, relatively low computational burden, and no requirement for exploration noise. Numerical experiments show the effectiveness of the results.
QC 20251019