1

Model-free inverse reinforcement learning algorithms for continuous-time and discrete-time zero-sum games

Authors

Asl H.J.; Le A.V.; Minh B.V.; Elara M.R.

Publisher

Elsevier B.V.

Publication year
2026
Abstract

This paper presents data-driven, model-free inverse optimal control (IOC) algorithms, also known as inverse reinforcement learning (IRL), for estimating the cost functions of two-player zero-sum games with deterministic continuous-time (CT) and discrete-time (DT) dynamics. Addressing the underexplored area of IOC in zero-sum games, the proposed methods reduce the complexity inherent in existing bi-level approaches by estimating all cost function terms, including those penalizing inputs. The method partitions unknown parameters into two sets: one is updated iteratively using a gradient-based scheme driven by policy errors, while the other is estimated via the Hamilton-Jacobi-Bellman equation after convergence of the first set. This eliminates the need to solve a forward IOC problem at each iteration, reducing computational load and ensuring policy stability. For DT systems, the approach parallels the CT method but requires adapted gains to manage additional complexities. Numerical simulations demonstrate the effectiveness of the algorithms in accurately estimating cost functions for two-player zero-sum games.

Index
WoS
Journals
Neurocomputing