Skip to main navigation Skip to search Skip to main content

Iteration Complexity for Robust CMDP for finite policy space

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

We consider the robust Constrained Markov decision (RCMDP) problem of learning a policy that will maximize the cumulative reward while satisfying a constraint against the worst possible stochastic model under the unknown uncertainty set. Such a problem is relevant when the simulated and the real environment differ. Such a problem poses significant additional challenges compared to the non-robust CMDP problem and the unconstrained robust MDP problem. We seek to characterize the number of iterations required to bound both the sub-optimality gap and the violations by at most ϵ. We observe that the primal-dual-based approaches that achieves iteration complexity bounds for non-robust CMDP cannot achieve the same in the robust CMDP case. We consider a modified problem where we consider the convex hull of the policy-spaces and the decision becomes the simplex over the policy space. We propose a primal-dual based approach and show that ϵ suboptimality gap and violation bound can be achieved after O(1/ϵ2) iterations. We also show that an extra-gradient based approach can achieve ϵ suboptimality gap and violation bound can be achieved after O(1/ϵ) iterations. This improves the existing bounds for robust CMDP problem OF O(1/ϵ4). Empirical evaluations show that our proposed approach can achieve feasible and yet optimal policies very fast.

Original languageEnglish (US)
Title of host publication2025 IEEE 64th Conference on Decision and Control, CDC 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages2713-2719
Number of pages7
ISBN (Electronic)9798331526276
DOIs
StatePublished - 2025
Event64th IEEE Conference on Decision and Control, CDC 2025 - Rio de Janeiro, Brazil
Duration: Dec 9 2025Dec 12 2025

Publication series

NameProceedings of the IEEE Conference on Decision and Control

Conference

Conference64th IEEE Conference on Decision and Control, CDC 2025
Country/TerritoryBrazil
CityRio de Janeiro
Period12/9/2512/12/25

ASJC Scopus subject areas

  • Control and Systems Engineering
  • Modeling and Simulation
  • Control and Optimization

Fingerprint

Dive into the research topics of 'Iteration Complexity for Robust CMDP for finite policy space'. Together they form a unique fingerprint.

Cite this