Sample-Efficient Constrained Reinforcement Learning with General Parameterization

Mondal, Washim Uddin; Aggarwal, Vaneet

Computer Science > Machine Learning

arXiv:2405.10624v1 (cs)

[Submitted on 17 May 2024 (this version), latest version 23 Jul 2024 (v2)]

Title:Sample-Efficient Constrained Reinforcement Learning with General Parameterization

Authors:Washim Uddin Mondal, Vaneet Aggarwal

View PDF HTML (experimental)

Abstract:We consider a constrained Markov Decision Problem (CMDP) where the goal of an agent is to maximize the expected discounted sum of rewards over an infinite horizon while ensuring that the expected discounted sum of costs exceeds a certain threshold. Building on the idea of momentum-based acceleration, we develop the Primal-Dual Accelerated Natural Policy Gradient (PD-ANPG) algorithm that guarantees an $\epsilon$ global optimality gap and $\epsilon$ constraint violation with $\mathcal{O}(\epsilon^{-3})$ sample complexity. This improves the state-of-the-art sample complexity in CMDP by a factor of $\mathcal{O}(\epsilon^{-1})$.

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2405.10624 [cs.LG]
	(or arXiv:2405.10624v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2405.10624

Submission history

From: Washim Mondal [view email]
[v1] Fri, 17 May 2024 08:39:05 UTC (39 KB)
[v2] Tue, 23 Jul 2024 12:04:52 UTC (39 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.LG

< prev | next >

new | recent | 2024-05

Change to browse by:

cs
cs.AI

References & Citations

export BibTeX citation

Computer Science > Machine Learning

Title:Sample-Efficient Constrained Reinforcement Learning with General Parameterization

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Sample-Efficient Constrained Reinforcement Learning with General Parameterization

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators