Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation

Min, Do June; Perez-Rosas, Veronica; Resnicow, Kenneth; Mihalcea, Rada

Computer Science > Computation and Language

arXiv:2403.13578 (cs)

[Submitted on 20 Mar 2024]

Title:Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation

Authors:Do June Min, Veronica Perez-Rosas, Kenneth Resnicow, Rada Mihalcea

View PDF HTML (experimental)

Abstract:In this paper, we study the problem of multi-reward reinforcement learning to jointly optimize for multiple text qualities for natural language generation. We focus on the task of counselor reflection generation, where we optimize the generators to simultaneously improve the fluency, coherence, and reflection quality of generated counselor responses. We introduce two novel bandit methods, DynaOpt and C-DynaOpt, which rely on the broad strategy of combining rewards into a single value and optimizing them simultaneously. Specifically, we employ non-contextual and contextual multi-arm bandits to dynamically adjust multiple reward weights during training. Through automatic and manual evaluations, we show that our proposed techniques, DynaOpt and C-DynaOpt, outperform existing naive and bandit baselines, showcasing their potential for enhancing language models.

Subjects:	Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2403.13578 [cs.CL]
	(or arXiv:2403.13578v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2403.13578

Submission history

From: Do June Min [view email]
[v1] Wed, 20 Mar 2024 13:24:41 UTC (639 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CL

< prev | next >

new | recent | 2024-03

Change to browse by:

cs
cs.LG

References & Citations

export BibTeX citation

Computer Science > Computation and Language

Title:Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators