Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' Rule

Kurita, Shuhei; Cho, Kyunghyun

Computer Science > Computation and Language

arXiv:2009.07783 (cs)

[Submitted on 16 Sep 2020 (v1), last revised 8 Oct 2020 (this version, v3)]

Title:Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' Rule

Authors:Shuhei Kurita, Kyunghyun Cho

View PDF

Abstract:Vision-and-language navigation (VLN) is a task in which an agent is embodied in a realistic 3D environment and follows an instruction to reach the goal node. While most of the previous studies have built and investigated a discriminative approach, we notice that there are in fact two possible approaches to building such a VLN agent: discriminative \textit{and} generative. In this paper, we design and investigate a generative language-grounded policy which uses a language model to compute the distribution over all possible instructions i.e. all possible sequences of vocabulary tokens given action and the transition history. In experiments, we show that the proposed generative approach outperforms the discriminative approach in the Room-2-Room (R2R) and Room-4-Room (R4R) datasets, especially in the unseen environments. We further show that the combination of the generative and discriminative policies achieves close to the state-of-the art results in the R2R dataset, demonstrating that the generative and discriminative policies capture the different aspects of VLN.

Comments:	13 pages, 8 figures
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2009.07783 [cs.CL]
	(or arXiv:2009.07783v3 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2009.07783

Submission history

From: Shuhei Kurita [view email]
[v1] Wed, 16 Sep 2020 16:23:17 UTC (2,886 KB)
[v2] Wed, 23 Sep 2020 18:57:09 UTC (2,887 KB)
[v3] Thu, 8 Oct 2020 17:16:49 UTC (2,891 KB)

Computer Science > Computation and Language

Title:Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' Rule

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' Rule

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators