Learning to Skip for Language Modeling

Zeng, Dewen; Du, Nan; Wang, Tao; Xu, Yuanzhong; Lei, Tao; Chen, Zhifeng; Cui, Claire

Computer Science > Computation and Language

arXiv:2311.15436 (cs)

[Submitted on 26 Nov 2023]

Title:Learning to Skip for Language Modeling

Authors:Dewen Zeng, Nan Du, Tao Wang, Yuanzhong Xu, Tao Lei, Zhifeng Chen, Claire Cui

View PDF

Abstract:Overparameterized large-scale language models have impressive generalization performance of in-context few-shot learning. However, most language models allocate the same amount of parameters or computation to each token, disregarding the complexity or importance of the input data. We argue that in language model pretraining, a variable amount of computation should be assigned to different tokens, and this can be efficiently achieved via a simple routing mechanism. Different from conventional early stopping techniques where tokens can early exit at only early layers, we propose a more general method that dynamically skips the execution of a layer (or module) for any input token with a binary router. In our extensive evaluation across 24 NLP tasks, we demonstrate that the proposed method can significantly improve the 1-shot performance compared to other competitive baselines only at mild extra cost for inference.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2311.15436 [cs.CL]
	(or arXiv:2311.15436v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2311.15436

Submission history

From: Tao Lei [view email]
[v1] Sun, 26 Nov 2023 21:45:53 UTC (881 KB)

Computer Science > Computation and Language

Title:Learning to Skip for Language Modeling

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Learning to Skip for Language Modeling

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators