Amino acid sequence encodes protein abundance shaped by protein stability at reduced synthesis cost

Filip Buric; Sandra Viknander; Xiaozhi Fu; Oliver Lemke; Oriol Gracia Carmona; Jan Zrimec; Lukasz Szyrwiel; Michael Mülleder; Markus Ralser; Aleksej Zelezniak

doi:10.1002/pro.5239

Amino acid sequence encodes protein abundance shaped by protein stability at reduced synthesis cost

Protein Sci. 2025 Jan;34(1):e5239. doi: 10.1002/pro.5239.

Authors

Filip Buric¹, Sandra Viknander¹, Xiaozhi Fu¹, Oliver Lemke², Oriol Gracia Carmona^{3

4}, Jan Zrimec^{1

5}, Lukasz Szyrwiel², Michael Mülleder⁶, Markus Ralser², Aleksej Zelezniak^{1

3

7}

Affiliations

¹ Department of Biology and Biological Engineering, Chalmers University of Technology, Gothenburg, Sweden.
² Department of Biochemistry, Charité - Universitätsmedizin Berlin, Berlin, Germany.
³ Randall Centre for Cell & Molecular Biophysics, King's College London, London, UK.
⁴ Institute of Structural and Molecular Biology, University College London, London, UK.
⁵ Department of Biotechnology and Systems Biology, National Institute of Biology, Ljubljana, Slovenia.
⁶ Core Facility High Throughput Mass Spectrometry, Charité - Universitätsmedizin Berlin, Berlin, Germany.
⁷ Institute of Biotechnology, Life Sciences Centre, Vilnius University, Vilnius, Lithuania.

Abstract

Understanding what drives protein abundance is essential to biology, medicine, and biotechnology. Driven by evolutionary selection, an amino acid sequence is tailored to meet the required abundance of a proteome, underscoring the intricate relationship between sequence and functional demand. Yet, the specific role of amino acid sequences in determining proteome abundance remains elusive. Here we show that the amino acid sequence alone encodes over half of protein abundance variation across all domains of life, ranging from bacteria to mouse and human. With an attempt to go beyond predictions, we trained a manageable-size Transformer model to interpret latent factors predictive of protein abundances. Intuitively, the model's attention focused on the protein's structural features linked to stability and metabolic costs related to protein synthesis. To probe these relationships, we introduce MGEM (Mutation Guided by an Embedded Manifold), a methodology for guiding protein abundance through sequence modifications. We find that mutations which increase predicted abundance have significantly altered protein polarity and hydrophobicity, underscoring a connection between protein structural features and abundance. Through molecular dynamics simulations we revealed that abundance-enhancing mutations possibly contribute to protein thermostability by increasing rigidity, which occurs at a lower synthesis cost.

Keywords: deep learning; explainable machine learning; language models; molecular dynamics; protein engineering; protein expression; protein sequence; protein stability; proteome.

MeSH terms

Amino Acid Sequence
Animals
Humans
Hydrophobic and Hydrophilic Interactions
Mice
Molecular Dynamics Simulation
Mutation
Protein Biosynthesis
Protein Stability*
Proteins / chemistry
Proteins / genetics
Proteins / metabolism

Substances

Proteins

Abstract

MeSH terms

Substances

Grants and funding