Machine learning-assisted amidase-catalytic enantioselectivity prediction and rational design of variants for improving enantioselectivity

Zi-Lin Li; Shuxin Pei; Ziying Chen; Teng-Yu Huang; Xu-Dong Wang; Lin Shen; Xuebo Chen; Qi-Qiang Wang; De-Xian Wang; Yu-Fei Ao

doi:10.1038/s41467-024-53048-0

Machine learning-assisted amidase-catalytic enantioselectivity prediction and rational design of variants for improving enantioselectivity

Nat Commun. 2024 Oct 10;15(1):8778. doi: 10.1038/s41467-024-53048-0.

Authors

Zi-Lin Li^#^{1

2}, Shuxin Pei^#³, Ziying Chen³, Teng-Yu Huang^{1

2}, Xu-Dong Wang¹, Lin Shen^{4

5}, Xuebo Chen^{6

7

8}, Qi-Qiang Wang^{1

2}, De-Xian Wang^{1

2}, Yu-Fei Ao^{9

10}

Affiliations

¹ Beijing National Laboratory for Molecular Sciences, CAS Key Laboratory of Molecular Recognition and Function, Institute of Chemistry, Chinese Academy of Sciences, Beijing, China.
² University of Chinese Academy of Sciences, Beijing, China.
³ Key Laboratory of Theoretical and Computational Photochemistry of Ministry of Education, College of Chemistry, Beijing Normal University, Beijing, China.
⁴ Key Laboratory of Theoretical and Computational Photochemistry of Ministry of Education, College of Chemistry, Beijing Normal University, Beijing, China. [email protected].
⁵ Yantai-Jingshi Institute of Material Genome Engineering, Yantai, China. [email protected].
⁶ Key Laboratory of Theoretical and Computational Photochemistry of Ministry of Education, College of Chemistry, Beijing Normal University, Beijing, China. [email protected].
⁷ Yantai-Jingshi Institute of Material Genome Engineering, Yantai, China. [email protected].
⁸ Shandong Laboratory of Yantai Advanced Materials and Green Manufacturing, Yantai, China. [email protected].
⁹ Beijing National Laboratory for Molecular Sciences, CAS Key Laboratory of Molecular Recognition and Function, Institute of Chemistry, Chinese Academy of Sciences, Beijing, China. [email protected].
¹⁰ University of Chinese Academy of Sciences, Beijing, China. [email protected].

^# Contributed equally.

Abstract

Biocatalysis is an attractive approach for the synthesis of chiral pharmaceuticals and fine chemicals, but assessing and/or improving the enantioselectivity of biocatalyst towards target substrates is often time and resource intensive. Although machine learning has been used to reveal the underlying relationship between protein sequences and biocatalytic enantioselectivity, the establishment of substrate fitness space is usually disregarded by chemists and is still a challenge. Using 240 datasets collected in our previous works, we adopt chemistry and geometry descriptors and build random forest classification models for predicting the enantioselectivity of amidase towards new substrates. We further propose a heuristic strategy based on these models, by which the rational protein engineering can be efficiently performed to synthesize chiral compounds with higher ee values, and the optimized variant results in a 53-fold higher E-value comparing to the wild-type amidase. This data-driven methodology is expected to broaden the application of machine learning in biocatalysis research.

MeSH terms

Amidohydrolases* / chemistry
Amidohydrolases* / genetics
Amidohydrolases* / metabolism
Biocatalysis*
Machine Learning*
Models, Molecular
Protein Engineering* / methods
Stereoisomerism
Substrate Specificity

Substances

Amidohydrolases
amidase

Abstract

MeSH terms

Substances

Grants and funding