Experimental Error, Kurtosis, Activity Cliffs, and Methodology: What Limits the Predictivity of Quantitative Structure-Activity Relationship Models?

Robert P Sheridan; Prabha Karnachi; Matthew Tudor; Yuting Xu; Andy Liaw; Falgun Shah; Alan C Cheng; Elizabeth Joshi; Meir Glick; Juan Alvarez

doi:10.1021/acs.jcim.9b01067

Experimental Error, Kurtosis, Activity Cliffs, and Methodology: What Limits the Predictivity of Quantitative Structure-Activity Relationship Models?

J Chem Inf Model. 2020 Apr 27;60(4):1969-1982. doi: 10.1021/acs.jcim.9b01067. Epub 2020 Apr 15.

Authors

Affiliations

¹ Computational and Structural Chemistry, Merck & Company Inc., Kenilworth, New Jersey 07033, United States.
² Computational and Structural Chemistry, Merck & Company Inc., West Point, Pennsylvania 19486, United States.
³ Biometrics Research, Merck & Company Inc., Rahway, New Jersey 07065, United States.
⁴ Computational and Structural Chemistry, Merck & Company Inc., South San Francisco, California 94080, United States.
⁵ Pharmacokinetics, Pharmacodynamics & Drug Metabolism, Merck & Company Inc., West Point, Pennsylvania 19486, United States.
⁶ Computational and Structural Chemistry, Merck & Company Inc., Boston, Massachusetts 02115, United States.

PMID: 32207612
DOI: 10.1021/acs.jcim.9b01067

Abstract

Given a particular descriptor/method combination, some quantitative structure-activity relationship (QSAR) datasets are very predictive by random-split cross-validation while others are not. Recent literature in modelability suggests that the limiting issue for predictivity is in the data, not the QSAR methodology, and the limits are due to activity cliffs. Here, we investigate, on in-house data, the relative usefulness of experimental error, distribution of the activities, and activity cliff metrics in determining how predictive a dataset is likely to be. We include unmodified in-house datasets, datasets that should be perfectly predictive based only on the chemical structure, datasets where the distribution of activities is manipulated, and datasets that include a known amount of added noise. We find that activity cliff metrics determine predictivity better than the other metrics we investigated, whatever the type of dataset, consistent with the modelability literature. However, such metrics cannot distinguish real activity cliffs due to large uncertainties in the activities. We also show that a number of modern QSAR methods, and some alternative descriptors, are equally bad at predicting the activities of compounds on activity cliffs, consistent with the assumptions behind "modelability." Finally, we relate time-split predictivity with random-split predictivity and show that different coverages of chemical space are at least as important as uncertainty in activity and/or activity cliffs in limiting predictivity.

MeSH terms

Quantitative Structure-Activity Relationship*
Scientific Experimental Error*
Structure-Activity Relationship
Uncertainty