Missing data in a long food frequency questionnaire: are imputed zeroes correct?

Gary E Fraser; Ru Yan; Terry L Butler; Karen Jaceldo-Siegl; W Lawrence Beeson; Jacqueline Chan

doi:10.1097/EDE.0b013e31819642c4

Missing data in a long food frequency questionnaire: are imputed zeroes correct?

Epidemiology. 2009 Mar;20(2):289-94. doi: 10.1097/EDE.0b013e31819642c4.

Authors

Gary E Fraser¹, Ru Yan, Terry L Butler, Karen Jaceldo-Siegl, W Lawrence Beeson, Jacqueline Chan

Affiliation

¹ Department of Epidemiology and Biostatistics, Loma Linda University, Loma Linda, CA, USA. [email protected]

Abstract

Background: Missing data are a common problem in nutritional epidemiology. Little is known of the characteristics of these missing data, which makes it difficult to conduct appropriate imputation.

Methods: We telephoned, at random, 20% of subjects (n = 2091) from the Adventist Health Study-2 cohort who had any of 80 key variables missing from a dietary questionnaire. We were able to obtain responses for 92% of the missing variables.

Results: We found a consistent excess of "zero" intakes in the filled-in data that were initially missing. However, for frequently consumed foods, most missing data were not zero, and these were usually not distinguishable from a random sample of nonzero data. Older, black, and less-well-educated subjects had more missing data. Missing data are more likely to be true zeroes in older subjects and those with more missing data. Zero imputation for missing data may create little bias except for more frequently consumed foods, in which case, zero imputation will be suboptimal if there is more than 5%-10% missing.

Conclusions: Although some missing data represent true zeroes, much of it does not, and data are usually not missing at random. Automatic imputation of zeroes for missing data will usually be incorrect, although there is [corrected] little bias unless the foods are frequently consumed. Certain identifiable subgroups have greater amounts of missing data, and require greater care in making imputations.

Publication types

Research Support, N.I.H., Extramural

MeSH terms

Aged
Aged, 80 and over
Algorithms
Bias*
Diet / statistics & numerical data
Diet Surveys*
Epidemiologic Studies
Female
Humans
Linear Models
Male
Middle Aged
Nutrition Assessment
Surveys and Questionnaires* / standards
United States

Abstract

Publication types

MeSH terms

Grants and funding