An Introduction to Probabilistic Record Linkage with a Focus on Linkage Processing for WTC Registries

Int J Environ Res Public Health. 2020 Sep 22;17(18):6937. doi: 10.3390/ijerph17186937.

Abstract

Since its post-World War II inception, the science of record linkage has grown exponentially and is used across industrial, governmental, and academic agencies. The academic fields that rely on record linkage are diverse, ranging from history to public health to demography. In this paper, we introduce the different types of data linkage and give a historical context to their development. We then introduce the three types of underlying models for probabilistic record linkage: Fellegi-Sunter-based methods, machine learning methods, and Bayesian methods. Practical considerations, such as data standardization and privacy concerns, are then discussed. Finally, recommendations are given for organizations developing or maintaining record linkage programs, with an emphasis on organizations measuring long-term complications of disasters, such as 9/11.

Keywords: 9/11 health; data matching; disaster epidemiology; epidemiology; interagency cooperation; probabilistic record linkage; record linkage.

Publication types

  • Research Support, Non-U.S. Gov't
  • Research Support, U.S. Gov't, P.H.S.
  • Review

MeSH terms

  • Algorithms*
  • Bayes Theorem
  • Medical Record Linkage*
  • Privacy
  • Registries*