Healthcare utilization is a collider: an introduction to collider bias in EHR data reuse

J Am Med Inform Assoc. 2023 Apr 19;30(5):971-977. doi: 10.1093/jamia/ocad013.

Abstract

Objectives: Collider bias is a common threat to internal validity in clinical research but is rarely mentioned in informatics education or literature. Conditioning on a collider, which is a variable that is the shared causal descendant of an exposure and outcome, may result in spurious associations between the exposure and outcome. Our objective is to introduce readers to collider bias and its corollaries in the retrospective analysis of electronic health record (EHR) data.

Target audience: Collider bias is likely to arise in the reuse of EHR data, due to data-generating mechanisms and the nature of healthcare access and utilization in the United States. Therefore, this tutorial is aimed at informaticians and other EHR data consumers without a background in epidemiological methods or causal inference.

Scope: We focus specifically on problems that may arise from conditioning on forms of healthcare utilization, a common collider that is an implicit selection criterion when one reuses EHR data. Directed acyclic graphs (DAGs) are introduced as a tool for identifying potential sources of bias during study design and planning. References for additional resources on causal inference and DAG construction are provided.

Keywords: EHR data reuse; cohort selection; collider bias; directed acyclic graphs; real-world data.

Publication types

  • Research Support, N.I.H., Extramural
  • Research Support, Non-U.S. Gov't

MeSH terms

  • Bias
  • Confounding Factors, Epidemiologic
  • Epidemiologic Methods
  • Patient Acceptance of Health Care*
  • Retrospective Studies