Analyzing longitudinal electronic health records data with clinically informative visiting process: possible choices and comparisons
Abstract
Analyzing longitudinal electronic health records (EHR) data is challenging due to clinically informative observation processes, where the timing and frequency of patient visits often depend on their underlying health condition. Traditional longitudinal data analysis rarely considers observation times as stochastic or models them as functions of covariates. In this work, we evaluate the impact of informative visiting processes on two common statistical tasks: (1) estimating the effect of an exposure on a longitudinal biomarker, and (2) assessing the effect of longitudinal biomarkers on a time-to-event outcome, such as disease diagnosis. The methods we consider range from using simple summaries of the observed longitudinal data, imputation, inversely weighting by the estimated intensity of the visit process, and joint modeling. To our knowledge, this is the most comprehensive evaluation of inferential methods specifically tailored for EHR-like visiting processes. For the first task, we improve certain methods around 18 times faster, making them more suitable for large-scale EHR data analysis. For the second task, where no method currently accounts for the visiting process in joint models of longitudinal and time-to-event data, we propose incorporating the historical number of visits to adjust for informative visiting. Using data from the longitudinal biobank at the University of Michigan Health System, we investigate two case studies: 1) the association between genetic variants and lab markers with repeated measures (known as the LabWAS); and 2) the association between cardiometabolic health markers and time to hypertension diagnosis. We show how accounting for informative visiting processes affects the analysis results. We develop an R package CIMPLE (Clinically Informative Missingness handled through Probabilities, Likelihood, and Estimating equations) that integrates all these methods.
- Publication:
-
arXiv e-prints
- Pub Date:
- October 2024
- DOI:
- 10.48550/arXiv.2410.13113
- arXiv:
- arXiv:2410.13113
- Bibcode:
- 2024arXiv241013113D
- Keywords:
-
- Statistics - Methodology;
- Statistics - Applications