Abstract
The US Food and Drug Administration recently approved sickle cell disease (SCD) gene therapies that require a contemporaneous cohort to compare their outcomes with those of individuals who did not receive gene therapy. Previously, we developed an automated algorithm to create a contemporaneous cohort of children and adults with SCD using electronic health record data. We tested the hypothesis that the updated Vanderbilt University Medical Center (VUMC) algorithm can identify individuals with SCD with sensitivity and specificity of >95%. In the VUMC health care system, we identified a cohort of 33 141 adults and children of primarily African ancestry who underwent β-globin gene sequencing, with a 1.4% prevalence of SCD (473/33 141). For SCD, the performance of the VUMC algorithm using International Classification of Diseases (ICD) and laboratory data showed a sensitivity of 97.6% (462/473), specificity of 99.9% (32 663/32 668), positive predictive value of 98.9% (462/467), and negative predictive value of 99.9% (32 663/32 674). Three SCD phenotypes account for 98.5% of the individuals in the cohort. The cohort had sensitivities of 96.8% (272/281), 93.7% (136/145), and 85% (34/40) for hemoglobin SS (HbSS)/Hb Sβ0-thalassemia (HbSβ0), HbSC, and HbSβ+, respectively. The cohort had a specificity of 99.9% across all 3 phenotypes: HbSS/HbSβ0, HbSC, and HbSβ+, respectively. We have developed an updated algorithm using ICD and laboratory data that accurately identifies individuals with SCD within a large health care system, and distinguishes HbSS/HbSβ0, HbSC, and HbSβ+.