Address lists from commercial vendors, enhanced with socio-demographic information, are widely used as the basis of address sampling frames, particularly when sampling subpopulations associated with the available socio-demographics is of interest. However, they come with missing data, requiring imputation.
The Study of Mindful Aging: Relationships and Thinking (SMART) is a nationally representative push-to-Web survey that targets adults aged 40 and older with an additional aim of oversampling Afro-Latino, Asian Indian, Chinese, Korean, and Vietnamese adults to examine the relationship between cognitive health and social relations. For SMART, a list of 635,373 addresses from 320 Census tracts was purchased with accompanying socio-demographic information. Age was missing for 129,450 addresses (20.4%), and race/ethnicity was missing for 161,743 addresses (25.5%).
We used the Light Gradient-Boosting Machine (LightGBM) algorithm to impute missing age and race/ethnicity. LightGBM is well suited for large-scale data containing a large number of both numeric and categorical predictors that may have nonlinear relationships with the variables being imputed. Our data included over 300 variables from the commercial vendors at the address level as well as the American Community Survey aggregated at the tract level.
Commercial data on household size and length of residence were important variables for imputing age, while important predictors differed across race/ethnicity categories. For example, language from commercial data was a strong predictor of granular Asian subgroups, whereas the geographic information (e.g., Census tract) was important for predicting subgroups such as White, Black, and Latino.
When applying a minimum predicted-probability threshold of 70%, age was imputed for 40,557 addresses and race/ethnicity for 61,506 addresses, reducing respective missing rates to 14% and 15.8%. SMART will use the age and race/ethnicity observed from the commercial data along with the imputed data for sampling and mailing invitation letters to prospective participants. We plan to report the results of these recruitment efforts in a future post.
Submitted by: Sunghee Lee, Jinseok Kim, Caleb Crouch, Felix Baez-Santiago
NIMLAS Topics: Case Prioritization
, Recruitment Across Contexts
, Differential Effectiveness of New Technologies
![]()