Table 1.
Overview of plasmode simulation scenarios reflecting varying exposure and outcome prevalences based on National Health and Nutrition Examination Survey (NHANES) Data Cycles (2013–2018).
Fig 1.
Comparison of Bias Across Different Methods in hdPS Analysis.
Fig 2.
Comparison of Coverage Probability Across Different Methods in hdPS Analysis.
Fig 3.
Comparison of Risk Differences (RD) and Odds Ratios (OR) with 95% confidence intervals for different methods used to evaluate the association between obesity and diabetes risk. The analysis is based on data from the National Health and Nutrition Examination Survey (NHANES) for the years 2013–2018. Methods are arranged by the number of variables used in the models.
Table 2.
Comparison of variable overlap of selected proxies across different methods used to evaluate the association between obesity and diabetes from the National Health and Nutrition Examination Survey (NHANES) for the years 2013–2018. Diagonal entries show the total number of proxies selected by each method, and off-diagonal entries represent the count of shared variables between method pairs. Most methods share a moderate number of proxies (typically 50-60 percent of the smaller set), indicating partial agreement in variable selection. Higher overlap is observed between closely related methods (e.g., LASSO and Elastic Net, or Hybrid with Bross/LASSO), while methods like XGBoost and Genetic Algorithm show lower overlap with others, reflecting divergent selection behavior in high-dimensional settings.
Table 3.
Comparison of the count and percentage of proxy variables selected by each methods in common with that by the Bross formula-based high-dimensional propensity score to evaluate the association between obesity and diabetes from the National Health and Nutrition Examination Survey (NHANES) for the years 2013–2018. The Hybrid method (Bross + LASSO) shows perfect agreement with the Bross-based hdPS by design. Other methods demonstrate overlap rates ranging from 0.69 to 0.79, indicating moderate consistency in variable selection. XGBoost shows the highest overlap (79 percent) among the non-hybrid methods, while the Genetic Algorithm shows the lowest (69 percent), reflecting greater divergence in selected proxies. These overlap patterns highlight methodological differences in how each approach prioritizes covariates in high-dimensional settings.
Fig 4.
Computing time for the real-world analysis for each algorithm under consideration. The analysis is based on data from the National Health and Nutrition Examination Survey (NHANES) for the years 2013–2018.