ORCID
0000-0002-1476-113X (Chakraborty)
Document Type
Article
Publication Date
2026
DOI
10.3389/fpubh.2026.1823225
Publication Title
Frontiers in Public Health
Volume
14
Pages
1823225
Abstract
Background: Asthma is one of the most prominent chronic diseases in children and one of the most challenging ailments to diagnose in infants and preschoolers in the United States. Predictive models can be instrumental in improving early diagnosis, personalized treatment strategies, and disease progression. By utilizing nationalized data, this study focuses on building and comparing high-performing analytical predictive models based on the relevant risk factors and identifying the most influential predictors.
Methods: We analyzed cross-sectional BRFSS Asthma Call-Back Survey data (2011-2020; N = 9,813) and randomly split participants into training and testing sets. An XGBoost model (hyperparameters tuned via grid search) was developed and compared with SVM, random forest, LASSO, and GBM using accuracy, AUC, precision, and recall. Calibration was evaluated with reliability plots and improved using Platt scaling and isotonic regression. Predictor contributions were examined using variable-importance (VIP) and Shapley Additive Explanations (SHAP) plot.
Results: Of the five predictive models, the XGBoost was found to be the best performing model with AUC: 0.95, followed by random forest (AUC: 0.9345), GBM (AUC: 0.9341), SVM (AUC 0.9304), and LASSO (AUC 0.88); however, the random forest model was found to have the highest sensitivity (0.9786), and hence preferred for initial screening of asthma. On the independent test set, calibration (10-bin reliability curves; Brier/ECE/intercept-slope) improved most with isotonic regression, specifically for Random Forest (ECE 0.0158 to 0.0086; intercept -0.174 to -0.010), whereas Platt scaling often worsened calibration, with AUC remaining largely stable across models (AUC ≈ 0.92-0.95). The top two contributing predictors were overnight hospitalization visits and time since the last asthma medication, accounting for 24.62 and 20.92%, respectively, of the asthma status, from the VIP.
Conclusion: The analytical methodology of model development was found to be instrumental in the discovery of behavioral health-risk knowledge and to visualize the significance of predictive modeling from a multidimensional behavioral health survey. These insights can be instrumental in predicting different types of chronic lung diseases affecting people of all ages and can be useful for clinicians to diagnose asthma at an early stage, allowing for early intervention and proactive management.
Rights
© 2026 Chakraborty and Bashar.
This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0). The use, distribution or reproduction in other forums is permitted, provided the original authors and the copyright owners are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
Data Availability
Article states: "The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author/s."
Original Publication Citation
Chakraborty, A., & Bashar, A. (2026). The crucial role of machine learning models in predicting current childhood asthma: Model comparison, calibration, and SHAP-based interpretation. Frontiers in Public Health, 14, Article 1823225. https://doi.org/10.3389/fpubh.2026.1823225
Repository Citation
Chakraborty, A., & Bashar, A. (2026). The crucial role of machine learning models in predicting current childhood asthma: Model comparison, calibration, and SHAP-based interpretation. Frontiers in Public Health, 14, Article 1823225. https://doi.org/10.3389/fpubh.2026.1823225
Included in
Artificial Intelligence and Robotics Commons, Diagnosis Commons, Disease Modeling Commons, Theory and Algorithms Commons