Variable selection based on the boosting algorithm in multilevel models

Document Type : Research Paper

Authors

1 Department of Statistics, SR.C., Islamic Azad University, Tehran, Iran.

2 Department of Statistics, Imam Khomeini International University, Qazvin, Iran.

Abstract

The existence of a natural hierarchy in all the multilevel problems are required for the proper exploration as specialized analytical ‎tools.‎ ‎Multilevel analytical tools provide the proper estimation of such type of problems‎. ‎The variance components models‎, ‎t‎he hierarchy linear models and random coefficient models are known as multilevel regression‎. ‎In this paper‎, ‎we consider random intercept and slope model and estimate parameters of model using the EM algorithm‎. ‎Also‎, ‎we propose a boosting algorithm based on the EM algorithm‎. ‎The performance of the proposed algorithm is studied and is compared to usual variable selection methods using simulation studies‎. ‎A real data analysis carried out to illustrate the procedures obtained theoretically.

Keywords

Main Subjects


[1] Aitkin, M. & Longford, N. (1986). Statistical modelling in school e ectiveness studies (with discussion). Journal of the Royal Statistical Society, 149, 1-43. https://doi.org/10.2307/2981882
[2] Barr, DJ. Levy, R. Scheepers, C. & Tily, HJ. (2013). Random e ects structure for con rmatory hypothesis testing: Keep it maximal. Journal of memory and language, 68(3), 255-278. https://doi.org/10.1016/j.jml.2012.11.001
[3] Bouwmeester, W. Twisk, JW. Kappen, TH. van Klei, WA. Moons, KG. & Vergouwe, Y. (2013). Prediction models for clustered data: comparison of a random intercept and standard regression model. BMC medical research methodology, 13, 1-10. https://doi.org/10.1186/1471-2288-13-19
[4] Buhlmann, P. (2006). Boosting for high-dimensional linear models. Annals of Statistics, 34(2), 559-583. https://doi.org/10.1214/009053606000000092
[5] Buhlmann, P. & Yu, B. (2003). Boosting with the L2 loss: Regression and classication. Journal of the American Statistical Association, 98(462), 324-339. https://doi.org/10.1198/016214503000125
[6] Chen, LP. & Qiu, B. (2023). Analysis of length-biased and partly interval-censored survival data with mismeasured covariates. Biometrics, 79, 3929-3940. https://doi.org/10.1111/biom.13898
[7] Chen, LP. (2024a). Variable selection and estimation for misclassi ed binary responses and multivariate error-prone predictors. Journal of Computational and Graphical Statistics, 33, 407-420. https://doi.org/10.1080/10618600.2023.2218428
[8] Chen, LP. (2024b). Accelerated failure time models with error-prone response and nonlinear covariates. Statistics and Computing, 34, 183. https://doi.org/10.1007/s11222-024-10491-9
[9] Choi, J. (2020). Nonnegative variance component estimation for mixed-e ects models. Communications for Statistical Applications and Methods, 27(5), 523-533. https://doi.org/10.29220/CSAM.2020.27.5.523
[10] Chu, J. Lee, TH. Ullah, A. & Wang, R. (2020). Boosting. Macroeconomic Forecasting in the Era of Big Data: Theory and Practice, 431-463. https://link.springer.com/book/10.1007/978-3-030-31150-6
[11] Fisher, RA. (1919). The correlation between relatives on the supposition of Mendelian inheritance. Earth and Environmental Science Transactions of the Royal Society of Edinburgh, 52(2), 399-433. https://doi.org/10.1017/S0080456800012163
[12] Friedman, JH. (2001). Greedy function approximation: a gradient boosting machine. Annals of statistics, 1189-1232. https://doi.org/10.1214/aos/1013203451
[13] Heiling, HM., Rashid, NU. Li, Q. Peng, XL., Yeh, JJ. & Ibrahim, JG. (2024). Ef- cient computation of high-dimensional penalized generalized linear mixed models by latent factor modeling of the random e ects. Biometrics, 80(1). https://doi.org/10.1093/biomtc/ujae016 
[14] Henderson, CR. Kempthorne, O. Searle, SR. & Von Krosigk, CM. (1959). The estimation of environmental and genetic trends from records subject to culling. Biometrics, 15(2), 192-218. https://www.cabidigitallibrary.org/doi/ull/10.5555/19590102115
[15] Longford, NT. (1987). A fast scoring algorithm for maximum likelihood estimation in unbalanced mixed models with nested random e ects. Biometrika, 74, 817-827. https://doi.org/10.2307/2336476
[16] McLean, RA., Sanders, WL. & Stroup, WW. (1991). A uni ed approach to mixed linear models. The American Statistician, 45(1), 54-64. https://doi.org/abs/10.1080/00031305.1991.10475767
[17] Mei, J. Zhang, Y. & Chen, M. (2022). A exible EM algorithm for mixture and nonlinear mixed-e ects models with complex random-e ects structures. Statistical Methods in Medical Research, 31(11), 2180{2198. https://doi.org/10.1177/09622802221098687
[18] Oberauer, K. (2022). The importance of random slopes in mixed models for Bayesian hypothesis testing. Psychological Science, 33(4), 648-665. 09567976211046884
[19] Raudenbush, SW. & Bryk, AS. (1986). A hierarchical model for studying school e ects. Sociology of Education, 12, 241-269. https://doi.org/10.2307/2112482
[20] Robinson, GK. (1991). That BLUP is a good thing: the estimation of random e ects. Statistical science, 15-32. https://doi.org/10.1214/ss/1177011926
[21] Schapire, R. (1990). The strength of weak learnability. Mach. Learn, 5, 197-227. https://doi.org/10.1007/BF00116037
[22] Schapire, R. (1999). Theoretical views of boosting and applications, in Proc. Comput. Learn. Theory, 1-10. https://doi.org/10.1007/3-540-46769-6
[23] Zakkour, A. Perret, C. & Slaoui, Y. (2023). Stochastic expectation maximization algorithm for linear mixed-e ects model with interactions in the presence of incomplete data. Entropy, 25(3), 473. https://doi.org/10.3390/e25030473

Articles in Press, Accepted Manuscript
Available Online from 10 February 2026
  • Receive Date: 10 June 2025
  • Revise Date: 09 December 2025
  • Accept Date: 10 February 2026