Null Distribution of the Test Statistic for Model Selection via Marginal Screening: Implications for Multivariate Regression Analysis
linear models, multiple comparisons, freedman's paradox, extreme value theory, order statistics, genetic risk score
Abstract
Marginal screening (MS) is the computationally simple and commonly used for the dimension reduction procedures. In it, a linear model is constructed for several top predictors, chosen according to the absolute value of marginal correlations with the dependent variable. Importantly, when k predictors out of m primary covariates are selected, the standard regression analysis may yield false-positive results if m >> k (Freedman's paradox). In this work, we provide analytical expressions describing null distribution of the test statistics for model selection via MS. Using the theory of order statistics, we show that under MS, the common F-statistic is distributed as a mean of k top variables out of m independent random variables having a 2 1 χ distribution. Based on this finding, we estimated critical p-values for multiple regression models after MS, comparisons with which of those obtained in real studies will help researchers to avoid falsepositive result. Analytical solutions obtained in the work are implemented in a free Excel spreadsheet program.
Downloads
How to Cite
References
Mohammad Ahsanullah, Valery Nevzorov, Mohammad Shakil (2013) An Introduction to Order Statistics.
K Alam, K Wallenius (1979) Distribution of a sum of order statistics. 6, 123-126.
Barry Arnold, N Balakrishnan, H Nagaraja (2008) A First Course in Order Statistics.
J Cohen (1988) Statistical Power Analysis for the Behavioral Sciences.
J Cohen (1992) A Power Primer. 112, 155-159.
George Diehr, Donald Hoflin (1974) Approximating the Distribution of the Sample R 2 in Best Subset Regressions. 16(2), 317.
J Fan, Q Shao, W Zhou (2017) Are Discoveries Spurious? Distributions of Maximum Spurious Correlations and Their Applications.
Dean Foster, Robert Stine (2006) Honest confidence intervals for the error variance in stepwise regression. 31(1-2), 89-102.
David Freedman (1983) A Note on Screening Regression Equations. 37(2), 152-155.
Christopher Genovese, Larry Wasserman (2009) Confidence sets for nonparametric wavelet regression. 33(2).
C Genovese, J Jin, L Wasserman, Z Yao (2012) Comparison of the lasso and marginal regression. 13, 2107-2143.
T Hastie, R Tibshirani (2003) Expression arrays and the n p >> problem.
J Lee, J Taylor (2014) Exact Post Model Selection Inference for Marginal Screening.
J Leek (2016) Everyday Ethics: Top 10 Ethical Considerations in Using Telepractice.
Paul Lukacs, Kenneth Burnham, David Anderson (2010) Model selection bias and Freedman's paradox. 62(1), 117-125.
Haikady Nagaraja (1980) Contributions to the theory of the selection differential and to order statistics.
H Nagaraja (1982) Some Nondegenerate Limit Laws for the Selection Differential. 10(4), 1306-1310.
H Nagaraja (1980) Order Statistics from Independent Exponential Random Variables and the Sum of the Top Order Statistics. 22, 49-53.
A Rubanovich, N Khromov-Borisov (2016) Genetic risk assessment of the joint effect of several genes: Critical appraisal. 52(7), 757-769.
David Salt, Subhash Ajmani, Ray Crichton, David Livingstone (2007) An Improved Approximation to the Estimation of the Critical F Values in Best Subset Regression. 47(1), 143-149.
Stephen Stigler (1973) The Asymptotic Distribution of the Trimmed Mean. 1(3), 472-477.
R Tibshirani, J Taylor, R Richard Lockhart, R Tibshirani (2016) Exact Post-Selection Inference for Sequential Regression Procedures. 111, 600-620.
(2016) Statistical Approaches to Gene X Environment Interactions for Complex Phenotypes.
N Wray, J Yang, B Hayes, N Wray, J Yang, B Hayes, A L Price, M Goddard, P Visscher (2013) Pitfalls of predicting complex traits from SNPs. 14, 507-515.
Published
2021-10-21
Issue
Section
License
Copyright (c) 2021 Authors and Global Journals Private Limited

This work is licensed under a Creative Commons Attribution 4.0 International License.