Module 14: Multiple and Logistic Regression
Model Selection
Barbara Illowsky & OpenStax et al.
In practice, the model that includes all available explanatory variables is often referred to as the full model. The full model may not be the best model, and if it isn’t, we want to identify a smaller model that is preferable.
Identifying variables in the model that may not be helpful
Adjusted R2 describes the strength of a model fit, and it is a useful tool for evaluating which predictors are adding value to the model, where adding value means they are (likely) improving the accuracy in predicting future outcomes.
Let’s consider two models, which are shown in Tables 1 and 2. The first table summarizes the full model since it includes all predictors, while the second does not include the duration variable.
df = 136
| Table 1. The fit for the full regression model, including the adjusted R2. | ||||
|---|---|---|---|---|
| Estimate | Std. Error | t value | Pr( >|t|) | |
| (Intercept) | 36.2110 | 1.5140 | 23.92 | 0.0000 |
| cond_new | 5.1306 | 1.0511 | 4.88 | 0.0000 |
| stock_photo | 1.0803 | 1.0568 | 1.02 | 0.3085 |
| duration | –0.0268 | 0.1904 | –0.14 | 0.8882 |
| wheels | 7.2852 | 0.5547 | 13.13 | 0.0000 |
| R2adj = 0.7108 |
| Table 2. The fit for the regression model for predictors cond_new, stock_photo, and wheels. | ||||
|---|---|---|---|---|
| Estimate | Std. Error | t value | Pr( >|t|) | |
| (Intercept) | 36.0483 | 0.9745 | 36.99 | 0.0000 |
| cond_new | 5.1763 | 0.9961 | 5.20 | 0.0000 |
| stock_photo | 1.1177 | 1.0192 | 1.10 | 0.2747 |
| wheels | 7.2984 | 0.5448 | 13.40 | 0.0000 |
| R2adj = 0.7128 |
Example
Which of the two models is better?
Solution:
We compare the adjusted R2 of each model to determine which to choose. Since the first model has an R2adj smaller than the R2adj of the second model, we prefer the second model to the first.
Will the model without duration be better than the model with duration? We cannot know for sure, but based on the adjusted R2, this is our best assessment.
Two model selection strategies
Two common strategies for adding or removing variables in a multiple regression model are called backward elimination and forward selection. These techniques are often referred to as stepwise model selection strategies, because they add or delete one variable at a time as they “step” through the candidate predictors.
Backward elimination starts with the model that includes all potential predictor variables. Variables are eliminated one-at-a-time from the model until we cannot improve the adjusted R2. The strategy within each elimination step is to eliminate the variable that leads to the largest improvement in adjusted R2.
Example
Results corresponding to the full model for the mario kart data are shown in Table 8.6. How should we proceed under the backward elimination strategy?
Solution:
Our baseline adjusted R2 from the full model is R2adj = 0.7108, and we need to determine whether dropping a predictor will improve the adjuste