3.5 Cautions about Regression
While regression is a very useful and powerful tool, it is also commonly misused. The main things we need to keep in mind when interpreting our results are:
-
- Linearity
- Correlation (or association) does not imply causation
- Extrapolation
- Outliers and influential points
Linearity
Remember, it is always important to plot a scatter diagram first. If the scatter plot indicates that there is a linear relationship between the variables, then it is reasonable to use the methods we are discussing.
Correlation Does Not Imply Causation
Even when there is an apparent linear relationship and a reasonable value of r, there can always be confounding or lurking variables at work. Be wary of spurious correlations and make sure the connection you are making makes sense!
There are also often situations where it may not be clear which variables affect each other. Does lack of sleep lead to higher stress levels, or do high stress levels lead to lack of sleep? Which came first, the chicken or the egg? Sometimes these may not be answerable, but at least we are able to show an association there.
Extrapolation
Remember, it is always important to plot a scatter diagram first. If the plot suggests the variables have a linear relationship, then it is reasonable to use a best-fit line to make predictions for y given x within the domain of x values in the sample data, though not necessarily for x values outside that domain. The process of predicting inside of the observed x values observed in the data is called interpolation. The process of predicting outside of the observed x values observed in the data is called extrapolation.
Recall our example from the previous section. You could use the line to predict the final exam score for a student who earned a grade of 73 on the third exam. You should NOT use the line to predict the final exam score for a student who earned a grade of 50 on the third exam, because 50 is not within the domain of the x values in the sample data, which are between 65 and 75.
To understand just how unreliable the prediction can be outside of the observed x values observed in the data, make the substitution x = 90 in the equation:
The final exam score is predicted to be 261.19. The largest a final exam score could be is 100.
Outliers and Influential Points
In some scatter plots, there may be points that stick out. How they stick out is important in the bivariate case. Outliers are points that stick out from the rest of the group in a single variable. We can identify outliers in univariate data using the fence rules.
In addition to outliers, a sample may contain one or more points that are called influential points. Influential points are observed data points that do not follow the trend of the rest of the data. These points could have a big effect on the slope of the regression line calculation. To begin to identify an influential point, you can remove it from the dataset and see if the slope of the regression line