Module 3: Examining Relationships: Quantitative Data
Module 3: Examining Relationships: Quantitative Data
Assessing the Fit of a Line (1 of 4)
Assessing the Fit of a Line (1 of 4)
Learning OUTCOMES
- Use residuals, standard error, and r2 to assess the fit of a linear model.
Introduction
Let’s take a moment to summarize what we have done up to this point in Examining Relationships: Quantitative Data. Our goal from the beginning was to examine the relationship between two quantitative variables. We started by looking at scatterplots to see if we could see any pattern between the explanatory and response variables. We focused early in the course on identifying those cases that were linear in form. At the same time, we assessed how strong the linear relationship was on the basis of visual inspection. As is our usual strategy, we turned from graphs to numeric measures, and in particular, we developed the correlation coefficient, r, as a measure of the strength of the linear relationship we observed in the graph.
Once we established that there was a linear relationship between explanatory and response variables, the next step was to find a line that fit the data: the best-fit line. Here we used the least-squares method to find the regression line. Finally, we used the equation of the regression line to predict the value of the response variable for a given value of the explanatory variable.
How Good Is the Best-Fit Line?
Now that we have a mathematical model (the least-squares regression line) that we can use to make predictions, we want to know: How good are these predictions, and how can we measure the error in a prediction?
Example
Highway Sign Visibility
Let’s begin our investigation by predicting the maximum distance that an 18-year-old driver can read a highway sign and then determining the error in our prediction.
We use the regression line equation:
Distance = 576 + (–3 * Age)
To predict the distance for an 18-year-old driver, we plug Age = 18 into the equation.
Predicted distance = 576 + (–3 * 18) = 522
Our prediction is that 522 feet is the maximum distance at which an 18-year-old driver can read a highway sign. Now let’s compare our prediction to the actual data for the 18-year-old driver: (18, 510).
The error in our prediction is 510 – 522 = –12.
This tells us that the actual distance for the 18-year-old driver is 12 feet closer than the prediction. In other words, our prediction is too large. It overestimates the actual distance by 12 feet.
So in general, we have Observed data value – Predicted value = Error.
If we use (x, y) to represent a typical data point and ŷ to represent the predicted value (obtained by using the regression equation), then we have
observed[latex]y -[/latex] predicted[latex]y =[/latex] error
[latex]y - \hat{y} =[/latex] error
Try It
Using this table showing “observed” and “predicted” distances for some drivers, find the following:
https://assessments.lumenlearning.com/assessments/3497
Now let’s look at the error from a different perspective. We can think of the error a