2.7 Measures of Spread
We’re now to the final key aspect of our acronym, SOCS:
- Shape
- Outliers
- Center
- Spread
A complement to the center of a distribution is the variation, variability, or spread of the data. In some data sets, the values are concentrated closely, while in others the are more spread out. Some rough measures of spread we have already seen are the range and IQR. The most common measure of spread is the standard deviation.
Similar to measures of center, the shape of the distribution and presence of extreme values can dictate what the most appropriate measure of spread is to describe the distribution.
The Interquartile Range
Recall the Interquartile Range (IQR):
IQR = Q3 – Q1.
In addition to helping us establish our fences and identify outliers, the IQR indicates the spread of the middle half or the middle 50% of the data. The IQR can be used as a somewhat rough but very robust measure of spread when outliers may be present. It is often used alongside the median to describe the center and spread of skewed distributions.
Simply showing the five number summary or a Box Plot can be a good way to get all of the information for a skewed dataset in one place
The Standard Deviation
The standard deviation is a measure of spread that measures how spread out values are from their mean. It is essentially the “average” deviation, or distance of each observation from the mean.
Not only does it provide a numerical measure of the overall amount of variation in a data set, it can also be used for other purposes
The lower case letter s represents the sample standard deviation and the lower case greek letter σ (sigma) represents the population standard deviation.
By extension, s² represents the sample variance and the lower case greek letter σ² represents the population variance. The variance is useful
The standard deviation is small when the data are all concentrated close to the mean, exhibiting little variation or spread. The standard deviation is larger when the data values are more spread out from the mean, exhibiting more variation. It must always greater than or equal to zero.
Suppose that we are studying the amount of time customers wait in line at the checkout at supermarket A and supermarket B. the average wait time at both supermarkets is five minutes. At supermarket A, the standard deviation for the wait time is two minutes; at supermarket B the standard deviation for the wait time is four minutes.
Because supermarket B has a higher standard deviation, we know that there is more variation in the wait times at supermarket B. Overall, wait times at supermarket B are more spread out from the average; wait times at supermarket A are more concentrated near the average.
Calculating the Standard Deviation
The procedure to calculate the standard deviation can be tedious and depends on whether the data are from the entire population or a sample. The calculations are similar, but not identical.
If x is a number, then the difference “x – mean”