2.4 Describing Quantitative Distributions
Consider the following exercise.
Your classmates write down the average time (in hours, to the nearest half-hour) they sleep per night and then create a simple dot plot of the data:
How would you interpret or explain this distribution? Where do your data appear to cluster? How might you interpret the clustering? If you did the same example in an English class with the same number of students, do you think the results would be the same? Why or why not?
The questions above ask you to analyze and interpret your data. It isn’t enough to just make graphs, we must be able to interpret the information with a critical eye.
Key Aspects of Quantitative Data
When describing a quantitative distribution we want to note at least four: the shape of the distribution, the presence of outliers, the center, and the spread. A helpful acronym for remembering this is SOCS:
Shape is the main characteristic we can determine by looking at a graph. We are often able to identify potential outliers visually as well. Center and spread can be roughly gauged visually, but there are also numerical calculations for them, which will be discussed in the following sections.
Shape
Shape is the first thing we should note since it will often dictate how to proceed with the rest of our analysis. We have already seen that most of our graphical methods can give us an idea of the shape of a distribution, but the best choice in most situations is a properly formatted histogram.
Symmetry vs. Skewness
Most of us are familiar with datasets that show roughly equal tails trailing off equally in both directions, which would be described as symmetric.
Consider the following alternative:
The figure above suggests that most loans have rates under 15%, while only a handful of loans have rates above 20%. When data trails off to the right in this way and has a longer right tail, the shape is said to be right-skewed. Datasets with the reverse characteristic—a long, thinner tail to the left—are said to be left-skewed. We also say that such a distribution has a long left tail.
Modality
In addition to looking at whether a distribution is skewed or symmetric, histograms can be used to identify the modality of a distribution. A mode is represented by a prominent peak in the distribution. There is only one prominent peak in the histogram of loan amounts. The definition of mode sometimes taught in math classes is the value with the most occurrences in the dataset. However, for many real-world datasets, it is common to have no observations with the same value in a dataset, making this definition impractical in data analysis. The figure below shows histograms that have one, two, or three prominent peaks.
Such distributions are called unimodal, bimodal, and multimodal, respectively. Any distribution with more than two prominent peaks is called multimodal. Notice that there was one prominent peak in the unimodal distribution, with a second less prominent peak that was not