Chapter 2 Wrap Up
Concept Check
Section Reviews
2.1 Introduction to Descriptive Statistics and Frequency Tables
Descriptive statistics are ways of organizing summarizing and presenting data. There are two main types: visual and numerical. Usually we want to first examine a dataset visually then describe it numerically. Appropriate methods often depend on the type of data you are working with, however frequency tables are a quick easy way to organize any type of data.
2.2 Displaying and Describing Categorical Data
Two basic visual methods we have for displaying categorical statistics are:
- Pie charts
- Bar charts
When describing a categorical distribution we want to note:
- Mode
- Level of variability (diversity)
2.3 Displaying Quantitative Data
The following are common methods of displaying quantitative data
- Stem-and-leaf plots
- Dot plots
- Line graphs
- Histograms
- Frequency polygons
- Time series plots
Some work better to show certain aspects, or for different sample sizes than others.
2.4 Describing Quantitative Distributions
When describing a quantitative distribution we want to at least note 4 things: the shape of the distribution, the presence of outliers, the center, and the spread. A helpful acronym to remember this is SOCS:
- Shape – Can be identified visually, want to note symmetry or lack thereof (skewness) and modality
- Outliers – Extreme outliers can be seen visually
- Center – Central tendency can be estimated visually
- Spread – Dispersion can be estimated visually and roughly quantified with the range
2.5 Measures of Location and Outliers
The values that divide a rank-ordered set of data into 100 equal parts are called percentiles. Percentiles are used to compare and interpret data. For example, an observation at the 50th percentile would be greater than 50 percent of the other observations in the set.
Where:
- i = the ranking or position of a data value,
- k = the kth percentile,
- n = total number of data.
Expression for finding the percentile of a data value:
Where:
- x = the number of values counting from the bottom of the data list up to but not including the data value for which you want to find the percentile,
- y = the number of data values equal to the data value for which you want to find the percentile,
- n = total number of data
Quartiles divide data into quarters. The first quartile (Q1) is the 25th percentile, the second quartile (Q2 or median) is 50th percentile, and the third quartile (Q3) is the the 75th percentile.
The interquartile range, or IQR, is the range of the middle 50 percent of the data values. The IQR is found by subtracting Q1 from Q3, and can help determine outliers by using the following fence rules.
- Upper fence = Q3 + IQR(1.5)
- Upper fence =Q1 – IQR(1.5)
Box plots are a type of graph that can help visually organize data. To graph a box plot the following data points must be calculated: the minimum value, the first quartile, the median, the third quartile, and the maximum value. Once the bo