7.3 The Sampling Distribution of the Sample Proportion
7.3 The Sampling Distribution of the Sample Proportion
We have now talked at length about the basics of inference on the mean of quantitative data. What if the variable we are interested in is categorical? We cannot calculate means, variances, and the like for categorical data. However, we can count the number of individuals that have a characteristic we are interested in and divide by the total number in our population to get the population proportion (p).
Suppose a poll suggested the US President’s approval rating is 45%. We would consider 45% to be a point estimate of the approval rating we might see if we collected responses from the entire population. This entire-population response proportion is generally referred to as the parameter of interest. When the parameter is a proportion, it is often denoted by p, and we often refer to the sample proportion as ˆp (pronounced “p-hat”). Unless we collect responses from every individual in the population, p remains unknown, and we use ˆp as our estimate of p. The difference we observe from the poll versus the parameter is called the error in the estimate.
Understanding the Variability of a Proportion
Suppose we know the proportion of American adults who support the expansion of solar energy is p = 0.88, which is our parameter of interest. If we were to take a poll of 1000 American adults on this topic, the estimate would not be perfect, but how close might we expect the sample proportion in the poll would be to 88%? We want to understand, how does the sample proportion, ˆp, behave when the true population proportion is 0.88. We can simulate responses we would get from a simple random sample of 1000 American adults, which is only possible because we know the actual support for expanding solar energy is 0.88. Here’s how we might go about constructing such a simulation:
- There were about 250 million American adults in 2018. On 250 million pieces of paper, write “support” on 88% of them and “not” on the other 12%.
- Mix up the pieces of paper and pull out 1000 pieces to represent our sample of 1000 American adults.
- Compute the fraction of the sample that say “support”.
Any volunteers to conduct this simulation? Probably not. Running this simulation with 250 million pieces of paper would be time-consuming and very costly, but we can simulate it using technology. In this simulation, one sample gave a point estimate of ˆp1 = 0.894. We know the population proportion for the simulation was p = 0.88, so we know the estimate had an error of 0.894 − 0.88 = +0.014. One simulation isn’t enough to get a great sense of the distribution of estimates we might expect in the simulation, so we should run more simulations. In a second simulation, we get ˆp2 = 0.885, which has an error of +0.005. In another, ˆp3 = 0.878 for an error of -0.002. And in another, an estimate of ˆp4 = 0.859 with an error of -0.021. With the help of a computer, we’ve run the simulation 10,000 times and created a histogram of the results from