Module 9: Inference for Two Proportions
Putting It Together: Inference for Two Proportions
Putting It Together: Inference for Two Proportions
Let’s Summarize
In Inference for Two Proportions, we learned two inference procedures to draw conclusions about a difference between two population proportions (or about a treatment effect): (1) a confidence interval when our goal is to estimate the difference and (2) a hypothesis test when our goal is to test a claim about the difference. Both types of inference are based on the sampling distribution.
Distribution of the Differences in Sample Proportions
In the section “Distribution of Differences in Sample Proportions,” we learned about the sampling distribution of differences between sample proportions.
We used simulation to observe the behavior of sample differences when we select random samples from two populations. Every simulation began with an assumption about the difference between the two population proportions. From the simulated sampling distribution, we could determine if a sample difference observed in the data was likely or unlikely. A data result that is unlikely to occur in the sampling distribution provides evidence that our original assumption about the difference in the population proportions is probably incorrect. This logic is similar to the logic of hypothesis testing.
Because samples vary, we do not expect sample differences to always equal the population difference. Every sample difference has some error. We used simulations to observe the amount of error we expected to see in sample differences. The “typical” amount of error in the sampling distribution connects to the margin of error in a confidence interval.
We also used simulations to describe the shape, center, and spread of the sampling distribution. Later we developed a mathematical model for the sampling distribution with formulas for the mean of the sample differences and the standard deviation of the sample differences. We call this standard deviation the standard error because it represents an estimate for the average error we see in sample differences.
The mean of sample differences between sample proportions is equal to the difference between the population proportions, p1 − p2.
The standard error of differences between sample proportions is related to the population proportions and the sample sizes.
[latex]\sqrt{\frac{p_{1}(1-p_{1})}{n_{1}}+\frac{p_{2}(1-p_{2})}{n_{2}}}[/latex]
A normal model is a good fit for the sampling distribution of differences between sample proportions under certain conditions. We use a normal model if the counts of expected successes and failures are at least 10. For those who like formulas, this translates into saying the following four calculations must all be at least 10.
- n1p1
- n1(1 − p1)
- n2p2
- n2(1 − p2)
Estimating the Difference between Two Population Proportions
In the section “Estimate the Difference between Population Proportions,” we learned how to calculate a confidence interva