7.4 Inference for a Proportion
If we are working with categorical data our parameter of interest is often the population proportion, p. The point estimate for p is = where x is the number of successes and n is the sample size. It is also sometimes denoted as . We saw previously that if we meet conditions, np ≥ 10 and n(1 − p) ≥ 10, we can apply the central limit theorem and assume:
~N
How do you know you are dealing with a proportion problem? First, the underlying distribution is a binomial distribution. This will be categorical data with no mention of a mean or average. If X is a binomial random variable, then X ~ B(n, p) where n is the number of trials and p is the probability of a success.
Hypothesis Tests for p
When you perform a hypothesis test of a single population proportion p, the steps are exactly the same as what we have seen before, however we will calculate our Test Statistic differently. When conducting a test for p, our hypotheses will look as follows:
- Ho: p = p0
- Ha: p (<,>,≠) p0
Recall, the general form of a test statistic is:
For the normal distribution of proportions, the z-score formula is as follows:
If ~Nthen the z-score formula is:
Intuitively, you might think we use this as our test statistic but remember two things:
- We do not actually know p
- In a hypothesis test we begin by assuming the null is true
Sure to these facts, we substitute in p0 for p in the standard error which gives us:
We then can find a p-value and make our decision as normal
Example
Joon believes that 50% of first-time brides in the United States are younger than their grooms. She performs a hypothesis test to determine if the percentage is the same or different from 50%. Joon samples 100 first-time brides and 53 reply that they are younger than their grooms. For the hypothesis test, she uses a 1% level of significance.
You Try It
Confidence Intervals for p
During an election year, we see articles in the newspaper that state confidence intervals in terms of proportions or percentages. For example, a poll for a particular candidate running for president might show that the candidate has 40% of the vote within three percentage points (if the sample is large enough). Often, election polls are calculated with 95% confidence, so, the pollsters would be 95% confident that the true proportion of voters who favored the candidate would be between 0.37 and 0.43: (0.40 – 0.03,0.40 + 0.03).
Investors in the stock market are interested in the true proportion of stocks that go up and down each week. Businesses that sell personal computers are interested in the proportion of households in the United States that own personal computers. Confidence intervals can be calculated for the true proportion of stocks that go up or down each week and for the true proportion of households in the United States that own personal computers.
Constructing Confidence Intervals for p
The structure of, and procedure to find the confidence interval for a proportion is similar to that f