26 7.4 Inference for a Proportion
[latexpage]
If we are working with categorical data our parameter of interest is often the population proportion, p. The point estimate for p is [latex]\hat{p}[/latex] = [latex]\frac{x}{n}[/latex] where x is the number of successes and n is the sample size. It is also sometimes denoted as [latex]\({p}^{\prime }[/latex]. We saw previously that if we meet conditions, np ≥ 10 and n(1 − p) ≥ 10, we can apply the central limit theorem and assume:
[latex]\hat{p}[/latex] ~N[latex]\left(p,\sqrt{\frac{p\cdot q}{n}}\right)[/latex]
How do you know you are dealing with a proportion problem? First, the underlying distribution is a binomial distribution. This will be categorical data with no mention of a mean or average. If X is a binomial random variable, then X ~ B(n, p) where n is the number of trials and p is the probability of a success.
Hypothesis Tests for p
When you perform a hypothesis test of a single population proportion p, the steps are exactly the same as what we have seen before, however we will calculate our Test Statistic differently. When conducting a test for p, our hypotheses will look as follows:
- Ho: p = p0
- Ha: p (<,>,≠) p0
Recall, the general form of a test statistic is:
[latex]\text{Z=}\frac{\text{point estimate - null value}}{\text{SE}}[/latex]
For the normal distribution of proportions, the z-score formula is as follows:
If [latex]\hat{p}[/latex] ~N[latex]\left(p,\sqrt{\frac{p\cdot q}{n}}\right)[/latex]then the z-score formula is:
\(z=\frac{\hat{p}\text{-p}}{\sqrt{\frac{pq}{n}}}\)
Intuitively, you might think we use this as our test statistic but remember two things:
- We do not actually know p
- In a hypothesis test we begin by assuming the null is true
Sure to these facts, we substitute in p0 for p in the standard error which gives us:
[latex]\sigma_{\hat{p}}\text{ = }\sqrt{\frac{p_o\text{(1-} p_o )}{n}}[/latex]
We then can find a p-value and make our decision as normal
Example
Joon believes that 50% of first-time brides in the United States are younger than their grooms. She performs a hypothesis test to determine if the percentage is the same or different from 50%. Joon samples 100 first-time brides and 53 reply that they are younger than their grooms. For the hypothesis test, she uses a 1% level of significance.
You Try It
Confidence Intervals for p
During an election year, we see articles in the newspaper that state confidence intervals in terms of proportions or percentages. For example, a poll for a particular candidate running for president might show that the candidate has 40% of the vote within three percentage points (if the sample is large enough). Often, election polls are calculated with 95% confidence, so, the pollsters would be 95% confident that the true proportion of voters who favored the candidate would be between 0.37 and 0.43: (0.40 – 0.03,0.40 + 0.03).
Investors in the stock market are interested in the true proportion of stocks that go up and down each week. Businesses that