Module 11: Chi-Square Tests
Goodness-of-Fit (2 of 2)
Goodness-of-Fit (2 of 2)
Learning outcomes
- Conduct a chi-square goodness-of-fit test. Interpret the conclusion in context.
Here we continue with the details of the chi-square goodness-of-fit hypothesis test. A goodness-of-fit test determines whether or not the distribution of a categorical variable in a sample fits a claimed distribution in the population. The chi-square test statistic is our measure of how much the sample distribution deviates from the population distribution.
As with other hypothesis tests, we need to be able to model the variability we expect in samples if the null hypothesis is true. Then we can determine whether the chi-square test statistic from the data is unusual or typical. An unusual χ2 value suggests that there are statistically significant differences between the sample data and the null distribution and provides evidence against the null hypothesis. This is the same logic we have been applying with hypothesis testing.
Example
Distribution of Color in Plain M&M Candies
Recall the claim made by the manufacturer of M&M candy: the color distribution for plain chocolate M&Ms is 13% brown, 13% red, 14% yellow, 24% blue, 20% orange, 16% green. We used this distribution as our null hypothesis.
- H0: The color distribution for plain M&Ms is 13% brown, 13% red, 14% yellow, 24% blue, 20% orange, 16% green.
- Ha: The color distribution for plain M&Ms is different from the distribution stated in the null hypothesis.
Suppose we buy a large bag of plain M&M candies to test these hypotheses. We randomly select 300 from the bag and view this as a random sample from the population of all plain M&M candies. Our observed counts along with the expected counts are shown in the following ribbon chart and the table. Recall that the expected counts come from the null hypothesis.
We see that the sample distribution is very close to the null distribution for some colors and not others. The deviation appears largest for blue and orange. When we calculate the chi-square statistic, we see that these colors contribute the most to the chi-square value.
[latex]X^2=\frac{(38-39)^2}{39}+\frac{(32-39)^2}{39}+\frac{(51-42)^2}{42}+\frac{(58-72)^2}{72} + \frac{(74-60)^2}{60} + \frac{(47-48)^2}{48}[/latex]
[latex]\approx 0.03+1.26+1.93+2.72+3.27+0.02=9.23[/latex]
What can we conclude? Is this chi-square value unusual or typical? To answer these questions, we must take many random samples from the population described by the null hypothesis. As we have done before, we use a simulation to take random samples. We do this in the next activity.
Try It
Reasoning from the Chi-Square Sampling Distribution
This simulation allows you to click a button and generate random samples of 300 M&Ms. The table shows the color distribution of the sample, and the resulting chi-squared sampling distribution of all samples is plotted below.
Try It
Reasoning from the Chi-Square Sampling Distribution
Recall the distribution of