Chapter 11: The Chi-Square Distribution
Chapter 11 Review
11.1 Review
The chi-square distribution is a useful tool for assessment in a series of problem categories. These problem categories include primarily (i) whether a data set fits a particular distribution, (ii) whether the distributions of two populations are the same, (iii) whether two events might be independent, and (iv) whether there is a different variability than expected within a population.
An important parameter in a chi-square distribution is the degrees of freedom df in a given problem. The random variable in the chi-square distribution is the sum of squares of df standard normal variables, which must be independent. The key characteristics of the chi-square distribution also depend directly on the degrees of freedom.
The chi-square distribution curve is skewed to the right, and its shape depends on the degrees of freedom df. For df > 90, the curve approximates the normal distribution. Test statistics based on the chi-square distribution are always greater than or equal to zero. Such application tests are almost always right-tailed tests.
Formula Review
χ2 = (Z1)2 + (Z2)2 + … (Zdf)2
chi-square distribution random variable
μχ2 = df chi-square distribution population mean
[latex]{\sigma }_{{\chi }^{2}}\text{=}\sqrt{2\left(df\right)}[/latex] Chi-Square distribution population standard deviation
If the number of degrees of freedom for a chi-square distribution is 25, what is the population mean and standard deviation?
Solution
mean = 25 and standard deviation = 7.0711
If df > 90, the distribution is _____________. If df = 15, the distribution is ________________.
When does the chi-square curve approximate a normal distribution?
Solution
when the number of degrees of freedom is greater than 90
Where is μ located on a chi-square curve?
Is it more likely the df is 90, 20, or two in the graph?
Solution
df = 2
11.2 Review
To assess whether a data set fits a specific distribution, you can apply the goodness-of-fit hypothesis test that uses the chi-square distribution. The null hypothesis for this test states that the data come from the assumed distribution. The test compares observed values against the values you would expect to have if your data followed the assumed distribution. The test is almost always right-tailed. Each observation or cell category must have an expected value of at least five.
Formula Review
[latex]\sum _{k}\frac{{\left(O-E\right)}^{2}}{E}[/latex]
goodness-of-fit test statistic where:
O: observed values
E: expected values
k: number of different data cells or categories
df = k − 1 degrees of freedom
Determine the appropriate test to be used in the next three exercises.
An archeologist is calculating the distribution of the frequency of the number of artifacts she finds in a dig site. Based on previous digs, the archeologist creates an expected distribution broken down by grid sections in the dig site. Once the site has been fully excavated, she compares the actual numb