180 Putting It Together: Inference for Means
Let’s Summarize
The focus of this module, Inference for Means, is inference for a population mean or a difference between two populations means. We began this module with a discussion of the sampling distribution of sample means. We then developed a probability model based on this sampling distribution. We used the probability model with an actual sample mean to test a claim about population mean in a hypothesis test or to estimate a population mean with a confidence interval. We then moved to inference for a difference in two population means (or a treatment effect.)
Sampling Distribution of Means
If we have a quantitative data set from a population with mean µ and standard deviation σ, the model for the theoretical sampling distribution of means of all random samples of size n has the following properties:
- The mean of the sampling distribution of means is µ.
- The standard deviation of the sampling distribution of means is [latex]σ/\sqrt{n}[/latex].
- Notice that as n grows, the standard error of the sampling distribution of means shrinks. That means that larger samples give more accurate estimates of a population mean.
- For large enough sample size, the sampling distribution of means is approximately normal (even if population is not normal). This is called the central limit theorem.
- If a variable has a skewed distribution for individuals in the population, a larger sample size is needed to ensure that the sampling distribution has a normal shape.
- The general rule is that if n is at least 30, then the sampling distribution of means will be approximately normal. However, if the population is already normal, then any sample size will produce a normal sampling distribution.
- We practiced finding a probability associated with a range of sample means, which is similar to finding a P-value in hypothesis testing. The process is as follows.
- Convert a sample mean X into a z-score: [latex]Z=\frac{\stackrel{¯}{x}-μ}{σ/\sqrt{n}}[/latex]
- Use technology to find a probability associated with a given range of z-scores.
Confidence Intervals
The Form
A confidence interval approximates a population mean by giving us a range of values that likely contains the population mean μ. The general form of the confidence interval is
[latex]\stackrel{¯}{x}±\mathrm{margin}\text{}\mathrm{of}\text{}\mathrm{error}=\stackrel{¯}{x}±(\mathrm{critical}\text{}\mathrm{value})⋅(\mathrm{standard}\text{}\mathrm{error})[/latex]
We covered three different types of confidence intervals:
One-sample Z-interval: [latex]\stackrel{¯}{x}±{Z}_{c}⋅σ/\sqrt{n}[/latex], where σ is the population standard deviation (when it is known).
One-sample T-interval: [latex]\stackrel{¯}{x}±{T}_{c}⋅s/\sqrt{n}[/latex], where s is the sample standard deviation.
Two-sample T-interval: [latex]({\stackrel{¯}{x}}_{1}\text{−}{\stackrel{¯}{x}}_{2})±{T}_{c}⋅\sqrt{\frac{{{s}_{1}}^{2}}{{n}_{1}}+\frac{{{s}_{2}}^{2}}{{n}_{2}}}[/latex], where we use the sample statisti