← Back to Book Detail

Part IV. ALGEBRAIC REASONING (12/7) -- Quantitative Problem Solving in Natural ...

Browse
171%

Part IV. ALGEBRAIC REASONING

Part IV. ALGEBRAIC REASONING 12. Relationships Between Variables In the previous chapter, our discussion of variables and functions largely assumed that relationships were known or developed independent of any measurement or data. However, functional relationships between variables can also be derived from data. Here, we explore two concepts that help us understand the strength and nature of systematic relationships between variables. 12.1 Correlation In common parlance, the word correlation suggests that two events or observations are linked with one another. In the analysis of data, the definition is much the same, but we can even be more specific about the manner in which events or observations are linked. The most straight-forward measure of correlation is the linear correlation coefficient, which is usually written r (and is, indeed, related to the r2 that we cite in assessing the fit of a regression equation). The value of r may range from -1 to 1, and the closer it is to the ends of this range (i.e., [latex]\mid r \mid \rightarrow 1[/latex]), the stronger the correlation. We may say that two variables are positively correlated if r is close to +1, and negatively correlated if r is close to -1. Poorly correlated or uncorrelated variables will have r closer to 0. To the right are two plots comparing life-history and reproductive traits of various mammals. In the first one, Figure 12.1, the arrangement of points in a band from lower left to upper right on the graph is relatively strong, corresponding to a relatively high r of 0.73. In contrast, the correlation between litter number per year and litter size in Figure 12.2 is (surprisingly?) weak, producing more of a shotgun pattern and r a modest 0.36. In the abstract, the mathematical formula for the correlation between two variables, x and y, can be written: [latex]\large{r = \frac{1}{n}\sum_{i=1}^n \bigg(\frac{x_i - \bar{x}}{\sigma_x}\bigg)\bigg(\frac{y_i - \bar{y}}{\sigma_y}\bigg),} \tag{12.1}[/latex] where the subscript i corresponds to the ith observation, the overbar indicates mean values, and [latex]\sigma_x[/latex] and [latex]\sigma_y[/latex] are standard deviations. The specifics of this formula are not of great interest to us. The important thing to understand is that when positive changes in one variable are clearly linked with positive changes in a second variable, this indicates a good, positive correlation, r > 0. The same is true if negative changes in one variable correspond to negative changes in the other. However, positive changes in one variable corresponding to negative changes in another indicate negative correlation, r < 0. It is also important to note that this is a good measure of correlation only for linear relationships, and even if two variables are closely interdependent, if their functional dependence is not linear, the r value will not be particularly helpful. Nevertheless, correlation can still help us identify key relationships when we first encounter a datase
← Previous Chapter Next Chapter →