← Back to Book Detail

Advanced Text Analysis In R (7/11) -- Digital Humanities Tools and Techniques ...

Browse
63%

Advanced Text Analysis In R

Advanced Text Analysis In R INTRODUCTION Lexical dispersion quantifies how frequently a word appears across the parts of a corpus. It is a unitless statistic used in corpus linguistics that quantifies how “evenly” or “uniformly” a word or term occurs throughout the texts of a corpus. Dispersion values are usually normalized to be in the interval 0 to 1. For a specific word w, the minimum dispersion value of zero occurs if and only if w occurs in exactly one text in the corpus; that is, the word or term does not occur uniformly throughout the corpus, and is concentrated in only one text or document of the corpus. The maximum dispersion value of one occurs if and only if w occurs with the same frequency in all texts of the corpus; that is, w is spread uniformly throughout all texts of the corpus (Burch & Egbert, 2019). Several different dispersion measures can be used. The choice of which one to use is dependent on the analysis being performed, or the research question under investigation. In the following discussion, the following nomenclature is used (Gries, 2019): w The word being analyzed, for which dispersion is computed l The number of words or terms in the entire corpus (the length of the corpus) n The number of parts or sub-corpora in the corpus s A vector (list, array) of n percentages of the sized of the n corpus parts (1, 2, …, n) f The overall frequency of w in the entire corpus v A vector (list, array) of n frequencies of w in each part of the corpus (1, 2, …, n) p A vector (list, array) of n percentages that w comprises of each corpus part (1, 2, …, n) In the following discussion, let mp denote the mean of the n p values. Following the development in(Gries, 2019), the biased standard deviation estimator is used (the R function sd calculates the unbiased estimate) (Gries, 2019): [latex]sd_{\textrm{p}} = \sqrt{\frac{\sum\limits_{i = 1}^{n}(p_i-\mu_{\textbf{p}})^2}{n}}[/latex] where [latex]\mu_{\textbf{p}} = \frac{\sum\limits_{i = 1}^{n}p_i}{n} \label{mu_p}[/latex] Several dispersion measures exist, and some of them are presented below (Gries, 2019). A widely used measure is Juilland’s D (D denotes dispersion), denoted simply as D. It is calculated with the standard deviation and mean of the p percentages (Gries, 2019). [latex]D = 1 - \frac{sd_{\textbf{p}}}{\mu_{\textbf{p}}\sqrt{n-1}} \label{D}[/latex] Carroll’s D2 measure, simply denoted as D2, is based on entropy, a concept from information theory (Gries, 2019) [latex]D_2 = \frac{- \sum\limits_{i = 1}^{n}\left(\frac{p_i}{ \sum\limits_{i = 1}^{n}p_i } \times \lg \left( \frac{p_i}{\sum\limits_{i = 1}^{n}p_i} \right)\right)}{\lg{n}} \label{D_2}[/latex] where lg(x) denotes the base-2 logarithm of x (e.g. lg(8) = 3, lg(1024) = 10). Note that if pi = 0, then lg(pi) is represented in R as -Inf. Then 0 x -Inf results in NaN (not a number). Therefore, using this term when computing entropy, 0 x lg(pi) is defined as zero for the case of pi = 0. This situation can be seen in R code as follows. (N
← Previous Chapter Next Chapter →