Analysis Of Culture Through Text Analysis
Culturomics is defined as “the application of high-throughput data collection and analysis to the study of human culture” (Michel et al., 2011) (7). The resources used in this analysis are books, newspapers, manuscripts, maps, work if art, and other cultural artifacts (Michel et al., 2011). Culturomic analyses involve millions of books simultaneously, analogous to distant reading in literary studies.
One of the first projects to address cultural questions that can be gleaned from large-scale analysis of printed material was the construction of a large corpus from over five million books to quantitatively analyze cultural trends. The work was performed by a collaborative team of researchers from Harvard University, Google, and elsewhere. The corpus contains over 500 billion words in seven languages – English, German, Hebrew, Chinese, French, Spanish, and Russian.
This analysis was enabled by computational methods related to distant reading, where computational methods are applied to a large body of material. These pre-processing, processing, and analysis steps were performed computationally because of the scale of the tasks and the vast volume of textual material. The pre-processing consisted of the following cleaning steps:
- Digitized books were obtained from the Google Books project, started in 2004, in which collections of over 40 large libraries, as well as books provided directly by publishers, have been digitized. As of 2011, there were 15 million volumes that have been digitized.
- The publication dates for the materials were corrected as needed.
- Volumes were selected based on an OCR (optical character recognition) quality score of the materials.
- Language detection, correction, and filtering were applied.
- The selection and filtering process continued through assessment of metadata field.
The result of the pre-processing is a cleaned source corpus. Processing consisted in several standard text analysis steps:
- Tokenization.
- N-gram construction.
- Statistical counting and aggregation by date of publication.
The result of the processing is the corpus of historical n-grams. A 1-gram is defined as a string of characters uninterrupted by a space or other punctuation. 1-grams include single words (including misspelled ones), abbreviations, and numbers. A sequence of 1-grams (or unigrams) comprises an n-gram, where n denotes the number of 1-grams. Thus, for instance, “barometric pressure” is a 2-gram (n = 2), known as a bigram or diagram, and the phrase “sleight of hand” is a 3-gram (n = 3), also called a trigram. For larger n, the n-gram is usually denoted by n, followed by the suffix “gram”, such as four-gram or ten-gram. In the research described here, n was restricted to 5, and only n-grams occurring at least 40 times in the corpus were considered. Usage frequency is a per-year statistic. It is the fraction of occurrences of an n-gram to the total number of words in a corpus for a particular y