← Back to Book Detail

High-Performance Computing for Large-Scale Digital Humanities Projects (84/67) -- Contemporary Digital Humanities

Browse
125%

High-Performance Computing for Large-Scale Digital Humanities Projects

High-Performance Computing for Large-Scale Digital Humanities Projects Very large digital humanities projects using “Big Data” can potentially benefit from high-performance computing resources. In this case, “high-performance” indicates the use of multiple processors, known as parallel processing, to simultaneously process large amounts of data. It also indicates large computational tasks being executed, which is known as parallel computing. Although the uptake of utilization of high-performance computational facilities has been substantial in digital cultural heritage, there have been different challenges of re-purposing these resources for textual analysis in the humanities. Historians have used these resources for analyzing census information, specifically for cleaning and managing the data, and for matching census records (Terras, 2009). There are many other potential applications of high-performance computing technology in history. For instance, there is a need for creating a longitudinal database, consisting of census records matched temporally. Such a database would enable tracking individuals, groups, and population changes over time. Another potential application is the creation of variant lists that would address variants and errors in data, which are mostly consequences of the data collection process. Such variant lists would allow normalizing the search process and provide probabilistic and statistical information for further analysis. High-performance computing can also be used to verify census records, which are often missing information. The development of OCR methods for copperplate script (a type of English calligraphic handwriting used in census data), and digitizing missing fields in these records, were needed (Terras, 2009). In other examples of high-performance computing applications to historical scholarship, two projects at University College London in collaboration with the British Library were undertaken to analyze 60,000 digitized fiction and non-fiction books from the 17th through 19th centuries. The data were encoded in compressed ALTO XML code, which describing OCR text and the layout of digitized material. The size of the dataset was approximately 224 GB. The high-performance systems were used for computational queries that could not be performed through standard search interfaces. Results returned from these queries obtained from high-performance computational methods were then provided to scholars for subsequent close reading, visualization, and analysis on local computing resources (Terras et al., 2018). History of Medicine One project involved geospatial and demographic analysis to study how the occurrence of diseases described in the published literature compare to known epidemics in the 19th century, or whether there are correlations between infectious disease outbreaks and mentions of these diseases in fiction and non-fiction literature. The goal was to better understand the cultural response to diseases by ex
← Previous Chapter Next Chapter →