← Back to Book Detail

Big Data And “Smart Data” In the Digital Humanities (9/17) -- Digital Humanities Tools and Techniques ...

Browse
52%

Big Data And “Smart Data” In the Digital Humanities

Big Data And “Smart Data” In the Digital Humanities INTRODUCTION Data (plural of singular datum) are fundamental to the digital humanities, and the relationship between humanities disciplines and data science is becoming more evident as advanced data analysis and visualization techniques are increasingly important for analyzing the large amount of data studied in the humanities. A challenging aspect of this humanities data is its heterogeneity, and that it is situated in a large number of small data sets (such as documents), instead of in a small number of large data sets, as is found in other “Big Data” applications in science, engineering, medicine, and finance. Additionally, much of the data from these latter disciplines are self-descriptive, meaning that the data do not need additional interpretation. For instance, the number “2” is straightforwardly interpreted as two. Numeric or character data have a natural representation in the binary code and is therefore naturally conducive to algorithmic manipulation. Data semantics are inherent in the data themselves. Humanities data, however, are mostly textual, and semantics must be subsequently added. The importance of data is underscored by its addition as the “Fourth Pillar” of scientific endeavor, and of scholarship in general. Traditionally, theory and experimentation have been, and are, paradigmatic for scholarly activities. In the latter part of the 20th century, computation was added as the “Third Pillar”, due to its decisive importance in processing, modeling, simulation, and analysis. More recently, with the explosive interest in data, particularly Big Data, and with the advent and widespread acceptance of data-driven research and the recognition that standard data analysis techniques and database management approaches are insufficient to address these new data-related challenges, data itself has been added as the “Fourth Pillar”. Data are actual representations of objects of interest, such as a person’s name or location. Metadata refers to information that describes data in a particular location or stored in a particular way. In other words, metadata is “data about the data”. For instance, metadata contains the date and time the data were created, what is in the data (e.g., names or address are stored), and general or specific information as to the referent of the data. In the humanities, data are “a digital, selectively constructed, machine-actionable abstraction representing some aspects of a given object of humanistic inquiry.” (Schöch, 2013) (4). WORKING WITH DATA To understand and work with data, it is necessary to describe how data are arranged, particularly when these data need to be processed computationally. There are various linear arrangements of data, including arrays and arrays in higher dimensions, such as matrices (2D arrays). 2D arrays are used in many applications. An image – specifically, a grey-scale image – is a canonical example of a 2D array. Images themselves are tw
← Previous Chapter Next Chapter →