← Back to Book Detail

Components of Text Mining (62/67) -- Contemporary Digital Humanities

Browse
92%

Components of Text Mining

Components of Text Mining In this phase, we are extracting text document from various types of external sources into a text index (for subsequent search) as well as a text corpus (for text mining). Document source can be a public web site, an internal file system, etc. For instance, one may need to use a search, examine, or download a predefined list of web sites. The text would then be parsed, and subsequently converted into multiple documents that are stored in a text index and/or a text corpus. Likewise, text may be searched, monitored, and extracted from Twitter for a specific topic. This text may also be stored into a text index and corpus. Additionally, texts in different languages may employ machine translation services, if necessary. To extract text from the Web, a technique known as web scraping is often employed. Web scraping is described in more detail in the section on Internet Studies . At this point, a concise explanation will suffice. It is often difficult to obtain data resources from the Internet and web services, and, consequently, humanities scholarship may not benefit from this vast data repository (Black, 2016). Web scraping is an approach to make this data more accessible by using specialized software programs and scripts to extract specific, meaningful data from web pages. Programs that are used for this purpose include web crawlers, bots, and Web APIs. These programs navigate web pages to obtain target data directly from the HTML that constitutes the web page. This process is quite different than automatically obtaining screen captures or limited amounts of text from the page. Web sites to be “crawled” are pre-selected in a systematic manner. The result is a large collection of unstructured data that must be subsequently transformed into a semi-structured (e.g. TEI) or structured format so that it can be used (Zeng, 2017). Web scrapers are available in Python and R, the two programming languages most frequently used in digital humanities scholarship. After the text of interest has been extracted, it must be transformed into a form conducive to further computational processing. Text transformation consists of several subtasks. Text normalization is generally performed to facilitate and improve text transformation. Conversion of tokens to lowercase, expanding contractions, and the removal of punctuation, numerical values, and stop words are part of this process. Another optional normalization procedure is to convert words to synonymous forms. Wordnet is a large database focused on semantic relationships between words in a large number of languages. Some of these relationships are synonyms, hyponyms – words in a type-of relationship with a hypernym, where the hypernym denotes a supertype, and the hyponym denotes a subtype of the hypernym (e.g. “truck” is a hyponym of the hypernym “vehicle”), and meronyms, which are in a part-of relationship with a corresponding holonym (e.g. “microprocessor” is a meronym to the holonym “compu
← Previous Chapter Next Chapter →