← Back to Book Detail

Implementing and Analyzing N-Grams in Python (11/11) -- Digital Humanities Tools and Techniques ...

Browse
100%

Implementing and Analyzing N-Grams in Python

Implementing and Analyzing N-Grams in Python PYTHON IMPLEMENTATION OF N-GRAMS To implement n-gram analysis, a machine learning model based on NLP is used. (Please refer to the section on n-grams in the previous course for a full explanation of n-grams.) The n-grams are first generated with NLP operations, such as the ngrams() function in the Python NLTK (Natural Language Toolkit) library. For instance, if words is a Python list data structure of words , the operation (note: this example will be presented in further detail below): nltk.ngrams(words, 2) returns a zip object of bigrams. To obtain the bigram elements, a Pandas series can be generated from the zip object. >>> bigram = pd.Series(nltk.ngrams(words, 2)) The bigram series can then be explored. >>> type(bigram) <class 'pandas.core.series.Series'> >>> bigram 0 (much, culture) 1 (culture, pass) 2 (pass, digital) 3 (digital, even) 4 (even, arent) ... 1276 (take, cue) 1277 (cue, technological) 1278 (technological, principle) 1279 (principle, digital) 1280 (digital, artifact) Length: 1281, dtype: object The individual elements in the bigram can be accessed either individually, or by slices, like any array. >>> bigram[0] ('much', 'culture') >>> bigram[0:10] 0 (much, culture) 1 (culture, pass) 2 (pass, digital) 3 (digital, even) 4 (even, arent) 5 (arent, looking) 6 (looking, screen) 7 (screen, there) 8 (there, good) 9 (good, chance) dtype: object For n-grams to be useful, the frequency of each individual word sequence must be known. The Python NLTK supplies the value_counts() function, which is invoked on an n-gram zip object and returns a Pandas series object. bigram_count = bigram.value_counts() This object can also be explored. >>> type(bigram_count) <class 'pandas.core.series.Series'> >>> bigram_count[0] 14 >>> bigram_count[1] 13 >>> len(bigram_count) 1207 >>> bigram_count[0:10] (binary, code) 14 (digital, technology) 13 (0, 1) 11 (page, number) 6 (digital, culture) 5 (code, 0) 4 (event, digital) 3 (sequence, 0) 3 (discrete, code) 2 (technological, determinism) 2 dtype: int64 The bigram frequencies can be accessed directly from the Pandas series. For instance: >>> bigram_count[0] 14 >>> bigram_count[100] 1 The frequencies can also be put into a Numpy array and explored as follows: >>> bigram_freqs = np.array(bigram_count) >>> type(bigram_freqs) <class 'numpy.ndarray'> >>> bigram_freqs array([14, 13, 11, ..., 1, 1, 1], dtype=int64) The length 2 word sequences can be put into a list with the keys() function, which returns the keys of the Pandas series. ## For clarity, extract the keys from the bigram_count series. >>> keysBigram = bigram_count.keys() The resulting data structure of keys can be explored. >>> type(keysBigram) <class 'pandas.core.indexes.base.Index'> >>> keysBigram Index([ ('binary', 'code'), ('digital', 'technology'), ('0', '1'), ('page', 'number'), ('digital', 'culture'), ('code', '0'), ('sequence', '0'), ('event', 'digital'), ('reading', 'possibility'), ('digital', 'artifact'),
← Previous Chapter Next Chapter →