Implementing and Analyzing N-Grams in Python
PYTHON IMPLEMENTATION OF N-GRAMS
To implement n-gram analysis, a machine learning model based on NLP is used. (Please refer to the section on n-grams in the previous course for a full explanation of n-grams.) The n-grams are first generated with NLP operations, such as the ngrams()
function in the Python NLTK (Natural Language Toolkit) library. For instance, if words is a Python list data structure of words
, the operation (note: this example will be presented in further detail below):
nltk.ngrams(words, 2)
returns a zip object of bigrams. To obtain the bigram elements, a Pandas series can be generated from the zip object.
>>> bigram = pd.Series(nltk.ngrams(words, 2))
The bigram series can then be explored.
>>> type(bigram)
<class 'pandas.core.series.Series'>
>>> bigram
0 (much, culture)
1 (culture, pass)
2 (pass, digital)
3 (digital, even)
4 (even, arent)
...
1276 (take, cue)
1277 (cue, technological)
1278 (technological, principle)
1279 (principle, digital)
1280 (digital, artifact)
Length: 1281, dtype: object
The individual elements in the bigram can be accessed either individually, or by slices, like any array.
>>> bigram[0]
('much', 'culture')
>>> bigram[0:10]
0 (much, culture)
1 (culture, pass)
2 (pass, digital)
3 (digital, even)
4 (even, arent)
5 (arent, looking)
6 (looking, screen)
7 (screen, there)
8 (there, good)
9 (good, chance)
dtype: object
For n-grams to be useful, the frequency of each individual word sequence must be known. The Python NLTK supplies the value_counts()
function, which is invoked on an n-gram zip object and returns a Pandas series object.
bigram_count = bigram.value_counts()
This object can also be explored.
>>> type(bigram_count)
<class 'pandas.core.series.Series'>
>>> bigram_count[0]
14
>>> bigram_count[1]
13
>>> len(bigram_count)
1207
>>> bigram_count[0:10]
(binary, code) 14
(digital, technology) 13
(0, 1) 11
(page, number) 6
(digital, culture) 5
(code, 0) 4
(event, digital) 3
(sequence, 0) 3
(discrete, code) 2
(technological, determinism) 2
dtype: int64
The bigram frequencies can be accessed directly from the Pandas series. For instance:
>>> bigram_count[0]
14
>>> bigram_count[100]
1
The frequencies can also be put into a Numpy array and explored as follows:
>>> bigram_freqs = np.array(bigram_count)
>>> type(bigram_freqs)
<class 'numpy.ndarray'>
>>> bigram_freqs
array([14, 13, 11, ..., 1, 1, 1], dtype=int64)
The length 2 word sequences can be put into a list with the keys()
function, which returns the keys of the Pandas series.
## For clarity, extract the keys from the bigram_count series.
>>> keysBigram = bigram_count.keys()
The resulting data structure of keys can be explored.
>>> type(keysBigram)
<class 'pandas.core.indexes.base.Index'>
>>> keysBigram
Index([ ('binary', 'code'), ('digital', 'technology'),
('0', '1'), ('page', 'number'),
('digital', 'culture'), ('code', '0'),
('sequence', '0'), ('event', 'digital'),
('reading', 'possibility'), ('digital', 'artifact'),