Text Clustering Based on Tf-Idf Features
INTRODUCTION
In this section, further examples of text clustering and other text processing and analysis are presented.
In previous sections, k-means clustering was performed on numerical values to illustrate the basic concepts. Additionally, text clustering was illustrated through feature sets of a corpus Greco-Roman texts, including known works, commentaries, modern editions, and manuscripts. The first three features were first selected and clustered – i.e., the 3D samples were clustered, and, subsequently, all four features were selected and clustered – i.e., 4D samples were clustered. The characteristics of the texts were already provided, and, in that example, were not directly derived from the texts, which were not provided. In the current example, features will be calculated from texts and clustered. The clustering and analysis are performed in Python. The example that follows is based on the demonstration of sentence clustering in.
K-MEANS CLUSTERING ON TF-IDF FEATURES
Parts of documents were obtained from the first few paragraphs of Wikipedia articles on empires in the ancient world:
- the Seluecid Empire;
- the Pala Empire;
- the Neo-Assyrian Empire;
- the Khmer Empire;
- the Durrani Empire; and
- the Majapahit empire.
To perform this task, the necessary functions are first imported from the Scikit Learn package.
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.cluster import KMeans
The documents can be read from text or into a data frame from a CSV file. However, in this example, the individual documents are short, and therefore the text can be directly incorporated into the Python code.
documents = [
"The Seleucid Empire was a Greek state in Western Asia, ... ",
"The Pala Empire was an imperial power ...",
"The Neo-Assyrian Empire was an Iron Age Mesopotamian empire ... ",
"The Khmer Empire are the terms that historians use to refer ... ",
"The Durrani Empire, also called the Sadozai Kingdom ... ",
"The Majapahit was a Javanese Hindu thalassocratic empire ... "
]
The TfidfVectorizer()
function is used to convert the documents to a matrix of TF-IDF features.
## Initialize a Tfidf vectorizer to convert the documents to TF-IDF features.
## The English stop words supplied by Scikit Learn are used.
vectorizer = TfidfVectorizer(stop_words = 'english')
The next step is for the vectorizer to learn the vocabulary and inverse document frequency (IDF) of the documents, and to generate the document-term matrix.
## Generate the document-term matrix.
X = vectorizer.fit_transform(documents)
The document-term matrix, X, is a sparse Numpy array (i.e., it contains many zeros) where the rows represent the documents, and the columns represent the terms that were determined from the documents. Because the matrix is sparse, it is stored in a compressed format. In the current example, the term frequency-inverse document frequency (TF-IDF) metric is computed on the texts result in a 6 x 627 spa