← Back to Book Detail

Introduction to T-Sne for High Dimensional Visualization (21/11) -- Digital Humanities Tools and Techniques ...

Browse
190%

Introduction to T-Sne for High Dimensional Visualization

Introduction to T-Sne for High Dimensional Visualization INTRODUCTION In the previous section, principal component analysis (PCA) was introduced as a method to reduce the complexity of high dimensional data. Specifically, PCA is used to reduce the dimension of the original data, while retaining most of the pertinent information in that data, and preserving as much of the variation in the original data as possible. PCA transforms this high dimensional data into a form that can be represented graphically with scatter plots in two or three dimensions. Both PCA and k-means clustering are unsupervised machine learning techniques. They both produce abstract results that may or may not be indicative of the actual characteristics of the data. However, both are extremely useful for analyzing and exploring large, complex data sets. K-means clustering can be used to determine the features that are most characteristic of the data set, and PCA is used to represent these complex data in a simpler form, and, for the purposes of the present discussion, facilitates visualization. Another technique to visualize high dimensional data is t-SNE, or t-distributed stochastic neighbour embedding, a technique widely applied in machine learning applications. Like PCA, t-SNE is an unsupervised method for dimensionality reduction – or, more formally, “embedding” high dimensional data into a low dimensional space – so that this lower dimensional data can be visualized. High dimensional data samples that have similar characteristics are modeled by 2D or 3D points that are close to each other. Those data samples whose characteristics are dissimilar are modeled, with a high probability, by 2D or 3D points that are distant to each other. Although both PCA and t-SNE are unsupervised dimensionality reduction techniques, there are some key differences. PCA is a linear technique wherein large pair-wise distances between data samples are preserved to maximize the variance between the transformed points. Data points that are dissimilar – i.e., they are distant in high dimensional space – are correspondingly distant in the transformed points that PCA returns. In contrast, t-SNE is a nonlinear technique wherein only small pairwise distances between similar points are preserved. T-SNE is widely employed in a variety of machine learning applications. For instance, it is used to visualize neural networks, and especially high dimensional convolutional neural network (CNN) feature maps, which, for the purpose of the current discussion, represent the output of the multiple layers of this network. T-SNE AND ITS PYTHON (SCIKIT LEARN) IMPLEMENTATION The results produced by t-SNE depend on the tuning (adjustment) of hyperparameters. For the current discussion, the most important of these is perplexity. Perplexity has a complex mathematical definition, but it generally can be considered as a variance metric. It is a measure of the number of nearest neighbours a sample has within a region defined b
← Previous Chapter Next Chapter →