← Back to Book Detail

Introduction to Principal Component Analysis (19/11) -- Digital Humanities Tools and Techniques ...

Browse
172%

Introduction to Principal Component Analysis

Introduction to Principal Component Analysis INTRODUCTION In previous sections, k-means clustering was presented as a way to categorize numerical data, and text and documents from the statistics and metrics calculated from this text. An example of clustering ancient Greco-Roman authors based on characteristics of their literary output, such as known works, commentaries, modern editions, and number of manuscripts was presented. Clustering results were analyzed through silhouette scores and silhouette plots. These statistical methods provide metrics to assess clustering performance and are usually sufficient for determining whether a different number of clusters (k) or a different set of features should be applied. In addition, it is often beneficial to explore the clusters produced by k-means visually. As demonstrated in the previous section, for clustering on two features (i.e., 2D data), a scatterplot can be generated wherein the 2D data are directly plotted on the 2D axes (x-axis and y-axis) and coloured according to the clusters to which each data sample belongs. For clustering on three features (i.e., 3D data), a 3D scatter plot (sometimes called a bubble plot) can be used to directly plot the 3D data on the three coordinate axes (x-axis, y-axis, and z-axis) and colour-coded according to their respective clusters. However, as seen in the previous section, interpreting the clusters becomes more difficult in 3D, even with the interactive plot generated in Plotly that allows users to rotate, zoom, and interact with the plot. On a 2D computer monitor, screen, or other output device, it is beneficial to have the capability to “map”, or “transform” data in higher dimensions (3D or higher) to a 2D plot so that the data can be intuitively analyzed. In other words, it would be beneficial to visualize high dimensional clusters in the two dimensions corresponding to computer screens. For instance, the data on Greco-Roman authors in the last section was clustered on both three dimensions (the feature set of known works, commentaries, and modern editions), and four dimensions (the feature set of known works, commentaries, modern editions, and number of manuscripts). Although the 3D clustering could be represented in an interactive 3D scatter plot, interpretation could possibly be facilitated through a 2D representation. The 4D clustering results cannot be directly visualized, although in some visualizations for scientific data, time is sometimes utilized as the fourth dimension. However, for larger corpora with more metadata, the data have higher dimensions, such as 5D, 10D, etc., and even 3D scatter plot representations are inadequate in those cases. Therefore, the most straightforward approach to resolving this problem is to somehow “map” or “transform” the data so that it can be visualized in 2D, as indicated above. Visualizing high dimensional data is an active research area in scientific visualization, where data acquired from experiments, sensors, o
← Previous Chapter Next Chapter →