Part 1. Putting Data Driven Research In Perspective
Part 1. Putting Data Driven Research In Perspective
1.3 Why is Structured Data Important?
Why Is Structured Data Important?
The terms “structured data” and “unstructured data” which refer to curated and unprocessed data, respectively, are useful to keep in mind as you work with and prepare spreadsheets. We are often charged with the task in research projects to take unstructured data and make it structured.
More than 80 percent of all data generated today is considered unstructured. When data is structured, it can be converted into different formats, joined with existing datasets, and transformed into data visualizations. Structured data is highly-organized and formatted in a way so it is easily searchable in relational databases. Data collection takes a lot of time. Therefore, drawing on already assembled and cleaned datasets can facilitate the research process. In addition, researchers can combine or add on to already existing datasets.
In this chapter, we will discuss the importance of structured data, understand how pre-assembled datasets can be manipulated, and explain how accessible datasets facilitate data driven research.
Keywords: Structured Data, Unstructured Data, Sustainability, Tidy Data README File, Data Dictionary
Tidy Data
Datasets are the bedrock of data driven research. Because data is produced everyday at an accelerated rate, archiving data is an important concept for researchers to consider. Consequently, Tidy data is an essential concept to consider for research in data analytics.
Hadley Wickham’s “Tidy Data” (2014) offers useful and straightforward principles for organizing data. Wickham defines “tidy data” as “a standardized way to link the structure of a dataset (its physical layout) with its semantics (its meaning).” According to Wickham, “Data preparation is not just a first step, but must be repeated over the course of analysis as new problems come to light or new data is collected.”
Even though Tidy data is a concept meant to help users organize data in R, the general concept also applies to how humanities and social science scholars collect data for our research. In tidy data (1) Each variable forms a column; (2) Each observation forms a row; (3) Each type of observational unit forms a table.
This concept applies to how we collect and organize information in spreadsheets. Many times, in the humanities and social sciences, we have to create our own datasets to fit our research needs. Creating more datasets related to academic subjects and making them publicly available can help spur interactions with data across disciplines. The organization and structure of data dictates how we can merge different types of files.
In order to transform our data into visualizations, we have to be mindful about data organization. Using the concept of Tidy Data, we can ensure that programs like ArcGIS or Tableau Public will be able to interpret our data properly.
Datasets and Sustainability
I have taught classes in digital humanities at t