← Back to Book Detail

Consider the following comma-delimited file: (7/8) -- Introduction to Geographic Information S...

Browse
87%

Consider the following comma-delimited file:

Consider the following comma-delimited file: city, sun, temp, precip Los Angeles, 300, 70, 10 London, 50, 55, 40 Singapore, 330, 80, 60 Looking at the contents of the file, we can see that it contains data about the cities of Los Angeles, London, and Singapore. A comma separates each field or attribute, and the file also contains a header row that describes the data contained in each column. Or does it? What does the column “sun” refer to? Is it the number of sunny days this year, last year, annually, or when? What about “temp”? Does this refer to the average daytime, evening, or annual temperature? For that matter, how is temperature measured? In Celsius? Fahrenheit? Kelvin? The column “precip” probably refers to precipitation, but again, what are the units or time frame for such measures and data? Finally, where did these data come from? Who collected them, when were they collected, and for what purpose? It is fantastic to think that such a small text file can lead to so many questions. Now let us extend the example to a file with one hundred records on ten variables, one thousand records on one hundred variables or, better yet, ten thousand records on one thousand variables. Through this rather simple example, many general but central issues that are related to data emerge. Such issues range from the relatively mundane naming conventions that are used to identify individual records (i.e., rows) and distinguish one field (i.e., column) from another, to the issue of providing documentation about what data are included in a given file; when the data were collected; for what purpose are the data to be used; who collected them; and, of course, where did the data come from? The previous simple text file illustrates how we cannot and should not take data and information for granted. It also highlights two essential concepts concerning the source of data and the contents of data files. Concerning data sources, data can be put into one of two distinct categories. The first category is called primary data. Primary data refer to data that are collected directly or on a firsthand basis. For example, if you wanted to examine the variability of local temperatures in May, and you recorded the temperature at noon every day in May, you would be constructing a primary data set. Conversely, secondary data refer to data collected by someone else or some other party. For instance, when we work with Census or economic data collected and distributed by the government, we are using secondary data. Several factors influence the decision behind the construction and use of primary data sets versus secondary data sets. Among the most critical factors are the costs associated with data acquisition in terms of money, availability, and time. The data acquisition and integration phase of most geographic information system (GIS) projects are often the most time-consuming. In other words, locating, obtaining, and putting together the data to be used for a GIS project, whether
← Previous Chapter Next Chapter →