Clodoaldo Polo Barrera1, María Martínez-Rojas1 and Juan Carlos Rubio-Romero1
Clodoaldo Polo Barrera1, María Martínez-Rojas1 and Juan Carlos Rubio-Romero1
1 Universidad de Málaga.
Keywords: Occupational accidents, Health and Safety, Information Systems, Construction Sector, Manufacturing Sector, KNIME platform.
1. Introduction
Occupational accidents are a problem that affects all work sectors. All of them involve a high human and economic cost both for companies and society, although not all of them are of the same severity [1]. In this work, we focus on the application of data mining techniques and data analysis to compare accidents in construction and manu- facturing sectors in order to learn more about previous workplace accidents. These tech- niques allow to contemplate a large number of possible variables related to each acci- dent and allow to find relational patterns between them [2].
2. Methodology
2.1. Database and filtering process
The database contains a high number of variables (58) which characterize the occur- rence of each of the occupational accidents that are collected. The database has been provided by the Ministry of Labor and Social Economy and it contains the accidents that occurred in Spain between 2009-2018. Each year occurs around half a million work accidents, so that in the data set there are about 6 million cases, each with its 58 asso- ciated characteristics. The data set has a size that implies a considerable computational demand for its analysis. To do this, KNIME [3] data mining platform will be used to filter and analyze data.
After the filtering process, the database is reduced from almost 6 million (5,920,749) to a total of 1,744,252 cases belonging to the mentioned sectors, which again filtered by the same type of node to separate construction and industry, obtaining 704,681 cases and 1,039,571 cases, respectively.
2.2. Variables
A set of variables that are of interest to the research community has been selected [4]. The variables are age, temporality, Physical activities, and Type of contact. By analyz- ing these variables, the aim is to respond questions like the 5 W (who, when, what, how, where) to generate knowledge.
3. Multivariable Analysis with decision tree technique
The following analysis will evaluate the variables contained in the data set to deter- mine which ones are the most relevant to predict the accident severity. To do this, de- cision tree technique which is a supervised data mining method that can serve as an effective tool for multivariate data analysis is used [5].
The technique creates a top-down branching structure, consisting of a root node that is split into several branches. This technique provides simplicity and ease of interpre- tation of the results, allowing them to be evaluated from the beginning to the end of the tree visually node by node. Furthermore, decision trees are useful for our evaluated dataset which contains both for quantitative and qualitative variables. Different tech- niques of decision trees modeling have been tested. The one with the best