Published May 17, 2021 | Version v1

Data Analysis for the Systematic Literature Review of DL4SE

  • 1. Washington and Lee University
  • 2. College of William and Mary

Description

Data Analysis is the process that supports decision-making and informs arguments in empirical studies. Descriptive statistics, Exploratory Data Analysis (EDA), and Confirmatory Data Analysis (CDA) are the approaches that compose Data Analysis (Xia & Gong; 2014). An Exploratory Data Analysis (EDA) comprises a set of statistical and data mining procedures to describe data. We ran EDA to provide statistical facts and inform conclusions. The mined facts allow attaining arguments that would influence the Systematic Literature Review of DL4SE.

The Systematic Literature Review of DL4SE requires formal statistical modeling to refine the answers for the proposed research questions and formulate new hypotheses to be addressed in the future. Hence, we introduce DL4SE-DA, a set of statistical processes and data mining pipelines that uncover hidden relationships among Deep Learning reported literature in Software Engineering. Such hidden relationships are collected and analyzed to illustrate the state-of-the-art of DL techniques employed in the software engineering context.

Our DL4SE-DA is a simplified version of the classical Knowledge Discovery in Databases, or KDD (Fayyad, et al; 1996). The KDD process extracts knowledge from a DL4SE structured database. This structured database was the product of multiple iterations of data gathering and collection from the inspected literature. The KDD involves five stages:

  1. Selection. This stage was led by the taxonomy process explained in section xx of the paper. After collecting all the papers and creating the taxonomies, we organize the data into 35 features or attributes that you find in the repository. In fact, we manually engineered features from the DL4SE papers. Some of the features are venue, year published, type of paper, metrics, data-scale, type of tuning, learning algorithm, SE data, and so on.
  2. Preprocessing. The preprocessing applied was transforming the features into the correct type (nominal), removing outliers (papers that do not belong to the DL4SE), and re-inspecting the papers to extract missing information produced by the normalization process. For instance, we normalize the feature “metrics” into “MRR”, “ROC or AUC”, “BLEU Score”, “Accuracy”, “Precision”, “Recall”, “F1 Measure”, and “Other Metrics”. “Other Metrics” refers to unconventional metrics found during the extraction. Similarly, the same normalization was applied to other features like “SE Data” and “Reproducibility Types”. This separation into more detailed classes contributes to a better understanding and classification of the paper by the data mining tasks or methods.
  3. Transformation. In this stage, we omitted to use any data transformation method except for the clustering analysis. We performed a Principal Component Analysis to reduce 35 features into 2 components for visualization purposes. Furthermore, PCA also allowed us to identify the number of clusters that exhibit the maximum reduction in variance. In other words, it helped us to identify the number of clusters to be used when tuning the explainable models.
  4. Data Mining. In this stage, we used three distinct data mining tasks: Correlation Analysis, Association Rule Learning, and Clustering. We decided that the goal of the KDD process should be oriented to uncover hidden relationships on the extracted features (Correlations and Association Rules) and to categorize the DL4SE papers for a better segmentation of the state-of-the-art (Clustering). A clear explanation is provided in the subsection “Data Mining Tasks for the SLR od DL4SE”. 5.Interpretation/Evaluation. We used the Knowledge Discover to automatically find patterns in our papers that resemble “actionable knowledge”. This actionable knowledge was generated by conducting a reasoning process on the data mining outcomes. This reasoning process produces an argument support analysis (see this link).

We used RapidMiner as our software tool to conduct the data analysis. The procedures and pipelines were published in our repository.

Overview of the most meaningful Association Rules. Rectangles are both Premises and Conclusions. An arrow connecting a Premise with a Conclusion implies that given some premise, the conclusion is associated. E.g., Given that an author used Supervised Learning, we can conclude that their approach is irreproducible with a certain Support and Confidence.

Support = Number of occurrences this statement is true divided by the amount of statements Confidence = The support of the statement divided by the number of occurrences of the premise

 

 

Files

[correlation] setask-processing.json

Files (2.9 MB)

Name Size Download all
md5:6769cfebd8a8ac06a8823b40f418cb63
6.5 kB Preview Download
md5:de5ae13d2a8c6b02113696b0789253af
6.3 kB Preview Download
md5:3a2cea8a8e8d450a6cf64b93e6c1f257
8.4 kB Preview Download
md5:aaf34722b6deceed06b08b86b0fe962f
8.2 kB Preview Download
md5:62f515637d85a55f4c7d5dd3768f1d18
7.9 kB Preview Download
md5:0aa688a34ea827647d89401bf66c9b8e
18.9 kB Preview Download
md5:3a9b2f235bd333800b7d492951f6f95b
64.2 kB Preview Download
md5:244038a1003177bd30af185f4e803b18
42.3 kB Preview Download
md5:9a06df16429a4fb6dd549c6d651a9514
311.2 kB Preview Download
md5:31318a32e763e5e56213f1a44488803a
29.1 kB Preview Download
md5:3ee09815a6815c13367639a1a3df97e2
16.5 kB Preview Download
md5:b16ed4262387e79640592fb65ae801bd
762.8 kB Preview Download
md5:1821f2fa7fd501824b4c45329bbdb3ea
60.7 kB Preview Download
md5:7aba14e2c75dbf88f809087d6e77cf3c
79.6 kB Preview Download
md5:d3bc95d636ffcfbc331fe0bba42364f2
64.1 kB Download
md5:32083dd342e84e74d15ff021f8c09a33
225.1 kB Preview Download
md5:c0b0949c9c99ae4eb623a0a97b0b4f9d
30.0 kB Preview Download
md5:7375238e4292d899e208606b08bf3e64
881 Bytes Preview Download
md5:d18f4ba64822573bcc4534d770bc9238
17.2 kB Preview Download
md5:b596b2f8f4fc04b889edc60f19404644
113.3 kB Preview Download
md5:1d5a187404383a76d2760a416ac7911a
171.1 kB Download
md5:55cb0c8bd223c7bf0213f983885e3b81
517.5 kB Download
md5:774f0cb26ba5f8de0c4846ca3f5261ce
26.4 kB Download
md5:faf560f56c16e6fc211f0455f2d58e32
78.5 kB Preview Download
md5:a843006ad1ad5dd981c358d07f08d028
78.6 kB Preview Download
md5:848f2aaf14e73ea3eec1e05ac29c56d4
172.6 kB Preview Download