Published January 13, 2020 | Version v1

Study replicability dataset 3

Authors/Creators

  • 1. Delft University of Technology

Description

This dataset contains all data collected to conduct the studies in Chapter 4 "Code comments for defect prediction" of the Ph.D. thesis "Augmented fine-grained defect prediction for code review".

 

Code comments are a key software component containing information about the underlying implementation. Several studies have shown that code comments enhance the readability of the code. Nevertheless, not all the comments have the same goal and target audience. In this paper, we investigate how 14 diverse Java open and closed source software projects use code comments, with the aim of understanding their purpose. Through our analysis, we produce a taxonomy of source code comments; subsequently, we investigate how often each category occurs by manually classifying more than 40,000 lines of code comments from the aforementioned projects. In addition, we investigate how to automatically classify code comments at line level into our taxonomy using machine learning; initial results are promising and suggest that an accurate classification is within reach, even when training the machine learner on projects different than the target one.

Files

Classifying Code Comments in Java Mobile Applications.zip

Files (89.2 MB)

Name Size Download all
md5:576e9532b53eb41c59b021b1cd0483cc
55.0 kB Preview Download
md5:58bb2e6ce13f72c3db9f68e449cda56f
89.1 MB Download
md5:89689b4a644c875c3ba9caed414beb69
5.2 kB Preview Download