Published April 4, 2019 | Version v1

Classifying code comments in Java software systems. Appendix

  • 1. Delft University of Technology
  • 2. Software Improvement Group
  • 3. University of Zurich

Description

This dataset refers to "Classifying code comments in Java software systems" paper. It contains a large sample of manual classified code comments. More in deep, code comments are a key software component containing information about the underlying implementation. Several studies have shown that code comments enhance the readability of the code. Nevertheless, not all the comments have the same goal and target audience. In this paper, we investigate how 14 diverse Java open and closed source software projects use code comments, with the aim of understanding their purpose. Through our analysis, we produce a taxonomy of source code comments; subsequently, we investigate how often each category occur by manually classifying more than 40,000 lines of code comments from the aforementioned projects. In addition, we investigate how to automatically classify code comments at line level into our taxonomy using machine learning; initial results are promising and suggest that an accurate classification is within reach, even when training the machine learner on projects different than the target one. Preprint: http://dx.doi.org/10.1007/s10664-019-09694-w

Files

Files (89.1 MB)

Name Size Download all
md5:b40a03a4f39260a8772cbca2f7473a15
89.1 MB Download

Additional details

Funding

European Commission
SENECA - Software ENgineering in Enterprise Cloud Applications systems 642954
Swiss National Science Foundation
Data-driven Contemporary Code Review PP00P2_170529