Software Artifact Mining in Software Engineering Conferences: A Meta-Analysis - Replication Package
Authors/Creators
- 1. Inria
- 2. LTCI, Télécom Paris, Polytechnic Institute of Paris
Description
Software Artifact Mining in Software Engineering Conferences: A Meta-Analysis - Replication Package
Contents
This is a replication package for the paper entitled ["A Systematic Mapping of Software Artifact Mining in Software Engineering Conferences"]. It contains 16 years of history for each of the following conferences:
- ICSE, International Conference on Software Engineering
- ASE, IEEE/ACM International Conference on Automated Software Engineering
- FSE, ACM SIGSOFT Symposium on the Foundations of Software Engineering
- ICSM, IEEE International Conference on Software Maintenance
- MSR, Working Conference on Mining Software Repositories
- WCRE, Working Conference on Reverse Engineering
- ICSME, International Conference on Software Maintenance and Evolution
- ICPC, IEEE International Conference on Program Comprehension
- SCAM, International Working Conference on Source Code Analysis & Manipulation
The data is stored in a PostgreSQL database (see the [PostgreSQL Dump] in folder db).
Alternatively, the database can be recreated from CSV files using Python and the SQLAlchemy Object Relational Mapper using the scripts included (more details below).
Data
- Papers and authors: the DBLP data dump. We used the data in dblp-2021-11-02.xml file.
Using the replication
Directly
Most simply, you can import the [SQL dump] in the folder db into your database management system and start querying.
Via Python
Alternatively, you can take a look at how the database was created using PostgreSQL, Python and SQLAlchemy, and use these mechanisms also for querying. This will allow you to easily extend the database or update its schema.
Dependencies and installation instructions
If you take this path, make sure you have Python and a PostgreSQL server installed before attempting anything. Follow the follwoing steps (tested on our OS 11.3 machine with Python 3.7.7):
- Install SQLAlchemy:
easy_install SQLAlchemy - Tweek
database.inifor your particular PostgreSQL user and password (the script assumes user root with empty password) - Install CERMINE [https://github.com/CeON/CERMINE] to extract content from PDF files or use cermine.jar file included here.
Python scripts
initDB.py: declares the database schema using Python classes (will be automatically mapped to tables by SQLAlchemy).populateDB.py: reads data about the papers for each conference and loads it into the database.downloadPdf.py: download the pdf of the papers using a modified version PyPaperBot (The source code of our PyPaperBot is in the replication package).cermine.py: Extract the text from the Pdf files into XML files the pdf.
How to use
Python files arguments:
| Arguments | Description | Type |
|---|---|---|
| --dir | Directory path in which to save the result | (str) |
Jupyter notebooks
The various Jupyter notebooks containing all the scripts used to gather the data and answer the research questions.
1_main.ipynb: Runs populateDB, downloadPdf and cermine files.2_SectionHeadersExtraction.ipynb: Extract the header of the section which helps us exclude irrelevant sections.3_NLP.ipynb: Generate n-grams and update the database.4_PreliminaryAnalysis.ipynb:5_RQ1-Artifacts.ipynb: Execute the code and generate the figures to answer RQ1.6_RQ2-ArtifactCombinations.ipynb: Execute the code and generate the figures to answer RQ2.7_RQ3_Purposes.ipynb: Execute the code and generate the figures to answer RQ3.
db : a repository containing all the data of our study after all processing steps.
figs : The generated figures from the notebooks for the paper.