The NLP4RE ID-Card
Authors/Creators
Description
# The NLP4RE ID-Card
Replication is an important aspect of empirical evaluation that involves repeating an experiment under similar conditions using a different subject population. Replicability is currently regarded as a major quality attribute in software engineering (SE) research, and it is one of the main pillars of Open Science.
The NLP4RE ID-Card reported in this repository is an artifact that fosters the replication of natural language processing for requirements engineering (NLP4RE) studies.
The NLP4RE ID-Card is a means of capturing a descriptive overview of an NLP4RE (Natural Language Processing for Requirements Engineering) paper by highlighting key information that can be useful in diverse scenarios such
as replication, teaching support, or secondary research.
The ID-Card consists of seven dimensions concerning
- (I) the problem tackled by the paper (RE Task),
- (II) the solution proposed (NLP task(s)),
- (III) the input and output of the solution,
- (IV) the raw data and the annotated dataset,
- (V) the annotators and annotation process,
- (VI) the implementation of the solution proposed in the paper (Tool), and finally
- (VII) how the solution is evaluated (Evaluation).
To get an informative summary, the ID-Card can be applied on one RE task at a time.
Multiple Cards can be used to describe a paper tackling multiple RE tasks (e.g., domain model generation and incompleteness detection). However, the ID-Card allows multiple answers in most of the questions, since the same RE task (e.g.,requirements classification) can be solved using multiple NLP tasks (e.g., information extraction and classification).
## Recommendations for use
In the following, we present some hints to facilitate generating an ID-Card for a particular paper:
- We included the field “Other/Comments” as a possible answer in all questions. This field can be used to add remarks when the question is not applicable, the answer
is not specified in the paper, introduce another alternative that is not among the possible answers, or further comments about the question.
- In Dimension III (NLP Task Details), we are interested in the “initial input” and the “final output” of the proposed solution. Different questions about the output
are adapted to suit the NLP task.
- In case a dimension or specific questions are not considered in the paper (e.g., the paper uses a dataset from the existing literature and so no details are given about
the annotation process), then the dimension/questions can be skipped.
- If a question with a free-text field has multiple answers (the paper is using multiple datasets in IV), then the respective answers should be provided in a commaseparated
format (e.g., “dataset1, dataset2, dataset3”).
## Reference Paper and How to cite
If you use the NLP4RE ID-Card, please cite the reference paper and the repository as below.
Sallam Abualhaija, F. Basak Aydemir, Fabiano Dalpiaz, Davide Dell'Anna, Alessio Ferrari, Xavier Franch, and Davide Fucci. 2024. Replication in Requirements Engineering: The NLP for RE Case. ACM Trans. Softw. Eng. Methodol. 33, 6, Article 151 (July 2024), 33 pages. https://doi.org/10.1145/3658669
Sallam Abualhaija, F. Basak Aydemir, Fabiano Dalpiaz, Davide Dell'Anna, Alessio Ferrari, Xavier Franch, and Davide Fucci. (2024). The NLP4RE ID-Card. Zenodo. https://doi.org/10.5281/zenodo.14197338
Files
NLP4REIDCard.zip
Files
(173.7 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:bf0934c6ad9827a569c88f8c8e68883b
|
173.7 kB | Preview Download |