Published May 6, 2022 | Version v1

Learning software requirements syntax: An unsupervised approach to recognize templates

  • 1. Higher Institute for Applied Sciences and Technology
  • 2. Arab International University

Description

Please cite this work as:

Sonbol, R., Rebdawi, G. and Ghneim, N. Learning software requirements syntax: An unsupervised approach to recognize templates . Knowledge-Based Systems Journal (2022). URL: https://doi.org/10.1016/j.knosys.2022.108933.

https://www.sciencedirect.com/science/article/abs/pii/S0950705122004518

Approach Details:

Requirements are textual representations of the desired software capabilities. Many templates have been used to standardize the structure of requirement statements such as Rupps, EARS, and User Stories. Templates provide a good solution to improve different Requirements Engineering (RE) tasks since their well-defined syntax facilitates the different text processing steps in RE automation researches. However, many empirical studies have concluded that there is a gap between these RE researches and their implementation in industrial and real-life projects. The success of RE automation approaches strongly depends on the consistency of the requirements with the syntax of the predefined templates. Such consistency cannot be guaranteed in real projects, especially in large development projects, or when one has little control over the requirements authoring environment.

In this paper, we propose an unsupervised approach to recognize templates from the requirements themselves by extracting their common syntactic structures. The resultant templates reflect the actual syntactic structure of requirements; hence it can recognize both standard and non-standard templates. Our approach uses techniques from Natural Language Processing and Graph Theory to handle this problem through three main stages (1) we formulate the problem as a graph problem, where each requirement is represented as a vertex and each pair of requirements has a structural similarity, (2) We detect main communities in the resultant graph by applying a hybrid technique combining limited dynamic programming and greedy algorithms, (3) finally, we reinterpret the detected communities as templates.

Our experiments show that the suggested approach can detect templates that follow well-known standards with a 0.90 F1-measure. Moreover, the approach can detect common syntactic features for non-standard templates in more than 73.5% of the cases. Our evaluation indicates that these results are robust regardless of the number and the length of the processed requirements.

 

Dataset Details:

  • The dataset consists of 8084 requirements statements, divided into 82 sets of requirements (for 82 projects), collected from publicly available data sets.
  • The original data sets have different formats: text files, PDFs, XMLs, and My SQL tables. We extracted requirements texts from each of them and applied a set of text cleaning steps. The final format is an xls file where each row represents one requirement.
  • we annotated each requirement with its matched template. We used 5 labels to annotate all sets: 4 of them for the well-known used templates (User Story, Rupp, EARS, Use Case), and an additional label for the remaining cases (Others).

Files

Learning_software_requirements_syntax.ipynb

Files (866.9 kB)

Name Size Download all
md5:5dca1e28b87f84c02c9dd89ea63b781b
388.5 kB Download
md5:5d0453d42925e419e35ee0601743aeb1
27.7 kB Preview Download
md5:a0b2867acd765e3213d7f2127014bb84
450.7 kB Download