DOI;Link;Publication year;Publication type;Authors;Title;Are the aims of the research clearly defined?;Are the performance measures used to assess the models clearly defined?;Are the performance measures used to assess the models considered credible?;Are the limitations or threats to validity of the study specified?;Is the proposed method or methods compared with other methods and/or baselines?;Are the findings of study clearly stated and supported by reported results?;Does the study provide convincing arguments about additional value given to academia or industry community?;Sum;PQ1: Was the modelling process explained in detail?;CC: PQ1;PQ2: Was a replication package published?;CC: PQ2;Verification: PQ2;PQ2 annotation;PQ3: How were projects selected?;CC: PQ3;PQ4: How were samples from projects for classification selected?;CC: PQ4;PQ5: How were code smells assessed?;CC: PQ5;Crosscheck: LM comment;Crosscheck resolution;Chart: RQ1.1;Chart: RQ1.2;Chart: RQ2.1;Chart: RQ2.2;Chart: RQ2.3 10.1016/j.infsof.2017.09.011;https://doi.org/10.1016/j.infsof.2017.09.011;2018;Journal;"Chen, Z.;Chen, L.;Ma, W.;Zhou, X.;Zhou, Y.;Xu, B.";Understanding metric-based detectable smells in Python software: A comparative study;Y;Y;P;Y;N;P;Y;5;Detailed;Detailed;Yes, without documentation;Yes, w/o doc;Yes, without documentation (OK);scripts & data;Manual selection (well-known, non-trivial projects);Selected 106 Python projects with most stars on GitHub. So there is a selection method behind it. DIFF: not quite manual selection;Strict thresholds (advisors);statistics-based thresholds were applied;Authors' assessment, verified by 3 PhD students and 2 MSc students (at least 3 years research experience), definitions provided by authors (but not detailed detection methods);"Authors' ""invited three masters and two Ph.D. students independent of the authors as volunteers to identify whether an example is a positive or negative smell instance."" - DIFF in number of students";Slight DIFFerences in PQ5 and PQ3;"PQ5 - adjusted, data acquisition error (not affecting results) PQ3 - changed ""Manual - criteria given"" to ""Automated""";Detailed;Scripts & data;Automated;Advisors;Authors & students 10.1109/ASE.2017.8115667;https://doi.org/10.1109/ASE.2017.8115667;2017;Conference;"Ocariza, F.S.;Pattabiraman, K.;Mesbah, A.";Detecting unknown inconsistencies in web applications;Y;P;P;P;N;Y;Y;4,5;Detailed;Detailed;Scripts - no, program - claimed yes, link dead;DIFF: Link is NOT dead! The tool is available, see http://ece.ubc.ca/~frolino/projects/holocron/;Scripts - no, program - yes (link OK - issue: wrong tilde character);scripts;Manual selection (taken from MVC application lists on GitHub);Manual selection of open-source web applications from each of the three main MVC frameworks (AngularJS, BackboneJS, and Em-ber.js) from GitHub;GitHub search in issue text;"GitHub's advanced search for GitHub issues that are given the label ""bug"", and whose status is ""closed""";From bug reports;From bug reports;DIFF in PQ2 the link is NOT dead! The tool is available, see http://ece.ubc.ca/~frolino/projects/holocron/;PQ2 - fixed during extra verification phase;Detailed;Scripts;Manual - criteria given;Advisors;Automated 10.1109/MOBILESoft.2017.29;https://doi.org/10.1109/MOBILESoft.2017.29;2017;Conference;"Kessentini, M.;Ouni, A.";Detecting Android Smells Using Multi-Objective Genetic Programming;P;Y;P;Y;Y;Y;Y;6;Detailed;Detailed;No;No;No (OK);;Manual selection (no criteria given);"""184 Android projects with source code hosted in GitHub"" but lack of selection criteria";Algorithm run;Not precisely defined.;34 graduate students - 13 full-time developer or manager (4-9 years), rest at least 2 years experience, lectured and tested about smells, majority vote;"""34 graduate students from a Software Quality Assurance course at the University of Michigan to analyze the results. Participants include 13 students who are working as full time developer or manager in software companies ranging from 4 to 9 years of programming experience. The remaining participants have minimum years of experience of 2 years in industry as programmer."", ""all the participants attended one lecture about Android smells and passed eight tests"", ""the majority of votes [was used] to determine if suggested Android smells are correct or not"".";DIFF in PQ4;"PQ4 - changed to ""Unknown""";Detailed;None;Manual - no criteria;Unknown;Students & developers 10.1007/s11219-016-9309-7;https://doi.org/10.1007/s11219-016-9309-7;2016;Journal;"Mansoor, U.;Kessentini, M.;Maxim, B.R.;Deb, K.";Multi-objective code-smells detection using good and bad design examples;Y;Y;P;P;Y;Y;Y;6;Detailed;Detailed;No;No;Unlikely - probable package would be at http://staff.unak.is/andy/StaticAnalysis0809/metrics/overview.html but link is dead;;From earlier literature;From earlier literature;Existing corpus (Khomh, Palomba, Kessentini);Existing corpus (Khomh, Kessentini, Palomba);15 people - 8 MSc students, 5 PhD students, 2 faculty members (experience 1 to 17 years);"15 student subjects ""Subjects included eight master students in software engineering, five Ph.D. students in software engineering and two faculty members in software engineering""";OK;OK;Detailed;Unresolvable;Earlier studies;Existing corpus;Authors & students 10.5220/0006338804740482;https://doi.org/10.5220/0006338804740482;2017;Conference;"Hozano, M.;Antunes, N.;Fonseca, B.;Costa, E.";Evaluating the accuracy of machine learning algorithms on detecting code smells for different developers;Y;P;P;P;N/A;Y;Y;4,5;Briefly (referenced reproduction with modifications);"Briefly (only names of the ML algorithms implemented in Weka are give, configuration details are not explicitly given). ""Initially, we applied these algorithms by following the same configuration adopted (Fontana et al., 2015) and then we tried other configurations in order to find one in which the analyzed algorithm could reach its highest efficiency""";No;"No (""we will make the dataset used in our ex- periments available in order to help other studies in smell detection"" but lack of URL)";"""we will make the dataset used in our experiments available in order to help other studies in smell detection"", but no link";;From earlier literature;Ad-hoc from earlier literure (GanttProject);Existing corpus (Fontana);"Sample of existing corpus (""The code snippets used in our study were extracted from GanttProject1 (v2.0.10), an open source Java project. We selected this project because they had been used in other studies related to code smells (Moha et al., 2010; Khomh et al., 2011b; Fon- tana et al., 2015). Moreover, such studies reported a variety of suspicious code smells in this project."")";40 developers with at least 3 years of experience;"""40 developers with at least 3 years experience in software development""";OK;OK;Briefly;Missing;Earlier studies;Existing corpus;Developers 10.1007/s10664-015-9378-4;https://doi.org/10.1007/s10664-015-9378-4;2016;Journal;"Arcelli Fontana, F.;Mantyla, M.V.;Zanoni, M.;Marino, A.";Comparing and experimenting machine learning techniques for code smell detection;P;Y;P;Y;N;Y;Y;5;Detailed;Detailed;Data - yes, models - yes, scripts - no;Data - theoretically yes but in practice only some are available (e.g., http://essere.disco.unimib.it/reverse/files/mlcsd_files/datasets/evaluation_dataset.zip is not available URL error 404), models - no, tools -yes;Data - yes, models - yes, creation tool - yes (published in a separate study, but available under same web page), tool configuration - no;scripts & data;Existing corpus (Qualitas Corpus) without non-compilable items;QualitasCorpus without non-compatible ones;External advisors (PMD and iPlasma);External advisors (PMD and iPlasma);3 MSc students trained and lectured for the task;Manual assesment by 3 MSc students trained by researchers;DIFFerence: PQ2 (some data not available AND lack of models);PQ2 - agreed, misclassification during initial data acquisition. Does not affect results;Detailed;Scripts & data;Existing corpus subset;Advisors;Students 10.1109/TSE.2015.2503740;https://doi.org/10.1109/TSE.2015.2503740;2016;Journal;"Liu, H.;Liu, Q.;Niu, Z.;Liu, Y.";Dynamic and Automatic Feedback-Based Threshold Adaptation for Code Smell Detection;Y;Y;P;Y;Y;Y;Y;6,5;Detailed;Detailed;Claimed yes, but the whole page is in Chinese so it's hard to say (no reference seems to match);No (server is not responding, URL http://sei.pku.edu.cn/~liuhui04/data/scripts.7z is not available);Yes - http://sei.pku.edu.cn/?liuhui04/data/scripts.7z (link OK - issue: wrong tilde character);scripts;Manual selection (no criteria given);"Ad hoc (""To collect data (code smells) for evaluation, we ask three engineers to collect and analyze code smells from five open- source applications from SourceForge"")";3 engineers reading through the applications;"""Three engineers manually analyzed every method in selected applications""";3 engineers (graduate students);"""Three engineers manually analyzed every method in selected applications""";DIFFerence: PQ2 - No (server is not responding, URL http://sei.pku.edu.cn/~liuhui04/data/scripts.7z is not available);PQ2 - server failure, not realated to paper, does not affect results;Detailed;Scripts;Manual - no criteria;Exhaustive search;Students 10.1109/ISSRE.2015.7381819;https://doi.org/10.1109/ISSRE.2015.7381819;2015;Conference;"Amorim, L.;Costa, E.;Antunes, N.;Fonseca, B.;Ribeiro, M.";Experience report: Evaluating the effectiveness of decision trees for detecting code smells;Y;Y;P;Y;Y;P;Y;6;Detailed;Detailed;No (data claimed yes, but link dead);Data: yes (link works);Data: yes (one of the links in HTML and both links in PDF work - proper one is not https://goo.gl/T2itbC but https://goo.gl/T2ifbC );data;Referred data set (Ptidej);Referred data set (Ptidej);Referred data set (Ptidej);Referred data set (Ptidej);Referred data set (Ptidej);Referred data set (Ptidej);Difference in PQ2;PQ2 - fixed during extra verification phase;Detailed;Data;Predefined;Predefined;Predefined 10.1007/978-3-319-47106-8_24;https://doi.org/10.1007/978-3-319-47106-8_24;2016;Conference;Mkaouer, M.W.;Interactive code smells detection: An initial investigation;Y;Y;P;N;Y;Y;Y;5,5;Briefly;Briefly;No;No;No (OK);;Manual selection (well-known, open source, also referred in literature);Manual selection of only 4 projects (well-known, open source, also referred in literature);Advisors (InCode, Mantyla et al.);Advisors (e.g., InCode);2 PhD students;two Ph.D. students was asked to evaluate, manually, whether the suggested code fragments do contain the reported smell;OK;OK;Briefly;None;Manual - criteria given;Advisors;Students 10.1109/ESEM.2015.7321194;https://doi.org/10.1109/ESEM.2015.7321194;2015;Conference;"Fu, S.;Shen, B.";Code Bad Smell Detection through Evolutionary Data Mining;Y;Y;P;N;P;Y;P;4,5;Briefly;Briefly;No;No;No (OK);;Manual selection (no criteria given);Ad-hoc convenience selection w/o specified criteria;Manual system assessment;Manual assesments of code shapshots by students;6 MSc students;"Six students were involved in assessement ""Three Master students from Shanghai Jiao Tong University were invited to manually detect the bad smells. They analyzed the snapshot of systems, and looked for bad smells according to their definitions, not aware of our detection approach in advance. And other three master students verified whether the found bad smells are correct or not. The verified bad smells forms the complete set.""";OK;OK;Briefly;None;Manual - no criteria;Exhaustive search;Students 10.1109/ICSE.2015.244;https://doi.org/10.1109/ICSE.2015.244;2015;Conference;Palomba, F.;Textual Analysis for Code Smell Detection;Y;P;P;N;Y;Y;Y;5;Very briefly (only refers to used methods);Briefly;No;No;No (OK);;Referred data set (Palomba);Ad-hoc convenience selection w/o specified criteria;Referred data set (Palomba);"""Details on how these smells have been manually identified can be found in the paper by Palomba et al. [21].""";Referred data set (Palomba);"""Details on how these smells have been manually identified can be found in the paper by Palomba et al. [21].""";"DIFFerence: PQ1 - Briefly instead of Very Briefly as we only have 3 levels and ""Briefly"" is the lowest one";PQ1 - agreed, fixed during data postprocessing;Briefly;None;Predefined;Predefined;Predefined 10.1109/TSE.2014.2331057;https://doi.org/10.1109/TSE.2014.2331057;2014;Journal;"Kessentini, W.;Kessentini, M.;Sahraoui, H.;Bechikh, S.;Ouni, A.";A Cooperative Parallel Search-Based Software Engineering Approach for Code-Smells Detection;Y;Y;P;Y;Y;Y;Y;6,5;Detailed;Detailed;No;No;No (OK);;Referred data set (Ouni, Kessentini, Moha);"Ad-hoc convenience selection of existing corpus/data set (""existing corpus [16], [17], [23]"", Moha et al, Moha and Gueheneuc [aka DECOR], Ouni et al.)";Referred data set (Ouni, Kessentini, Moha);identified manually - ad-hoc convenience sampling, referred data set;Referred data set (Ouni, Kessentini, Moha);Referred data set;OK;OK;Detailed;None;Predefined;Predefined;Predefined 10.1145/2675067;https://doi.org/10.1145/2675067;2014;Journal;"Sahin, D.;Kessentini, M.;Bechikh, S.;Deb, K.";Code-smell detection as a bilevel problem;Y;Y;P;Y;Y;Y;Y;6,5;Detailed;Detailed;No;No;No (OK);;Referred data set (Moha);Referred data set;Referred data set (Moha);Referred data set;Referred data set (Moha);Referred data set;OK;OK;Detailed;None;Predefined;Predefined;Predefined 10.1109/ICSM.2013.56;https://doi.org/10.1109/ICSM.2013.56;2013;Conference;"Fontana, F.A.;Zanoni, M.;Marino, A.;Mantyla, M.V.";Code smell detection: Towards a machine learning-based approach;P;P;P;N;N;P;P;2,5;Briefly;Briefly;No;No;No (OK);;Existing corpus (Qualitas Corpus) without non-compilable items;Existing corpus (Qualitas Corpus) without non-compilable items;External advisors (PMD, iPlasma, Anti-Pattern Scanner, Fluid Tool, Marinescu rules);External advisors (PMD, iPlasma, Anti-Pattern Scanner, Fluid Tool, Marinescu detection rule);3 MSc students trained and lectured for the task;3 MSc students trained for the task;OK;OK;Briefly;None;Existing corpus subset;Advisors;Students 10.1109/ASE.2013.6693086;https://doi.org/10.1109/ASE.2013.6693086;2013;Conference;"Palomba, F.;Bavota, G.;Di Penta, M.;Oliveto, R.;De Lucia, A.;Poshyvanyk, D.";Detecting bad smells in source code using change history information;P;Y;P;Y;Y;Y;Y;6;Briefly;Moderate;No;No (Claimed yes, but page dead: http://www.rcost.unisannio.it/mdipenta/papers/ase2013);Claimed yes, but page dead: http://www.rcost.unisannio.it/mdipenta/papers/ase2013;;Manual selection (no criteria given);Manual selection (no criteria given);Manual system assessment;"Manual (""A Master's student from the University of Salerno manually identified instances of the five considered smells in each of the systems' snapshots"")";2 MSc students;"2 MSc students (""A Master's student from the University of Salerno manually identified instances of the five considered smells in each of the systems' snapshots. Starting from the definition of the five smells reported in literature, the student manually analyzed each snapshot's source code looking for instances of those smells... A second Master's student validated the produced oracle..."")";"DIFFerence: PQ1 - ""Briefly"" vs ""Moderate""";"PQ1 - agreed, changed to ""Moderate""";Moderate;Unresolvable;Manual - no criteria;Exhaustive search;Students 10.1109/ICSE.2013.6606670;https://doi.org/10.1109/ICSE.2013.6606670;2013;Conference;"Gauthier, F.;Merlo, E.";Semantic smells and errors in access control models: A case study in PHP;P;N;N;N;N;P;P;1,5;Detailed;Detailed/Moderate;No;"Partly (on the positive side, used tool is available vide http://gibbslda.sourceforge.net and was used with default parameters ""LDA modeling was performed with the GibbsLDA++ [11] tool with default parameters"")";No (OK);;Manual selection (no criteria given);Manual selection of one project (Moodle) (no criteria given);Algorithm run;Algorithm run;Authors' assessment;Authors' assessment;DIFFerence: PQ2 - Partly (tool is available vide http://gibbslda.sourceforge.net), Verification: PQ2, ~PQ1;"PQ1 - left ""detailed"" PQ2 - changed ""None"" to ""Scripts""";Detailed;Scripts;Manual - no criteria;Advisors;Authors 10.1007/978-3-642-39742-4_6;https://doi.org/10.1007/978-3-642-39742-4_6;2013;Conference;"Boussaa, M.;Kessentini, W.;Kessentini, M.;Bechikh, S.;Ben Chikha, S.";Competitive coevolutionary code-smells detection;Y;Y;P;N;Y;Y;Y;5,5;Detailed;Detailed;No;No;No (OK);;Referred data set (Moha);Referred data set (Moha);Referred data set (Moha);Referred data set (Moha);Referred data set (Moha);Referred data set (Moha);OK;OK;Detailed;None;Predefined;Predefined;Predefined 10.1109/WCRE.2012.56;https://doi.org/10.1109/WCRE.2012.56;2012;Conference;"Maiga, A.;Ali, N.;Bhattacharya, N.;Sabane, A.;Gueheneuc, Y.-G.;Aimeur, E.";SMURF: A SVM-based incremental anti-pattern detection approach;Y;P;P;P;Y;P;Y;5;Detailed;Detailed;No;"Data: Yes http://www.ptidej.net/download/experiments/wcre12a/ Scripts: No";"Data: Yes, but only link in PDF valid ( http://www.ptidej.net/download/experiments/wcre12a/ not http://www.ptidej.net/download/experiments/were12a/ )";data;Referred data set (Moha);"Referred Khomh's data set: [13] F. Khomh, S. Vaucher, Y.-G. Gu ?eh ?eneuc, and H. Sahraoui,""Bdtex: A gqm-based bayesian approach for the detection ofantipatterns,""J. Syst. Softw., vol. 84, no. 4, pp. 559-572, Apr.2011";Referred data set (Moha);"Referred Khomh's data set: [13] F. Khomh, S. Vaucher, Y.-G. Gu ?eh ?eneuc, and H. Sahraoui,""Bdtex: A gqm-based bayesian approach for the detection ofantipatterns,""J. Syst. Softw., vol. 84, no. 4, pp. 559-572, Apr.2011";Referred data set (Moha);"Referred Khomh's data set: [13] F. Khomh, S. Vaucher, Y.-G. Gu ?eh ?eneuc, and H. Sahraoui,""Bdtex: A gqm-based bayesian approach for the detection ofantipatterns,""J. Syst. Softw., vol. 84, no. 4, pp. 559-572, Apr.2011";Difference in PQ2, PQ3, PQ4, PQ5;"PQ2 - fixed during extra verification phase PQ3,4,5 - error during initial data acquisition, does not affect results";Detailed;Data;Predefined;Predefined;Predefined 10.1145/2338965.2336785;https://doi.org/10.1145/2338965.2336785;2012;Conference;"Pradel, M.;Heiniger, S.;Gross, T.R.";Static detection of brittle parameter typing;Y;P;P;N;N;Y;Y;4;Detailed;Detailed;Yes, with brief documentation;Yes;Yes, with brief documentation (http://mp.binaervarianz.de/issta2012/);scripts & data;Manual selection (no criteria given);Manual selection (no criteria given);Algorithm run + bug seeding;Algorithm run;Authors' assessment + developer confirmation;Authors' assessment + developer confirmation;DIFFerence in PQ2 Verification (URL);OK;Detailed;Scripts & data;Manual - no criteria;Bug seeding;Authors 10.1109/QUATIC.2010.61;https://doi.org/10.1109/QUATIC.2010.61;2010;Conference;"Hassaine, S.;Khomh, F.;Gueheneucy, Y.-G.;Hamel, S.";IDS: An immune-inspired approach for the detection of software design smells;Y;Y;P;P;P;P;Y;5;Briefly;Briefly;No;No;Only description of used metrics: http://www.ptidej.net/downloads/replications/quatic10/;;Manual selection (subset from Moha data set);Manual selection (subset from Moha data set);Manual system assessment;Manual system assessment;Authors' assessment;Authors' assessment;OK;OK;Briefly;None;Manual - no criteria;Exhaustive search;Authors 10.1109/QUATIC.2010.60;https://doi.org/10.1109/QUATIC.2010.60;2010;Conference;"Bryton, S.;Brito E Abreu, F.;Monteiro, M.";Reducing subjectivity in code smells detection: Experimenting with the Long Method;Y;Y;Y;P;N;P;P;4,5;Detailed;Detailed;No;No;No (OK);;Manual selection (widely used, small enough for manual assessment);Manual selection;Manual system assessment;Manual selection;Authors' assessment;Authors' assessment;OK;OK;Detailed;None;Manual - criteria given;Exhaustive search;Authors 10.1109/QSIC.2009.47;https://doi.org/10.1109/QSIC.2009.47;2009;Conference;"Khomh, F.;Vaucher, S.;Gueeheeneuc, Y.-G.;Sahraoui, H.";A bayesian approach for the detection of code and design smells;Y;P;P;N;Y;Y;Y;5;Detailed;Detailed;Data - yes, scripts - no;Data only (http://www.ptidej.net/downloads/replications/qsic09/);Data - yes ( http://www.ptidej.net/downloads/replications/qsic09/ ), scripts - no;data;Manual selection (small enough for manual assessment, subset of Moha);Manual selection;Manual system assessment;Manual selection;2 undergraduate students, 2 graduate students;"two undergraduate students and two graduate students (""To build the corpus, we asked two undergraduate students and two graduate students to detect occur- rences of the Blob in the two programs."")";OK;OK;Detailed;Data;Manual - criteria given;Exhaustive search;Students 10.1016/j.entcs.2005.02.059;https://doi.org/10.1016/j.entcs.2005.02.059;2005;Journal;Kreimer, J.;Adaptive detection of design flaws;P;P;P;N;N;P;P;2,5;Detailed;Detailed;No;"No (""For evaluating the concept of section 3 the prototype tool ""It's Your Code"" (IYC)4 has been developed as a plug-in for the Java development environment of the Eclipse platform"" - lack of URL; at the end of the paper there is also URL http://ag-kastens.upb.de/iyc which is not working)";No (OK);;Manual selection (well-known to the authors);Manual selection (well-known to the author);Random sampling;chosen randomly;Authors' assessment;Author's assessment;OK;OK;Detailed;None;Manual - criteria given;Random sampling;Authors 10.1016/j.jss.2019.110486;https://doi.org/10.1016/j.jss.2019.110486;2020;Journal;;A machine-learning based ensemble method for anti-patterns detection;Y;Y;Y;Y;Y;P;Y;6,5;Detailed;Detailed;Yes, with brief documentation;Yes;Yes, apps from https://github.com/ptidejteam/v5.2 , replication at https://github.com/antoineBarbez/SMAD/ ;scripts & data;Referred data set (combined Moha + Palomba);Referred data set (combined Moha + Palomba);Referred data set (combined Moha + Palomba) for GC, Advisors (HIST, InCode, JDeodorant) for FE;"GC: Referred data set (combined Moha + Palomba) with two constraints: (1) the full history of the system must be available through Git or SVN, and (2) the occurrences reported must be relevant, i.e., we kept only the systems for which we agreed with the occurrences tagged. After filtering, over the 15 systems available in these replication packages, they retained eight to construct the oracle. FE: Three advisors (HIST, InCode, and JDeodorant), adjusting their detection thresholds to produce a number of candidate per system proportional to the systems sizes.";Referred data set (combined Moha + Palomba) for GC, Authors, 9 MSc and PhD students, 2 software engineers for FE;"GC: Referred data set (combined Moha + Palomba) FE: three different groups of people manually checked each candidate of this set: (1) The authors of this paper, (2) nine M.Sc. and Ph.D. students, and (3) two software engineers.";Very minor difference in PQ4 (filtering...);PQ4 - agreed, does not affect results;Detailed;Scripts & data;Predefined;Predefined + Advisors;Predefined + Students & developers 10.1007/s11219-020-09498-y;https://doi.org/10.1007/s11219-020-09498-y;2020;Journal;;Code smell detection using multi-label classification approach;P;Y;P;P;Y;Y;Y;5,5;Detailed;Detailed;Data - yes, models - yes, scripts - no (N/A?);Data - yes, models - no, scripts - no;Data, models - yes, scripts - no ( https://figshare.com/articles/Detecting_Code_Smells_using_Machine_Learning_Techniques_Are_We_There_Yet_/5786631 );data;Referred data set (Fontana);Slightly modified referred data set (Fontana);Referred data set (Fontana);Slightly modified referred data set (Fontana);Referred data set (Fontana);Slightly modified referred data set (Fontana);DIFFerence: PQ2 (models yes vs no);PQ2 - agreed, misclassification during initial data acquisition. Does not affect results;Detailed;Data;Predefined;Predefined;Predefined 10.1145/3361242.3361257;https://doi.org/10.1145/3361242.3361257;2019;Conference;;Deep semantic-based feature envy identification;Y;Y;P;N;Y;N;N;3,5;Briefly;Moderate;No;No, (Data - reused data set by Fontana et al. https://essere.disco.unimib.it/machine-learning-for-code-smell-detection/);No (OK);;Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);DIFFerence: PQ1 (Briefly vs Moderate);"PQ1 - agreed, changed to ""Moderate""";Moderate;None;Predefined;Predefined;Predefined 10.1109/ISMSIT.2019.8932855;https://doi.org/10.1109/ISMSIT.2019.8932855;2019;Conference;;Comparison of Multi-Label Classification Algorithms for Code Smell Detection;Y;Y;P;N;N;Y;N;3,5;Briefly;Briefly;No;No;No (OK);;Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);OK;OK;Briefly;None;Predefined;Predefined;Predefined 10.1109/ICSME.2019.00021;https://doi.org/10.1109/ICSME.2019.00021;2019;Conference;;Deep Learning Anti-Patterns from Code Metrics History;Y;Y;P;Y;Y;Y;Y;6,5;Moderate;Moderate;Yes (https://github.com/antoineBarbez/CAME/ https://github.com/antoineBarbez/RepositoryMiner/), with brief documentation;Yes, with brief documentation;Yes - https://github.com/antoineBarbez/CAME/tree/master/experiments/tuning/;scripts & data;Manual selection (usage in prior studies);Manual selection;Merged data sets (Ptidej + mdipenta);Merged data sets (Ptidej + mdipenta);Merged data sets (Ptidej + mdipenta);Merged data sets (Ptidej + mdipenta);OK;OK;Moderate;Scripts & data;Manual - criteria given;Predefined;Predefined 10.1109/IJCNN.2019.8851854;https://doi.org/10.1109/IJCNN.2019.8851854;2019;Conference;;Deep Representation Learning for Code Smells Detection using Variational Auto-Encoder;Y;Y;P;Y;Y;P;P;5,5;Moderate;Moderate;No;No;No (OK);;Referred data set (Landfill);Referred data set (Landfill);Referred data set (Landfill);Referred data set (Landfill);Referred data set (Landfill);Referred data set (Landfill);OK;OK;Moderate;None;Predefined;Predefined;Predefined 10.1109/ICOIACT46704.2019.8938487;https://doi.org/10.1109/ICOIACT46704.2019.8938487;2019;Conference;;Software quality prediction using data mining techniques;P;P;N;N;Y;N;N;2;Briefly;Briefly;No;No;No (OK);;Referred corpus (Qualitas Corpus);Referred corpus (Qualitas Corpus);Advisors (iPlasma, JDeodorant, JMove);Advisors (iPlasma, JDeodorant, JMove);Unknown;Unknown;OK;OK;Briefly;None;Existing corpus subset;Advisors;Unknown 10.1109/MOBILESoft.2019.00025;https://doi.org/10.1109/MOBILESoft.2019.00025;2019;Conference;;Sniffing android code smells: An association rules mining-based approach;P;Y;P;N;N;P;Y;3,5;Moderate;Moderate;No;No;No (OK);;Unknown;Manual convenience selection of 48 opensource apps from F-Droid;Advisors (Unknown);Advisors (software quality metrics);Authors' assessment;Authors' assessment;minor differences in 2 columns: CC: PQ3 i CC: PQ4;"PQ3 - changed ""Unknown"" to ""Manual - no criteria"" PQ4 - does not affect results";Moderate;None;Manual - no criteria;Advisors;Authors 10.1109/IEMECONX.2019.8877008;https://doi.org/10.1109/IEMECONX.2019.8877008;2019;Conference;;An empirical framework for web service anti-pattern prediction using machine learning techniques;Y;Y;P;N;N;P;P;3,5;Briefly;Briefly;No;No;No (OK);;Referred data set (Ouni);Referred data set (Ouni);Referred data set (Ouni);Referred data set (Ouni);Referred data set (Ouni);Referred data set (Ouni);OK;OK;Briefly;None;Predefined;Predefined;Predefined 10.1007/s00521-019-04175-z;https://doi.org/10.1007/s00521-019-04175-z;2020;Journal;;SP-J48: a novel optimization and machine-learning-based approach for solving complex problems: special application in software engineering for detecting code smells;P;Y;P;P;Y;P;P;4,5;Detailed;Detailed;No;No;No (OK);;Manual selection (used earlier in literature - Moha?);Manual selection of 3 projects;Unknown (Moha?);Unknown;Unknown (Moha?);Unknown;OK;OK;Detailed;None;Manual - criteria given;Unknown;Unknown 10.1109/TSE.2019.2936376;https://doi.org/10.1109/TSE.2019.2936376;2019;Journal;;Deep Learning Based Code Smell Detection;P;Y;Y;P;Y;P;P;5;Detailed;Detailed;No;YES ( data and scripts vide https://github.com/liuhuigmail/DeepSmellDetection);Yes;scripts & data;Manual selection (open source, popular, well designed, successful, manually verified, SonarQube checks);Manual selection (open source, well-known, popular, high quality, manually verified, SonarQube checks);Smell-introducing refactoring;Smell-introducing refactoring;Smell-introducing refactoring;Smell-introducing refactoring;strong disagreement in PQ2;PQ2 - fixed during extra verification phase;Detailed;Scripts & data;Manual - criteria given;Bug seeding;Automated 10.1109/UBMK.2018.8566561;https://doi.org/10.1109/UBMK.2018.8566561;2018;Conference;;Comparison of Machine Learning Methods for Code Smell Detection Using Reduced Features;P;N;P;N;N;P;N;1,5;Very briefly (only refers to used methods);Briefly (almost lack!!!);No;No;No (OK);;Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);OK;OK;Briefly;None;Predefined;Predefined;Predefined 10.1145/3238147.3238166;https://doi.org/10.1145/3238147.3238166;2018;Conference;;Deep learning based feature envy detection;Y;Y;P;P;Y;P;Y;5,5;Briefly;Moderate;No;YES (data and scripts vide https://github.com/liuhuigmail/FeatureEnvy);Yes;scripts & data;Manual selection (open source, popular, well designed, successful, manually verified, SonarQube checks);Manual selection according to several subjective rules (e.g., Java apps, popularity, LOC etc.);Smell-introducing refactoring;Smell-introducing refactoring;Smell-introducing refactoring;Smell-introducing refactoring;minor difference in PQ1, major difference in PQ2;"PQ2 - fixed during extra verification phase PQ1 - agreed, error during initial data acquisition. Changed to ""Moderate""";Moderate;Scripts & data;Manual - criteria given;Bug seeding;Automated 10.1109/SANER.2018.8330265;https://doi.org/10.1109/SANER.2018.8330265;2018;Conference;;Keep it simple: Is deep learning good for linguistic smell detection?;Y;Y;Y;Y;Y;Y;Y;7;Moderate;Moderate;Data - yes, scripts - no;Data - yes, scripts - no;Data - yes, scripts - no;data;Manual selection (various domains);Manual selection (various domains);Advisors (LAPD);Advisors (LAPD);two evaluators;two evaluators (the kappa values range from 0.80 to 0.88);OK;OK;Moderate;Data;Manual - criteria given;Advisors;Unknown 10.1109/MLDS.2017.8;https://doi.org/10.1109/MLDS.2017.8;2017;Conference;;A Support Vector Machine Based Approach for Code Smell Detection;P;Y;P;N;Y;Y;P;4,5;Moderate;Moderate (close to Briefly);No;No;No (OK);;Manual selection (used earlier in literature - Moha?);"Manual convencience selection (""The systems which are considered in this study are Xerces v 2.7.0 and ArgoUML v0.19.8. These datasets are selected on the basis of various factors: Firstly, these systems are open source softwares and are freely available. Secondly, the researchers can replicate the study and other researchers who have used same systems allow comparison."")";Unknown (Moha?);Unknown;Unknown (Moha?);Unknown;OK;OK;Moderate;None;Manual - criteria given;Unknown;Unknown 10.14419/ijet.v7i2.27.14635;https://doi.org/10.14419/ijet.v7i2.27.14635;2018;Journal;;Design of testing framework for code smell detection (OOPS) using BFO algorithm;N;P;P;N;Y;N;N;2;Briefly;Briefly;No;No;No (OK);;Unknown;Unknown;Unknown;Unknown;Unknown;Unknown;OK;OK;Briefly;None;Unknown;Unknown;Unknown ;http://ceur-ws.org/Vol-2201/UYMS_2018_paper_80.pdf;2018;Conference;;Automatic detection of feature envy using machine learning techniques;P;P;P;N;N;P;N;2;Briefly;Briefly;No;No;No (OK);;Manual selection (no criteria given);"Manual convenience selection (no criteria given, ""We have conducted our experiments on two open source projects that are givenin Table II..."")";Unknown;Unknown;Unknown;Unknown;OK;OK;Briefly;None;Manual - no criteria;Unknown;Unknown 10.5220/0006709801370146;https://doi.org/10.5220/0006709801370146;2018;Conference;;A hybrid approach to detect code smells using deep learning;Y;Y;P;Y;P;Y;P;5,5;Briefly;Briefly;No;No;No (OK);;Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);OK;OK;Briefly;None;Predefined;Predefined;Predefined 10.1007/978-3-030-34706-2_8;https://doi.org/10.1007/978-3-030-34706-2_8;2019;Journal;;Code smell prediction employing machine learning meets emerging Java language constructs;Y;P;P;N;Y;Y;Y;5;Detailed;Detailed;No;Data - yes, scripts - no;Data - yes, scripts - no;data;Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);Referred data set (Fontana);OK (minor difference wrt. data was corrected via Verification: PQ2);PQ2 - fixed during extra verification phase;Detailed;Data;Predefined;Predefined;Predefined 10.1109/IEMECONX.2019.8877082;https://doi.org/10.1109/IEMECONX.2019.8877082;2019;Conference;;An empirical framework for code smell prediction using extreme learning machine∗;Y;P;Y;N;Y;;;3,5;Briefly;Briefly (close to Moderate);No;No;No - only metrics description;No - only metrics description;Unknown;Unknown;Unknown;Unknown;Unknown;Unknown;OK;OK;Briefly;None;Unknown;Unknown;Unknown 10.11591/ijece.v7i6.pp3613-3621;https://doi.org/10.11591/ijece.v7i6.pp3613-3621;2017;Journal;Kim, D.K.;Finding bad code smells with neural network models;P;P;P;N;N;P;N;2;Briefly;Briefly;No;No;No (OK);;Manual (no constraints given);Manual (no constraints/number of forks given);Advisors (Rules);Advisors (Rules);Automated (advisor accepted);Automated (advisor accepted);OK;OK;Briefly;None;Manual - no criteria;Advisors;Automated 10.1109/TENCON.2019.8929628;https://doi.org/10.1109/TENCON.2019.8929628;2019;Conference;Ananta Kumar Das, Shikhar Yadav, Subhasish Dhal;Detecting Code Smells using Deep Learning;N;P;N;N;N;N;N;0,5;Briefly;Briefly;No;No;No (OK);;Manual (no constraints given);Manual (no constraints given);Advisors (iPlasma);Advisor (iPlasma);Automated (advisor accepted);Automated (advisor accepted);OK;OK;Briefly;None;Manual - no criteria;Advisors;Automated 10.1016/j.knosys.2017.04.014;"https://doi.org/10.1016/j.knosys.2017.04.014 ";2017;Journal;Francesca Arcelli Fontana, Marco Zanoni;Code smell severity classification using machine learning techniques;Y;Y;P;Y;Y;P;P;5,5;Detailed;Detailed;Yes (data);"Claimed yes (data), but link does not work ""404 Not Found""";Claimed yes, but page dead: http://essere.disco.unimib.it/wiki/research/mlcsd;Data not accessible;Existing corpus (Qualitas Corpus) without non-compilable items;Part of Qualitas Corpus (76 of 111 systems);Advisors (several);Advisors (several);3 MSc students;Three MsC students performed the labelling process;OK;OK;Detailed;Unresolvable;Existing corpus subset;Advisors;Students