index,title,authors,year,abstract,link,relevant,queries,citations,methods,sources,contributions,categories,targets,venue,mapping-study-result,macro-areas,best,backward,subcategories,covered,reference_index,doi,type,keywords,focuses
0,Machine Learning for Performance Prediction of Spark Cloud Applications,"['A. Maros', ' F. Murai', ' A. P. Couto da Silva', ' J. M. Almeida', ' M. Lattuada', ' E. Gianniti', ' M. Hosseini', ' D. Ardagna']",2019,"Big data applications and analytics are employed in many sectors for a variety of goals: improving customers satisfaction, predicting market behavior or improving processes in public health. These applications consist of complex software stacks that are often run on cloud systems. Predicting execution times is important for estimating the cost of cloud services and for effectively managing the underlying resources at runtime. Machine Learning (ML), providing black box solutions to model the relationship between application performance and system configuration without requiring in-detail knowledge of the system, has become a popular way of predicting the performance of big data applications. We investigate the cost-benefits of using supervised ML models for predicting the performance of applications on Spark, one of today's most widely used frameworks for big data analysis. We compare our approach with Ernest (an ML-based technique proposed in the literature by the Spark inventors) on a range of scenarios, application workloads, and cloud system configurations. Our experiments show that Ernest can accurately estimate the performance of very regular applications, but it fails when applications exhibit more irregular patterns and/or when extrapolating on bigger data set sizes. Results show that our models match or exceed Ernest's performance, sometimes enabling us to reduce the prediction error from 126-187% to only 5-19%.",https://ieeexplore.ieee.org/document/8814514,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 5}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 200}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 42}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 92}]",1.0,"['multilayer-perceptron', 'random-forest', 'linear-regression', 'decision-tree']",['kpis'],['novel-use'],['workload-prediction'],['spark'],2019 IEEE 12th International Conference on Cloud Computing (CLOUD),True,['resource-provisioning'],,,,,,,,,
1,A Self-Learning Scheduling in Cloud Software Defined Block Storage,"['B. Ravandi', ' I. Papapanagiotou']",2017,"Software Defined Storage (SDS) separates the control layer from the data layer allowing the automation of data management and deployment of Commercial Off-The-Shelf (COTS) storage media rather than expensive traditional hardware-based solutions. Cloud block storage services lack an SDS framework that allows customization of block storage, policy enforcement, automate provisioning, and storage management. SDS decreases the human intervention and improves the resource utilization. Moreover, SDS allows cloud tenants to define customized functionalities based on their needs with guaranteed performance and high availability that meets Service Level Agreements (SLAs). However, maintaining SLAs requirements in cloud block storage is challenging due to the storage cluster features, the workload interference, the workload characteristics and other indirect related latent variables. To address the mentioned issues, cloud providers often over-provision the storage resources. Moving towards SDS, we initiate a framework for cloud block storage as an active storage system. Our framework provides customization of block storage services and optimized scheduling decisions based on the workload characteristics and performance of the underlying data layer leveraging a self-learning scheduler. The proposed scheduler treats the storage backend nodes as a black box and requires zero knowledge of their internal states. We showcase a practical application of the proposed scheduler in our private OpenStack deployment.",https://ieeexplore.ieee.org/document/8030616,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 16}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 979}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 21}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 666}]",0.0,,,,['scheduling'],,2017 IEEE 10th International Conference on Cloud Computing (CLOUD),True,['resource-provisioning'],,,,,,,,,
2,Using of Machine Learning into Cloud Environment (A Survey): Managing and Scheduling of Resources in Cloud Systems,"['E. Hormozi', ' H. Hormozi', ' M. K. Akbari', ' M. S. Javan']",2012,"Cloud computing is a model for delivering information technology services in which resources are retrieved from the internet through web-based tools and applications, rather than a direct connection to a server. Many companies, such as Amazon, Google, Microsoft and so on, are developing cloud computing systems and enhancing their services to provide for a larger amount of users. This technology holds a vast scope of using the various aspects of machine learning for increased performance and solving some of the challenges in front of the research community. in this survey, we investigate the effects using the concepts of machine learning on cloud environments, e.g. automated resource allocation mechanism, intelligently managing and allocating resources with SmartSLA, resources scheduling, etc.",https://ieeexplore.ieee.org/document/6362996,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 22}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 764}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 184}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 78}]",17.0,,,,"['scheduling', 'resource-consolidation']",,"2012 Seventh International Conference on P2P, Parallel, Grid, Cloud and Internet Computing",True,['resource-provisioning'],True,,,,,,,,
3,StackInsights: Cognitive Learning for Hybrid Cloud Readiness,"['M. Qiao', ' L. Bathen', ' S. Genot', ' S. Lee', ' R. Routray']",2018,"Hybrid cloud is an integrated cloud computing environment utilizing a mix of public cloud, private cloud, and on-premise IT infrastructures. Workload awareness, defined as a detailed full range understanding of each individual workload, is essential in implementing the hybrid cloud. While it is critical to perform an accurate analysis to determine which workloads are appropriate for on-premise deployment versus which workloads can be migrated to a cloud off-premise, the assessment is mainly performed by rule or policy based approaches. In this paper, we introduce StackInsights, a novel cognitive system to automatically analyze and predict the cloud readiness of workloads for an enterprise. Our system harnesses the critical metrics across the entire stack: (1) infrastructure metrics, (2) data relevance metrics, and (3) application taxonomy, to identify workloads that have characteristics of (a) low sensitivity with respect to business security, criticality and compliance, and (b) low response time requirements and access patterns. Since the capture of the data relevance metrics involves an intrusive and in-depth scanning of the content of storage objects, a machine learning model is applied to perform the business relevance classification by learning from the meta level metrics harnessed across stack. In contrast to traditional methods, StackInsights significantly reduces the total time for hybrid cloud readiness assessment by orders of magnitude.",https://ieeexplore.ieee.org/document/8457808,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 32}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 634}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 907}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 443}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 39}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 667}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 15}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 756}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 637}, {'database': 'arxiv', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 44}, {'database': 'arxiv', 'search_string': ""'classification' AND ('cloud')"", 'index': 56}, {'database': 'arxiv', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1278}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 57}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 71}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 385}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 95}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 80}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1015}]",0.0,"['clustering', 'random-forest']","['host-metrics', 'topology']",,['resource-consolidation'],['cloud'],2018 IEEE 11th International Conference on Cloud Computing (CLOUD),True,['resource-provisioning'],,True,,,,,,,
4,Detecting Anomalous Behavior of Black-Box Services Modeled with Distance-Based Online Clustering,"['A. Gulenko', ' F. Schmidt', ' A. Acker', ' M. Wallschläger', ' O. Kao', ' F. Liu']",2018,"Reliable deployment of services is especially challenging in virtualized infrastructures, where the deep tech-nological stack and the multitude of components necessitate automatic anomaly detection and remediation mechanisms. Traditional monitoring solutions observe the system and generate alarms when the collected metrics exceed predefined thresholds. The fixed thresholds rely on expert knowledge and can lead to numerous false alarms, while abnormal behavior that spans over multiple metrics, components, or system layers, may not be detected. We propose to use an unsupervised online clustering algorithm to create a model of the normal behavior of each monitored component with minimal human interaction and no impact on the monitored system. When an anomaly is detected, a human administrator or automatic remediation system can subsequently revert the component into a normal state. An experimental evaluation resulted in a high accuracy of our approach, indicating that it is suitable for anomaly detection in productive systems.",https://ieeexplore.ieee.org/document/8457902,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 41}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 143}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 346}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 21}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 277}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 114}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 591}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('remediation' OR 'recovery')"", 'index': 50}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 86}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1316}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 471}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('remediation' OR 'recovery')"", 'index': 208}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 291}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 64}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 7}]",4.0,['clustering'],"['host-metrics', 'network-metrics']",['novel-use'],['failure-detection'],"['vm', 'openstack']",2018 IEEE 11th International Conference on Cloud Computing (CLOUD),True,['failure-management'],,True,['anomaly-detection'],,,,,,
5,FailureSim: A System for Predicting Hardware Failures in Cloud Data Centers Using Neural Networks,"['N. A. Davis', ' A. Rezgui', ' H. Soliman', ' S. Manzanares', ' M. Coates']",2017,"Hardware failures in cloud data centers may cause substantial losses to cloud providers and cloud users. Therefore, the ability to accurately predict when failures occur is of paramount importance. In this paper, we present FailureSim, a simulator based on CloudSim that supports failure prediction. FailureSim obtains performance related information from the cloud and classifies the status of the hardware using a neural network. Performance information is read from hosting hardware and stored in a variable length windowing vector. At specified stages, various aggregation methods are applied to the windowing vector to obtain a single vector that is fed as input to the trained classification algorithm. Using conservative host failure behavior models, FailureSim was able to successfully predict host failure in the cloud with roughly 89% accuracy.",https://ieeexplore.ieee.org/document/8030632,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 44}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 27}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 265}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 97}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 537}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 32}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 134}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 667}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 104}]",8.0,"['multilayer-perceptron', 'rnn']","['host-metrics', 'network-metrics']",,['failure-prediction'],"['hardware', 'vm']",2017 IEEE 10th International Conference on Cloud Computing (CLOUD),True,['failure-management'],True,True,['hardware-failure-prediction'],True,39.0,,,,
6,Improving Energy Efficiency in NFV Clouds with Machine Learning,"['L. M. Moreira Zorello', ' M. G. Torres Vieira', ' R. A. Girani Tejos', ' M. A. Torres Rojas', ' C. Meirosu', ' T. C. Melo de Brito Carvalho']",2018,"Widespread deployments of Network Function Virtualization (NFV) technology will replace many physical appliances in telecommunication networks with software executed on cloud platforms. Setting compute servers continuously to high-performance operating modes is a common NFV approach for achieving predictable operations. However, this has the effect that large amounts of energy are consumed even when little traffic needs to be forwarded. The Dynamic Voltage-Frequency Scaling (DVFS) technology available in Intel processors is a known option for adapting the power consumption to the workload, but it is not optimized for network traffic processing workloads. We developed a novel control method for DVFS, based observing the ongoing traffic and online predictions using machine learning. Our results show that we can save up to 27% compared to commodity DVFS, even when including the computational overhead of machine learning.",https://ieeexplore.ieee.org/document/8457866,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 47}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 994}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 398}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 112}]",3.0,,,,['power-management'],"['vm', 'network']",2018 IEEE 11th International Conference on Cloud Computing (CLOUD),True,['resource-provisioning'],,,,,,,,,
7,Migration scheme based machine learning for QoS in cloud computing: Survey and research challenges,"['A. Son', ' E. Huh', ' S. Na', ' P. Lee']",2017,"VM migration has become a hot topic in the Cloud Dater Centers (CDCs). Numerous VM migration schemes are proposed for QoS aimed to improve various metrics affecting the CDCs. Also, it combines a Machine learning approach with modeling and predicting in CDC. Migration scheme through prediction based metric can greatly improve the physical machines resource utilization. Also, effective VM migration scheme can reduce the power consumption and time of the data centers. Thus, it needs to consider metrics which may impact the migration performance and energy efficiency. In this paper, we summarize and classify previous approaches of migration in CDCs. Furthermore, we conclude with a discussion of research problems in this area. In the future work, we will study on live migration mechanism to improve the live migration performance and energy efficiency in the variety of CDCs.",https://ieeexplore.ieee.org/document/8320748,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 52}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1997}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 106}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 290}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 396}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 12}]",0.0,,,['survey'],['resource-consolidation'],,2017 4th International Conference on Computer Applications and Information Processing Technology (CAIPT),True,['resource-provisioning'],,,,,,,,,
8,A Survey of Machine Learning Applications for Energy-Efficient Resource Management in Cloud Computing Environments,['M. Demirci'],2015,"Ensuring energy efficiency in data centers is a crucial objective in modern cloud computing because it reduces operating costs and complies with the goals of green computing. Researchers strive to develop optimal policies for resource management in the cloud, which has many components such as virtual machine placement, task scheduling, workload consolidation, and so on. Machine learning has a major role to play in these efforts. In this paper, we provide a detailed survey of recent works in the literature which have employed machine learning (ML) to offer solutions for energy efficiency in cloud computing environments. We also present a comparative classification of the proposed methods. Furthermore, we enrich this survey by studying non-ML proposals to energy conservation in data centers, and also how ML has been applied towards other objectives in the cloud.",https://ieeexplore.ieee.org/document/7424481,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 53}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 316}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 358}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1079}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 1098}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 10}]",28.0,,,['survey'],"['power-management', 'resource-consolidation']",,2015 IEEE 14th International Conference on Machine Learning and Applications (ICMLA),True,['resource-provisioning'],True,,,,,,,,
9,An Approach to Failure Prediction in a Cloud Based Environment,"['H. Adamu', ' B. Mohammed', ' A. B. Maina', ' A. Cullen', ' H. Ugail', ' I. Awan']",2017,"Failure in cloud system is defined as an even that occurs when the delivered service deviates from the correct intended service. As the cloud computing systems continue to grow in scale and complexity, there is an urgent need for cloud service providers (CSP) to guarantee a reliable on-demand resource to their customers in the presence of faults thereby fulfilling their service level agreement (SLA). Component failures in cloud systems are very familiar phenomena. However, large cloud service providers' data centers should be designed to provide a certain level of availability to the business system. Infrastructure-as-a-service (Iaas) cloud delivery model presents computational resources (CPU and memory), storage resources and networking capacity that ensures high availability in the presence of such failures. The data in-production-faults recorded within a 2 years period has been studied and analyzed from the National Energy Research Scientific computing center (NERSC). Using the real-time data collected from the Computer Failure Data Repository (CFDR), this paper presents the performance of two machine learning (ML) algorithms, Linear Regression (LR) Model and Support Vector Machine (SVM) with a Linear Gaussian kernel for predicting hardware failures in a real-time cloud environment to improve system availability. The performance of the two algorithms have been rigorously evaluated using K-folds cross-validation technique. Furthermore, steps and procedure for future studies has been presented. This research will aid computer hardware companies and cloud service providers (CSP) in designing a reliable fault-tolerant system by providing a better device selection, thereby improving system availability and minimizing unscheduled system downtime.",https://ieeexplore.ieee.org/document/8114482,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 54}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 57}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 128}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 845}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 58}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 119}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('remediation' OR 'recovery')"", 'index': 101}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 289}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 145}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 72}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 41}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 46}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 105}, {'database': 'IEEE', 'search_string': ""'regression' AND ('remediation' OR 'recovery')"", 'index': 279}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 150}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 968}]",4.0,"['linear-regression', 'support-vector-machine']",['host-metrics'],['novel-use'],['failure-prediction'],"['cpu', 'memory', 'hard-drive']",2017 IEEE 5th International Conference on Future Internet of Things and Cloud (FiCloud),True,['failure-management'],,True,['hardware-failure-prediction'],,,,,,
10,Workload and Resource Aware Proactive Auto-scaler for PaaS Cloud,"['R. S. Shariffdeen', ' D. T. S. P. Munasinghe', ' H. S. Bhathiya', ' U. K. J. U. Bandara', ' H. M. N. D. Bandara']",2016,"Elasticity is a key feature in Cloud Computing where virtualized resources are provisioned and de-provisioned via auto-scaling. However, auto-scaling in most Platform-as-a-Service (PaaS) systems is based on reactive, threshold-driven approaches. Such systems are incapable of catering to rapidly varying workloads, unless the associated thresholds are sufficiently low. Alternatively, maintaining low thresholds leads to resource over-provisioning under relatively stable workloads. Moreover, thresholds are not a good indication of QoS compliance, which is a key performance indicator of a cloud application. Hence, it is nontrivial to determine an optimum threshold while minimizing costs and meeting QoS demands. We propose inteliScaler, a proactive and cost-aware auto-scaling solution to address these issues by combining a predictive model, cost model, and a smart killing feature. An ensemble workload prediction mechanism is introduced based on time series and machine learning techniques for making accurate predictions on drastically different workload patterns. Utility of the solution is demonstrated using both simulations and empirical evaluations using Apache Stratos PaaS (deployed on the AWS EC2), as well as RUBiS and real-world workload traces. Results show significant QoS improvements and cost reductions by inteliScaler compared to a typical reactive and threshold-based PaaS auto-scaling solution.",https://ieeexplore.ieee.org/document/7820029,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 66}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 887}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 422}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 88}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 82}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 64}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1487}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 245}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 595}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 248}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 305}]",3.0,"['multilayer-perceptron', 'autoregression']",['traces'],['novel-use'],"['resource-consolidation', 'workload-prediction']","['rubis', 'apache']",2016 IEEE 9th International Conference on Cloud Computing (CLOUD),True,['resource-provisioning'],,True,,,,,,,
11,A Practical Approach to Hard Disk Failure Prediction in Cloud Platforms: Big Data Model for Failure Management in Datacenters,"['S. Ganguly', ' A. Consul', ' A. Khan', ' B. Bussone', ' J. Richards', ' A. Miguel']",2016,"Large scale cloud platforms can benefit from a service that runs a machine learning model to predict disk drive failures. Unlike previous studies in this space, we have combined multiple data inputs for the model and obtained a better model performance compared to earlier published models. In this paper we explain how we developed and deployed the predictive model in a large scale cloud service. To build the model, we used a combination of two open data sources - Self-Monitoring, Analysis and Reporting technology (S.M.A.R.T or SMART) data and Windows performance counters. The nature of both these data sources is different and complex. The paper provides unique ways of parsing and transforming the data to make it most suited for a classification problem. Trails with different machine learning (ML) and statistical modeling techniques led us to the best performing two-stage ensemble model. We implemented this model to be configurable such that it could be deployed on large scale distributed cloud management systems and iterated on with minimal code impact. We provide a glimpse of the complex cloud hardware ecosystem and how a predictive model would impact such an ecosystem. Although our study focused on hard disk drives, we believe a similar modeling approach can apply to other hardware components as well. A successfully executed hard disk failure prediction model can pre-empt negative impact to client workloads and improve the economics of running a large scale cloud service. We provide the details of our model as a possible template for future extensions and improvements towards building more robust hardware fault prediction services. Finally we give a staged approach to operationalizing the model in large scale cloud systems.",https://ieeexplore.ieee.org/document/7474362,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 74}, {'database': 'IEEE', 'search_string': ""'classification' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 35}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 8}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 78}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 62}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 1572}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 517}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 476}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 19}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 74}, {'database': 'IEEE', 'search_string': ""'classification' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 45}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 12}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 58}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 432}]",12.0,"['entropy-selection', 'logistic-regression', 'decision-tree']",['host-metrics'],['new-method'],['failure-prediction'],['hard-drive'],2016 IEEE Second International Conference on Big Data Computing Service and Applications (BigDataService),True,['failure-management'],False,,['hardware-failure-prediction'],,,,,,
12,iCSI: A Cloud Garbage VM Collector for Addressing Inactive VMs with Machine Learning,"['I. K. Kim', ' S. Zeng', ' C. Young', ' J. Hwang', ' M. Humphrey']",2017,"According to a recent study, 30% of VMs in private cloud data centers are ""comatose"", in part because there is generally no strong incentive for their human owners to delete them at an appropriate time. These inactive VMs are still scheduled and executed on physical cloud resources, taking valuable access away from productive VMs. In an extreme, cloud infrastructure may deny legitimate requests for new VMs because capacity limits have been hit. It is not sufficient for cloud infrastructure to identify such inactive VMs by monitoring resource utilization (e.g., CPU utilization) - e.g., management processes (e.g. virus-scan, software update) on inactive VMs often consume high CPU and memory resources, and active VMs with lightweight jobs (e.g. text editing) show almost zero resource utilization. To properly detect and address such inactive VMs, we present iCSI: a cloud garbage VM collector to improve resource utilization and cost efficiency of enterprise data centers. iCSI includes three main components, a lightweight data collector, a VM identification model and a recommendation engine. The data collector periodically gathers primitive information from VMs. The identification model infers the purpose of a VM from the data collection and extracts the most relevant features associated with the purpose. The recommendation engine offers proper actions to end users i.e., suspending or resizing VMs. In this prototype phase, iCSI is deployed into multiple data centers in IBM and manages more than 750 production VMs. iCSI achieves 20% better accuracy (90%) in identifying active/inactive VMs compared with state-of-the-art methods. With recommendations to end users, our estimation results show that iCSI can improve internal cost efficiency with 23% and resource utilization more than 45%.",https://ieeexplore.ieee.org/document/7923783,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 103}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 9}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1325}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 80}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 341}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 37}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 82}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 181}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 451}]",3.0,[],[],,"['resource-consolidation', 'failure-detection']","['vm', 'cloud']",2017 IEEE International Conference on Cloud Engineering (IC2E),True,"['resource-provisioning', 'failure-management']",,True,['anomaly-detection'],,,,,,
13,Reverse Engineering Technique (RET) to Predict Resource Allocation in a Google Cloud System,"['B. R. Ray', ' S. Chowdhury']",2018,"This paper presents a reverse engineering machine learning technique for resource allocation in cloud computing system. Efficient and timely resource allocation is a crucial task for complex operations in a large scale distributed system like cloud computing. Furthermore, to support Service Level Agreement (SLA) like priority, latency, and efficiency, the resource provisioning should be highly influenced by SLA requirements of the system. Therefore in this paper, we propose the Reverse Engineering Technique (RET) which highly influence by priority to improve resource allocation accuracy. The paper used neural network based deep learning and Levenberg-Marquardt training algorithm for resource allocation prediction. The dataset of Google cloud computing system, which is publicly available dataset for research, is used to test the proposed RET. Our experiment shows that the proposed technique improves resource provisioning accuracy for cloud based systems.",https://ieeexplore.ieee.org/document/8442524,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 106}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 507}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 59}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 65}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 350}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 432}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 383}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 96}]",2.0,['multilayer-perceptron'],,,['resource-consolidation'],['cloud'],"2018 8th International Conference on Cloud Computing, Data Science & Engineering (Confluence)",True,['resource-provisioning'],,True,,,,,,,
14,Machine Learning-Based Framework for Resource Management and Modelling for Video Analytic in Cloud-Based Hadoop Environment,"['M. Al-Rawahi', ' E. A. Edirisinghe', ' T. Jeyarajan']",2016,"Hadoop framework has recently been adapted for use by the video analytics community for intensive, distributed video processing, storage. However, the challenge is to estimate the required amount of resources to be used in such an environment to fulfil the requirements of a user with requirements constraints. Therefore, it is important to understand how to model the performance of a Hadoop based implementation of video analytic applications in terms of meeting their performance goals. In this paper we propose the use of machine learning approachs in modelling the execution time based on the given resources. The prediction is based on parameters related to typical video analytic applications such as video file characteristics (e, g, resolution, file size, frame rate. etc.), cluster resource consumption,, Hadoop configuration values (reducer slots, tasks). The investigation carried out compares the use of different machine learning classifiers with regard to their best obtainable performance accuracies, show that a decision based model (M5P) outperforms a Linear Regression model, while the Ensemble Classifier, Bagging, out-performs these standard single classifiers. The research conducted bridges an existing research gap in video analytic-related performance predictions, whereby current research focuses on different application types, is largely limited to using standard learning algorithms such as SVM, Linear Regression, Multilayer Perceptron (MLP).",https://ieeexplore.ieee.org/document/7816924,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 134}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 218}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1334}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 244}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 860}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 797}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 349}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 617}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 287}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 234}]",1.0,"['decision-tree', 'linear-regression']","['host-metrics', 'configuration', 'requests']","['novel-use', 'comparison']",['resource-consolidation'],['cloud'],"2016 Intl IEEE Conferences on Ubiquitous Intelligence & Computing, Advanced and Trusted Computing, Scalable Computing and Communications, Cloud and Big Data Computing, Internet of People, and Smart World Congress (UIC/ATC/ScalCom/CBDCom/IoP/SmartWorld)",True,['resource-provisioning'],,True,,,,,,,
15,Online QoS Modeling in the Cloud: A Hybrid and Adaptive Multi-learners Approach,"['T. Chen', ' R. Bahsoon', ' X. Yao']",2014,"Given the on-demand nature of cloud computing, managing cloud-based services requires accurate modeling for the correlation between their Quality of Service (QoS) and cloud configurations/resources. The resulted models need to cope with the dynamic fluctuation of QoS sensitivity and interference. However, existing QoS modeling in the cloud are limited in terms of both accuracy and applicability due to their static and semi-dynamic nature. In this paper, we present a fully dynamic multi-learners approach for automated and online QoS modeling in the cloud. We contribute to a hybrid learners solution, which improves accuracy while keeping model complexity adequate. To determine the inputs of QoS model at runtime, we partition the inputs space into two sub-spaces, each of which applies different symmetric uncertainty based selection techniques, and we then combine the sub-spaces results. The learners are also adaptive, they simultaneously allow several machine learning algorithms to model QoS function and dynamically select the best model for prediction on the fly. We experimentally evaluate our models using RUBiS benchmark and realistic FIFA 98 workload. The results show that our multi-learners approach is more accurate and effective in contrast to the other state-of-the-art approaches.",https://ieeexplore.ieee.org/document/7027509,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 138}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 418}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 324}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 177}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 345}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud')"", 'index': 862}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 68}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 205}]",6.0,,"['host-metrics', 'kpis']",,['resource-consolidation'],"['cloud', 'rubis', 'vm']",2014 IEEE/ACM 7th International Conference on Utility and Cloud Computing,True,['resource-provisioning'],,,,,,10.1109/UCC.2014.42,,,
16,"A Host-Agnostic, Supervised Machine Learning Approach to Automated Overload Detection in Virtual Machine Workloads","['E. M. Dow', ' J. N. Matthews']",2017,"This paper evaluates a mechanism for applying machine learning (ML) to identify over-constrained IaaS virtual machines (VMs). Herein, over-constrained VMs are defined as those who are not given sufficient system resources to meet their workload specific objective functions. To validate our approach, a variety of workload-specific benchmarks inspired by common Infrastructure-as-a-Service (IaaS) cloud workloads were used. Workloads were run while regularly sampling VM resource consumption features exposed by the hypervisor. Datasets were curated into nominal or over-constrained and used to train ML classifiers to determine VM over-constraint rules based on one-time workload analysis. Rules learned on one host are transferred with the VM to other host environments to determine portability. Key contributions of this work include: demonstrating which VM resource consumption metrics (features) prove most relevant to learned decision trees in this context, and a technique required to generalize this approach across hosts while limiting required up front training expenditure to a single VM and host. Other contributions include a rigorous explanation of the differences in learned rulesets as a function of feature sampling rates, and an analysis of the differences in learned rulesets as a function of workload variation. Feature correlation matrices and their corresponding generated rule sets demonstrate individual features comprising rule sets tend to show low cross-correlation (below 0.4) while no individual feature shows high direct correlation with classification. Our system achieves workload-specific error percentages below 2.4% with a mean error across workloads of 1.43% (and strong false negative bias) for a variety of synthetic, representative, cloud workloads tested.",https://ieeexplore.ieee.org/document/8118413,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 139}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 234}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1950}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 331}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 229}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1344}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 175}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 589}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 881}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 362}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 447}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 184}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1458}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 212}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 518}]",1.0,['rule-mining'],['host-metrics'],,"['workload-prediction', 'failure-detection']","['vm', 'cloud']",2017 IEEE International Conference on Smart Cloud (SmartCloud),True,['resource-provisioning'],,True,,,,,,,
17,Intelligent Cloud Resource Management with Deep Reinforcement Learning,"['Y. Zhang', ' J. Yao', ' H. Guan']",2017,"The cloud provides low-cost and flexible IT resources (hardware and software) across the Internet. As more cloud providers seek to drive greater business outcomes and the environments of the cloud become more complicated, it is evident that the era of the intelligent cloud has arrived. The intelligent cloud faces several challenges, including optimizing the economic cloud service configuration and adaptively allocating resources. In particular, there is a growing trend toward using machine learning to improve the intelligence of cloud management. This article discusses an architecture of intelligent cloud resource management with deep reinforcement learning. The deep reinforcement learning makes clouds automatically and efficiently negotiate the most appropriate configuration, directly from complicated cloud environments. Finally, we give an example to evaluate and conclude the remarkable ability of the intelligent cloud with deep reinforcement learning.",https://ieeexplore.ieee.org/document/8260812,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 144}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 32}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 656}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 20}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 199}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 465}]",10.0,['reinforcement-learning'],,,['resource-consolidation'],,IEEE Cloud Computing,True,['resource-provisioning'],,,,,,,,,
18,ML-NA: A Machine Learning Based Node Performance Analyzer Utilizing Straggler Statistics,"['X. Ouyang', ' C. Wang', ' R. Yang', ' G. Yang', ' P. Townend', ' J. Xu']",2017,"Current Cloud clusters often consist of heterogeneous machine nodes, which can trigger performance challenges such as the task straggler problem, whereby a small subset of parallel tasks running abnormally slower than the other sibling ones. The straggler problem leads to extended job response and deteriorates system throughput. Poor performance nodes are more likely to engender stragglers, and can undermine straggler mitigation effectiveness. For example, as the dominant mechanism for straggler alleviation, speculative execution functions by creating redundant task replicas on other machine nodes as soon as a straggler is detected. When speculative copies are assigned onto the poor performance nodes, it is hard for them to catch up with the stragglers compared to replicas run on fast nodes. And due to the fact that the performance heterogeneity is caused not only by static attribute variations such as physical capacity, but also dynamic characteristic uctuations such as contention level, analyzing node performance is important yet challenging. In this paper we develop ML-NA, a Machine Learning based Node performance Analyzer. By leveraging historical parallel tasks execution log data, ML-NA classies cluster nodes into different categories and predicts their performance in the near future as a scheduling guide to improve speculation effectiveness and minimize task straggler generation. We consider MapReduce as a representative framework to perform our analysis, and use the published OpenCloud trace as a case study to train and to evaluate our model. Results show that ML-NA can predict node performance categories with an average accuracy up to 92.86%.",https://ieeexplore.ieee.org/document/8368350,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 145}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1103}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1041}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 1631}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 54}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1753}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 674}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 11}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 56}]",1.0,"['clustering', 'naive-bayes']","['kpis', 'host-metrics']",['comparison'],"['resource-consolidation', 'failure-detection']",['node'],2017 IEEE 23rd International Conference on Parallel and Distributed Systems (ICPADS),True,['resource-provisioning'],,,,,,,,,
19,Online machine learning for cloud resource provisioning of microservice backend systems,"['H. Alipour', ' Y. Liu']",2017,"Microservices are bundled and generating traffic on the backend systems that need to scale on demand. When microservices generate variant and unexpected, the challenge is to classify the workload on the backend systems and adjust the scaling policy to reflect the resource demand timely and accurately. In this paper, we propose a microservice architecture that encapsulates functions of monitoring metrics and learning workload pattern. Then this service architecture is used to predict the future workload for decision making on resource provisioning. We deploy two machine learning algorithms and predict the resource demand of the backend systems of microservices emulated by a Netflix workload benchmark application. This service architecture presents an integrated solution of implementing self-managing cloud data services under variant workload.",https://ieeexplore.ieee.org/document/8258201,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 147}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1259}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1924}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 995}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1338}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1017}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 320}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 184}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 890}]",1.0,,,,,,2017 IEEE International Conference on Big Data (Big Data),True,['resource-provisioning'],,True,,,,,,,
20,Managing Uncertainty in Autonomic Cloud Elasticity Controllers,"['P. Jamshidi', ' C. Pahl', ' N. C. Mendonça']",2016,"Elasticity allows a cloud system to maintain an optimal user experience by automatically acquiring and releasing resources. Autoscaling-adding or removing resources automatically on the fly-involves specifying threshold-based rules to implement elasticity policies. However, the elasticity rules must be specified through quantitative values, which requires cloud resource management knowledge and expertise. Furthermore, existing approaches don't explicitly deal with uncertainty in cloud-based software, where noise and unexpected events are common. The authors propose a control-theoretic approach that manages the behavior of a cloud environment as a dynamic system. They integrate a fuzzy cloud controller with an online learning mechanism, putting forward a framework that takes the human out of the dynamic adaptation loop and can cope with various sources of uncertainty in the cloud.",https://ieeexplore.ieee.org/document/7503491,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 150}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 624}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 218}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 502}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 435}]",20.0,,,,['resource-consolidation'],,IEEE Cloud Computing,True,['resource-provisioning'],,,,,,,,,
21,Cloud antivirus cost model using machine learning,"['A. A. Hamzah', ' S. M. Khattab', ' S. S. El-Gamal']",2014,"An important cloud computing is a new generation of computing and is based on virtualization technology. More and more applications are being deployed in cloud environments. Malware detection or antivirus software has been recently provided as a service in the cloud. A cloud antivirus provider hosts a number of virtual machines each running the same or different antivirus engines on potentially different sets of workloads (files). From the provider's perspective, the problem of optimally allocating physical resources to these virtual machines is crucial to the efficiency of the infrastructure. We propose a search-based optimization approach for solving the resource allocation problem in cloud-based antivirus deployments. An elaborate cost model of the file scanning process in antivirus programs is instrumental to the proposed approach. The general architecture is presented and discussed, and a preliminary experimental investigation into the antivirus cost model is described. The cost model depends on many factors, such as total file size, size of code segment, and count and type of embedded files within the executable. However, not a single parameter of these can be reliably used alone to predict file scanning time. Thus, a machine-learning approach that combines all these parameters as features is used to build a classifier for antivirus file scanning time. The best results we obtained were using the Decision Tree classifier. The highest F-measure value was 0.91, the highest F-measure value using logitboost was 0.87, the highest F-measure value using support vector machine was 0.85 and the highest F-measure value using naïve Bayes was 0.82. We evaluated the accuracy of the classification model versus linear regression model using the Root Mean Square (RMS) measure. We found that the classification model is more accurate than linear regression model, whereas the values average of RMS were 0.988 second and 2.44 second for classification model and linear regression model, respectively.",https://ieeexplore.ieee.org/document/7036708,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 152}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 108}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 933}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 149}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 841}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 809}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 258}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 409}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 247}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 639}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 408}]",296.0,['search'],,,['resource-consolidation'],['antivirus'],2014 9th International Conference on Informatics and Systems,True,['resource-provisioning'],,True,,,,,,,
22,FIM: Performance Prediction for Parallel Computation in Iterative Data Processing Applications,"['J. Bhimani', ' N. Mi', ' M. Leeser', ' Z. Yang']",2017,"Predicting performance of an application running on high performance computing (HPC) platforms in a cloud environment is increasingly becoming important because of its influence on development time and resource management. However, predicting the performance with respect to parallel processes is complex for iterative, multi-stage applications. This research proposes a performance approximation approach FiM to model the computing performance of iterative, multi-stage applications running on a master-compute framework. FiM consists of two key components that are coupled with each other: 1) Stochastic Markov Model to capture non-deterministic runtime that often depends on parallel resources, e.g., number of processes. 2) Machine Learning Model that extrapolates the parameters for calibrating our Markov model when we have changes in application parameters such as dataset. Our new modeling approach considers different design choices along multiple dimensions, namely (i) process level parallelism, (ii) distribution of cores on multi-core processors in cloud computing, (iii) application related parameters, and (iv) characteristics of datasets. The major contribution of our prediction approach is that FiM is able to provide an accurate prediction of parallel computation time for the datasets which have much larger size than that of the training datasets. Such calculation prediction provides data analysts a useful insight of optimal configuration of parallel resources (e.g., number of processes and number of cores) and also helps system designers to investigate the impact of changes in application parameters on system performance.",https://ieeexplore.ieee.org/document/8030609,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 161}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 58}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 419}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 167}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 119}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 405}]",4.0,"['linear-regression', 'markov-model']",[],,['workload-prediction'],,2017 IEEE 10th International Conference on Cloud Computing (CLOUD),True,['resource-provisioning'],,,,,,,,,
23,Predicting Cloud Resource Utilization,"['M. Borkowski', ' S. Schulte', ' C. Hochreiner']",2016,"A major challenge in Cloud computing is resource provisioning for computational tasks. Not surprisingly, previous work has established a number of solutions to provide Cloud resources in an efficient manner. However, in order to realize a holistic resource provisioning model, a prediction of the future resource consumption of upcoming computational tasks is necessary. Nevertheless, the topic of prediction of Cloud resource utilization is still in its infancy stage. In this paper, we present an approach for predicting Cloud resource utilization on a per-task and per-resource level. For this, we apply machine learning-based prediction models. Based on extensive evaluation, we show that we can reduce the prediction error by 20% in a typical case, and improvements above 89% are among the best cases.",https://ieeexplore.ieee.org/document/7881613,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 194}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 421}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 440}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 178}, {'database': 'ACM', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 60}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 45}, {'database': 'ACM', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 43}, {'database': 'ACM', 'search_string': ""'classification' AND ('cloud')"", 'index': 23}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 48}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 23}, {'database': 'ACM', 'search_string': ""'regression' AND ('cloud')"", 'index': 7}, {'database': 'ACM', 'search_string': ""'clustering' AND ('cloud')"", 'index': 21}]",9.0,,,,,,2016 IEEE/ACM 9th International Conference on Utility and Cloud Computing (UCC),True,['resource-provisioning'],,,,,,10.1145/2996890.2996907,Conference Paper,"['usage prediction', 'cloud computing', 'resource usage', 'machine learning']",
24,Dependency analysis of cloud applications for performance monitoring using recurrent neural networks,"['S. Y. Shah', ' Z. Yuan', ' S. Lu', ' P. Zerfos']",2017,"Performance monitoring of cloud-native applications that consist of several micro-services involves the analysis of time series data collected from the infrastructure, platform, and application layers of the cloud software stack. The analysis of the runtime dependencies amongst the component microservices is an essential step towards performing cloud resource management, detecting anomalous behavior of cloud applications, and meeting customer Service Level Agreements (SLAs). Finding such dependencies is challenging due to the non-linear nature of interactions, aberrant data measurements and lack of domain knowledge. In this paper, we propose a novel use of the modeling capability of Long-Short Term Memory (LSTM) recurrent neural networks, which excel in capturing temporal relationships in multi-variate time series data and being resilient to noisy pattern representations. Our proposed technique looks into the LSTM model structure, to uncover dependencies amongst performance metrics, which were learned during training. We further apply this technique in three monitoring use cases, namely finding the strongest performance predictors, discovering lagged/temporal dependencies, and improving the accuracy of forecasting for a given metric. We demonstrate the viability of our approach, by comparing the results of our proposed method in the three use cases with those obtained from previously proposed methods, such as Granger causality and the classical statistical time series analysis models, such as ARIMA and Holt-Winters. For our experiments and analysis, we use performance monitoring data collected from two sources: a controlled experiment involving a sample cloud application that we deployed in a public cloud infrastructure and cloud monitoring data collected from the monitoring service of an operational, public cloud service provider.",https://ieeexplore.ieee.org/document/8258087,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 206}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 415}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 515}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 168}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1373}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1099}]",4.0,['rnn'],"['kpis', 'host-metrics']",['novel-use'],"['failure-prediction', 'failure-detection', 'root-cause-analysis']",,2017 IEEE International Conference on Big Data (Big Data),True,['failure-management'],,,"['system-failure-prediction', 'anomaly-detection', 'rca-others']",,,,,,
25,A Learning-Based Service for Cost and Performance Management of Cloud Databases,"['R. Marcus', ' S. Semenova', ' O. Papaemmanouil']",2017,"Data management applications deployed on IaaS cloud environments must simultaneously strive to minimize cost and provide good performance. Balancing these two goals requires complex decision-making across a number of axes: resource provisioning, query placement, and query scheduling. While previous works have addressed each axis in isolation for specific types of performance goals, this demonstration showcases WiSeDB, a cloud workload management advisor service that uses machine learning techniques to address all dimensions of the problem for customizable performance goals. In our demonstration, attendees will see WiSeDB in action for a variety of workloads and performance goals.",https://ieeexplore.ieee.org/document/7930073,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 228}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 584}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1073}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 343}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 311}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 725}]",1.0,,,,,,2017 IEEE 33rd International Conference on Data Engineering (ICDE),True,['resource-provisioning'],,,,,,,,,
26,Hi-Clust: Unsupervised Analysis of Cloud Latency Measurements Through Hierarchical Clustering,"['P. Mulinka', ' P. Casas', ' L. Kencl']",2018,"Latency is nowadays one of the most relevant network and service performance metrics reflecting end-user experience. With the wide adoption and deployment of delay-sensitive applications in the Cloud (e.g., gaming, interactive video conferencing, corporate services, etc.), monitoring and analysis of Cloud service latency is becoming increasingly relevant for Cloud service providers, tenants and even users. Traditional network monitoring approaches based on time-series analysis and thresholding are capable of raising alarms when anomalous events arise, but are not applicable to detect correlations among multiple monitored dimensions, necessary to provide an adequate interpretation of an anomaly. In this paper we present Hi-Clust, an unsupervised-based approach for analyzing and interpreting anomalies in multi-dimensional network data, through the application of hierarchical clustering techniques. While Hi-Clust is applicable to the analysis of different types of nested or hierarchically structured data, we particularly focus on the analysis of Cloud service latency, using active measurements collected from geographically distributed vantage points. We implement and benchmark multiple density-based clustering approaches for Hi-Clust over four weeks of real multidimensional Cloud service latency measurements. Using the most robust underlying clustering algorithm from the benchmark, we show how to automatically extract and interpret anomalous Cloud service behavior with Hi-Clust. In addition, we show the advantages of Hi-Clust over traditional threshold-based approaches for detecting and interpreting anomalous behavior, through practical examples over the collected measurements.",https://ieeexplore.ieee.org/document/8549558,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 240}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 210}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 316}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 237}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 304}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 87}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1277}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 781}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 804}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 44}]",1.0,['clustering'],['kpis'],['novel-use'],['failure-detection'],[],2018 IEEE 7th International Conference on Cloud Networking (CloudNet),True,['failure-management'],,True,,,,,,,
27,Study on a migration scheme by fuzzy-logic-based learning and decision approach for QoS in cloud computing,"['A. Son', ' E. Huh']",2017,"Migration contributes to efficient resource management in cloud computing environment. Therefore, migration is used in many areas such as load balancing in Cloud Data Centers (CDCs). However, most of the previous research has concentrated on minimization of migration time. Also, previous works generally did not consider balance between QoS metrics. Due to importance of balance of metrics, we have to consider combining multiple metrics. In this paper, we present QoS metrics of migration and proposed migration scheme based on fuzzy logic. The main goal of this paper is to apply fuzzy logic and machine learning technique for advanced migration.",https://ieeexplore.ieee.org/document/7993836,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 242}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 654}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 98}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 55}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 11}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 211}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1028}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 3}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 39}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 34}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 7}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 72}]",1.0,['fuzzy-logic'],,,,['vm'],2017 Ninth International Conference on Ubiquitous and Future Networks (ICUFN),True,['resource-provisioning'],,,,,,,,,
28,HYPER-VINES: A HYbrid Learning Fault and Performance Issues ERadicator for Virtual NEtwork Services over Multi-Cloud Systems,"['L. Gupta', ' T. Salman', ' R. Das', ' A. Erbad', ' R. Jain', ' M. Samaka']",2019,"Fault and performance management systems, in the traditional carrier networks, are based on rule-based diagnostics that correlate alarms and other markers to detect and localize faults and performance issues. As carriers move to Virtual Network Services, based on Network Function Virtualization and multi-cloud deployments, the traditional methods fail to deliver because of the intangibility of the constituent Virtual Network Functions and increased complexity of the resulting architecture. In this paper, we propose a framework, called HYPER-VINES, that interfaces with various management platforms involved to process markers through a system of shallow and deep machine learning models. It then detects and localizes manifested and impending fault and performance issues. Our experiments validate the functionality and feasibility of the framework in terms of accurate detection and localization of such issues and unambiguous prediction of impending issues. Simulations with real network fault datasets show the effectiveness of its architecture in large networks.",https://ieeexplore.ieee.org/document/8685496,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 247}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 135}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 232}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 379}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1865}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 248}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 969}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 523}]",3.0,"['autoencoder', 'support-vector-machine']",,,['root-cause-analysis'],['network'],"2019 International Conference on Computing, Networking and Communications (ICNC)",True,['failure-management'],,True,['fault-localization'],,,,,,
29,Resource Allocation of Cloud Application Through Machine Learning: A Case Study,"['J. Lin', ' Y. Dai', ' X. Chen', ' Y. Wu']",2017,"With the rapid development of WEB applications, the demand for dynamically adjusting computing resources based on the load variation is increasing. However, most of the traditional WEB systems have limited ability to respond to load changes. In order to solve the problem, software self-adaptation technology has been applied to the resource management of WEB systems. Many researchers have tried to propose various software self-adaptation models, all of which contain the control loop ""Monitor-Analyze-Plan-Execute"". However, the ""Knowledge-base"" of these models is based on predefined strategy or configuration, which increases the complexity of system development and maintenance. In this paper, machine learning is applied to software self-adaptation through a case study, and the ""Knowledge-base"" is given by machine learning, which greatly reduces the workload of system maintenance and rule configuration. The results show the case based on machine learning can construct the corresponding management strategy according to WEB system runtime status, and conduct software self-adaptive management.",https://ieeexplore.ieee.org/document/8117114,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 248}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1371}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 348}]",1.0,,,,['resource-consolidation'],,2017 International Conference on Green Informatics (ICGI),True,['resource-provisioning'],,,,,,,,,
30,Predicting the End-to-End Tail Latency of Containerized Microservices in the Cloud,"['J. Rahman', ' P. Lama']",2019,"Large-scale web services are increasingly adopting cloud-native principles of application design to better utilize the advantages of cloud computing. This involves building an application using many loosely coupled service-specific components (microservices) that communicate via lightweight APIs, and utilizing containerization technologies to deploy, update, and scale these microservices quickly and independently. However, managing the end-to-end tail latency of requests flowing through the microservices is challenging in the absence of accurate performance models that can capture the complex interplay of microservice workflows with cloudinduced performance variability and inter-service performance dependencies. In this paper, we present performance characterization and modeling of containerized microservices in the cloud. Our modeling approach aims at enabling cloud platforms to combine resource usage metrics collected from multiple layers of the cloud environment, and apply machine learning techniques to predict the end-to-end tail latency of microservice workflows. We implemented and evaluated our modeling approach on NSF Cloud's Chameleon testbed using KVM for virtualization, Docker Engine for containerization and Kubernetes for container orchestration. Experimental results with an open-source microservices benchmark, Sock Shop, show that our modeling approach achieves high prediction accuracy even in the presence of multi-tenant performance interference.",https://ieeexplore.ieee.org/document/8790059,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 250}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 846}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 262}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1982}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 755}]",1.0,,,,['workload-prediction'],,2019 IEEE International Conference on Cloud Engineering (IC2E),True,['resource-provisioning'],,,,,,,,,
31,Evaluating machine learning algorithms for anomaly detection in clouds,"['A. Gulenko', ' M. Wallschläger', ' F. Schmidt', ' O. Kao', ' F. Liu']",2016,"Critical services in the field of Network Function Virtualization require elaborate reliability and high availability mechanisms to meet the high service quality requirements. Traditional monitoring systems detect overload situations and outages in order to automatically scale out services or mask faults. However, faults are often preceded by anomalies and subtle misbehaviors of the services, which are overlooked when detecting only outages. We propose to exploit machine learning techniques to detect abnormal behavior of services and hosts by analysing metrics collected from all layers and components of the cloud infrastructure. Various algorithms are able to compute models of a hosts normal behavior that can be used for anomaly detection at runtime. An offline evaluation of data collected from anomaly injection experiments shows that the models are able to achieve very high precision and recall values.",https://ieeexplore.ieee.org/document/7840917,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 268}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1811}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 905}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1698}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 287}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 200}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 1052}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 97}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 568}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 272}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1361}]",21.0,"['random-forest', 'naive-bayes', 'rule-mining', 'support-vector-machine', 'decision-tree']","['host-metrics', 'network-metrics']",['comparison'],['failure-detection'],,2016 IEEE International Conference on Big Data (Big Data),True,['failure-management'],,,['anomaly-detection'],,,,,,
32,Machine Learning Based Prediction and Classification of Computational Jobs in Cloud Computing Centers,"['Z. Zhu', ' P. Fan']",2019,"With the rapid growth of the data volume and the fast increasing of the computational model complexity in the scenario of cloud computing, it becomes an important topic that how to handle users' requests by scheduling computational jobs and assigning the resources in data center.In order to have a better perception of the computing jobs and their requests of resources, we analyze its characteristics and focus on the prediction and classification of the computing jobs with some machine learning approaches. Specifically, we apply LSTM neural network to predict the arrival of the jobs and the aggregated requests for computing resources. Then we evaluate it on Google Cluster dataset and it shows that the accuracy has been improved compared to the current existing methods. Additionally, to have a better understanding of the computing jobs, we use an unsupervised hierarchical clustering algorithm, BIRCH, to make classification and get some interpretability of our results in the computing centers.",https://ieeexplore.ieee.org/document/8766558,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 269}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 139}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 442}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 154}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 354}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 170}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 274}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 197}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1274}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 699}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 57}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 451}, {'database': 'arxiv', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 7}, {'database': 'arxiv', 'search_string': ""'classification' AND ('cloud')"", 'index': 124}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 172}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 13}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 111}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 2}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 87}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 4}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 304}]",0.0,"['rnn', 'clustering']",,['novel-use'],"['scheduling', 'workload-prediction']",,2019 15th International Wireless Communications & Mobile Computing Conference (IWCMC),True,['resource-provisioning'],,True,,,,10.1109/IWCMC.2019.8766558,,,
33,Using Quantile Regression for Reclaiming Unused Cloud Resources While Achieving SLA,"['J. Dartois', ' A. Knefati', ' J. Boukhobza', ' O. Barais']",2018,"Although Cloud computing techniques have reduced the total cost of ownership thanks to virtualization, the average usage of resources (e.g., CPU, RAM, Network, I/O) remains low. To address such issue, one may sell unused resources. Such a solution requires the Cloud provider to determine the resources available and estimate their future use to provide availability guarantees. This paper proposes a technique that uses machine learning algorithms (Random Forest, Gradient Boosting Decision Tree, and Long Short Term Memory) to forecast 24-hour of available resources at the host level. Our technique relies on the use of quantile regression to provide a flexible trade-off between the potential amount of resources to reclaim and the risk of SLA violations. In addition, several metrics (e.g., CPU, RAM, disk, network) were predicted to provide exhaustive availability guarantees. Our methodology was evaluated by relying on four in production data center traces and our results show that quantile regression is relevant to reclaim unused resources. Our approach may increase the amount of savings up to 20% compared to traditional approaches.",https://ieeexplore.ieee.org/document/8590999,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 281}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 49}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1507}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 161}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 491}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 662}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 659}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 49}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1206}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 491}, {'database': 'IEEE', 'search_string': ""'regression' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 97}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 182}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 559}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 311}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 805}]",2.0,"['random-forest', 'rnn', 'linear-regression', 'decision-tree']",['host-metrics'],['novel-use'],"['resource-consolidation', 'workload-prediction']",,2018 IEEE International Conference on Cloud Computing Technology and Science (CloudCom),True,['resource-provisioning'],,True,,,,,,,
34,Towards Self-Managing Cloud Storage with Reinforcement Learning,"['R. R. Noel', ' R. Mehra', ' P. Lama']",2019,"Cloud storage services are often associated with various performance issues due to load imbalance, interference from background tasks such as data scrubbing, backfilling, recovery, and the difference in processing capabilities of heterogeneous servers in a datacenter. This has a significant impact on a broad range of applications that are characterized by massive working sets and real-time constraints. However, it is challenging and burdensome for human operators to hand-tune various control-knobs in a cloud-scale storage cluster for maintaining optimal performance under diverse workload conditions. Our study on an open-source object-based storage system, Ceph, shows that common load balancing strategies are ineffective unless they are adapted according to workload characteristics. Furthermore, positive effects of an applied strategy may not be immediately visible. To address these challenges, we developed a machine learning based system adaptation technique that enables a cloud storage system to manage itself through load balancing and data migration with the aim of delivering optimal performance in the face of diverse workload patterns and resource bottlenecks. In particular, we applied a stochastic policy gradient based reinforcement learning technique to track performance hotspots in the storage cluster, and take appropriate corrective actions to maximize future performance under a variety of complex scenarios. For this purpose, we leveraged system-level performance monitoring and commonly available control-knobs in object-based cloud storage systems. We implemented the developed techniques to build an Adaptive Resource Management (ARM) system for object based storage cluster, and evaluated its performance on NSF Cloud's Chameleon testbed. Experiments using Cloud Object Storage Benchmark (COSBench) show that, ARM improves the average read and write response time of Ceph storage cluster by upto 50% and 33% respectively, compared to the default case. It also outperforms a state-of-the-art dynamic load rebalancing technique in terms of read and write performance of Ceph storage by 43% and 36% respectively.",https://ieeexplore.ieee.org/document/8789906,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 286}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 67}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 1286}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 863}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 9}, {'database': 'IEEE', 'search_string': ""'clustering' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 55}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('remediation' OR 'recovery')"", 'index': 773}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 28}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 416}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 53}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 88}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('remediation' OR 'recovery')"", 'index': 43}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1281}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 468}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 1771}]",1.0,['reinforcement-learning'],,,['resource-consolidation'],['storage'],2019 IEEE International Conference on Cloud Engineering (IC2E),True,['resource-provisioning'],,,,,,,,,
35,A Combined Analytical Modeling Machine Learning Approach for Performance Prediction of MapReduce Jobs in Cloud Environment,"['E. Ataie', ' E. Gianniti', ' D. Ardagna', ' A. Movaghar']",2016,"Nowadays MapReduce and its open source implementation, Apache Hadoop, are the most widespread solutions for handling massive dataset on clusters of commodity hardware. At the expense of a somewhat reduced performance in comparison to HPC technologies, the MapReduce framework provides fault tolerance and automatic parallelization without any efforts by developers. Since in many cases Hadoop is adopted to support business critical activities, it is often important to predict with fair confidence the execution time of submitted jobs, for instance when SLAs are established with end-users. In this work, we propose and validate a hybrid approach exploiting both queuing networks and support vector regression, in order to achieve a good accuracy without too many costly experiments on a real setup. The experimental results show how the proposed approach attains a 21% improvement in accuracy over applying machine learning techniques without any support from analytical models.",https://ieeexplore.ieee.org/document/7829644,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 290}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 212}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1346}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1373}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 215}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 264}]",7.0,,,,,,2016 18th International Symposium on Symbolic and Numeric Algorithms for Scientific Computing (SYNASC),True,['resource-provisioning'],,True,,,,,,,
36,Predicting cloud resource provisioning using machine learning techniques,"['A. A. Bankole', ' S. A. Ajila']",2013,"In order to meet Service Level Agreement (SLA) requirements, Virtual Machine (VM) resources must be provisioned few minutes ahead due to the VM boot-up time. One way to do this is by predicting future resource demands. In this research, we have developed and evaluated cloud client prediction models for TPCW benchmark web application using three machine learning techniques: Support Vector Machine (SVM), Neural Networks (NN) and Linear Regression (LR). We included the SLA metrics for Response Time and Throughput to the prediction model with the aim of providing the client with a more robust scaling decision choice. Our results show that Support Vector Machine provides the best prediction model.",https://ieeexplore.ieee.org/document/6567848,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 296}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 139}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1927}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 527}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1991}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 153}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1390}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 192}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 440}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 164}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 226}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 403}]",20.0,,,,,,2013 26th IEEE Canadian Conference on Electrical and Computer Engineering (CCECE),True,['resource-provisioning'],,True,,,,,,,
37,Predicting Application Failure in Cloud: A Machine Learning Approach,"['T. Islam', ' D. Manivannan']",2017,"Despite employing the architectures designed for high service reliability and availability, cloud computing systems do experience service outages and performance slowdown. In addition to these, large-scale cloud systems experience failures in their hardware and software components which often result in node and application (e.g., jobs and tasks) failures. Therefore, to build a reliable cloud system, it is important to understand and characterize the observed failures. The goal of this work is to identify the key features that correlate to application failures in cloud and present a failure prediction model that can correctly predict the outcome of a task or job before it actually finishes, fails or gets killed. To accomplish this, we perform a failure characterization study of the Google cluster workload trace. Our analysis reveals that, there is a significant consumption of resources due to failed and killed jobs. We further explore the potential for failure prediction in cloud applications so that we can reduce the wastage of resources by better managing the jobs and tasks that ultimately fail or get killed. For this, we propose a prediction method based on a special type of Recurrent NeuralNetwork (RNN) named Long Short-Term Memory Network(LSTM) to identify application failures in cloud. It takes resource usage measurements or performance data for each job and task, and the goal is to predict the termination status (e.g., failed and finished etc.) of them. Our algorithm can predict task failures with 87%accuracy and achieves a true positive rate of 85% and false positive rate of 11%.",https://ieeexplore.ieee.org/document/8029219,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 301}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 496}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 988}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1393}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 540}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 187}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 935}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 122}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 325}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 457}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 580}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 174}]",13.0,['rnn'],"['kpis', 'host-metrics']",['novel-use'],['failure-prediction'],['task'],2017 IEEE International Conference on Cognitive Computing (ICCC),True,['failure-management'],True,,['system-failure-prediction'],True,49.0,,,,
38,Self-Adaptive and Online QoS Modeling for Cloud-Based Software Services,"['T. Chen', ' R. Bahsoon']",2017,"In the presence of scale, dynamism, uncertainty and elasticity, cloud software engineers faces several challenges when modeling Quality of Service (QoS) for cloud-based software services. These challenges can be best managed through self-adaptivity because engineers' intervention is difficult, if not impossible, given the dynamic and uncertain QoS sensitivity to the environment and control knobs in the cloud. This is especially true for the shared infrastructure of cloud, where unexpected interference can be caused by co-located software services running on the same virtual machine; and co-hosted virtual machines within the same physical machine. In this paper, we describe the related challenges and present a fully dynamic, self-adaptive and online QoS modeling approach, which grounds on sound information theory and machine learning algorithms, to create QoS model that is capable to predict the QoS value as output over time by using the information on environmental conditions, control knobs and interference as inputs. In particular, we report on in-depth analysis on the correlations of selected inputs to the accuracy of QoS model in cloud. To dynamically selects inputs to the models at runtime and tune accuracy, we design self-adaptive hybrid dual-learners that partition the possible inputs space into two sub-spaces, each of which applies different symmetric uncertainty based selection techniques; the results of sub-spaces are then combined. Subsequently, we propose the use of adaptive multi-learners for building the model. These learners simultaneously allow several learning algorithms to model the QoS function, permitting the capability for dynamically selecting the best model for prediction on the fly. We experimentally evaluate our models in the cloud environment using RUBiS benchmark and realistic FIFA 98 workload. The results show that our approach is more accurate and effective than state-of-the-art modelings.",https://ieeexplore.ieee.org/document/7572219,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 306}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1307}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 723}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 590}]",13.0,"['multilayer-perceptron', 'autoregression', 'regression-tree']",,['novel-use'],['workload-prediction'],"['cloud', 'vm', 'rubis']",IEEE Transactions on Software Engineering,True,['resource-provisioning'],,,,,,,,,
39,An RBM Anomaly Detector for the Cloud,"['C. Monni', ' M. Pezzè', ' G. Prisco']",2019,"Failures are unavoidable in complex software systems, and the intrinsic characteristics of cloud systems amplify the problem. Predicting failures before their occurrence by detecting anomalies in system metrics is a viable solution to enable failure preventing or mitigating actions. The most promising approaches for predicting failures exploit statistical analysis or machine learning to reveal anomalies and their correlation with possible failures. Statistical analysis approaches result in far too many false positives, which severely hinder their practical applicability, while accurate machine learning approaches need extensive training with seeded faults, which is often impossible in operative cloud systems. In this paper, we propose EmBeD, Energy-Based anomaly Detection in the cloud, an approach to detect anomalies at runtime based on the free energy of a Restricted Boltzmann Machine (RBM) model. The free energy is a stochastic function that can be used to efficiently score anomalies for detecting outliers. EmBeD analyzes the system behavior from raw metric data, does not require extensive training with seeded faults, and classifies the relation of anomalous behaviors with future failures with very few false positives. The experimental results presented in this paper confirm that EmBeD can precisely predict failure-prone behavior without training with seeded faults, thus overcoming the main limitations of current approaches.",https://ieeexplore.ieee.org/document/8730157,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 311}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1860}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 198}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1129}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 570}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 119}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 530}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1045}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 545}]",2.0,['boltzmann-machine'],['host-metrics'],['novel-use'],['failure-detection'],,"2019 12th IEEE Conference on Software Testing, Validation and Verification (ICST)",True,['failure-management'],,True,,,,,,,
40,Identifying Changed or Sick Resources from Logs,"['A. Poghosyan', ' N. Grigoryan', ' N. Kushmerick', ' H. Beybutyan', ' A. Harutyunyan']",2018,"The identification of important changes in a complex distributed system is a challenging data science problem. Solving this problem is critical for tools for managing modern cloud infrastructure stacks and other large complex distributed systems. In this paper, we investigate two specific approaches to using log data to solve this problem. The first approach is comparing a source's current and past behavior. Some solutions that perform anomaly detection on numeric data from the data center are inevitably relying on global change point detection concepts. On the other hand, while log data promises a significantly different perspectives and dimensions to accomplish a similar task, state-of-the-art of solutions lack a capability to automatically detect significant change points in the log stream of an event source through learning its behavioral patterns. Such change points indicate the most important times when the source's behavior significantly differs from the past. A second complementary approach to real-time change detection involves comparing a source's current behavior with the current behavior of its peers in a population of sources serving a common role in the data center. Employing the concept of event types of log messages introduced earlier, we propose algorithms for each of these approaches that apply classical statistical and machine learning techniques to data capturing the distribution of those constructs. We demonst.rate experimental results from our prototype algorithms.",https://ieeexplore.ieee.org/document/8599537,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 319}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 36}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 845}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 60}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 57}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 127}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 470}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 742}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1059}]",2.0,"['similarity-matching', 'clustering']",['logs'],['novel-use'],['failure-detection'],,2018 IEEE 3rd International Workshops on Foundations and Applications of Self* Systems (FAS*W),True,['failure-management'],,True,['anomaly-detection'],,,,,,
41,Cloud Client Prediction Models for Cloud Resource Provisioning in a Multitier Web Application Environment,"['A. A. Bankole', ' S. A. Ajila']",2013,"In order to meet Service Level Agreement (SLA) requirements, efficient scaling of Virtual Machine (VM) resources must be provisioned few minutes ahead due to the VM boot-up time. One way to proactively provision resources is by predicting future resource demands. In this research, we have developed and evaluated cloud client prediction models for TPC-W benchmark web application using three machine learning techniques: Support Vector Machine (SVM), Neural Networks (NN) and Linear Regression (LR). We included the SLA metrics for Response Time and Throughput to the prediction model with the aim of providing the client with a more robust scaling decision choice. Our results show that Support Vector Machine provides the best prediction model.",https://ieeexplore.ieee.org/document/6525518,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 326}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 122}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 492}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1810}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1030}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 760}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1980}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 714}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 98}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 771}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 282}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 913}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 163}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 225}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 813}]",27.0,,,,['workload-prediction'],,2013 IEEE Seventh International Symposium on Service-Oriented System Engineering,True,['resource-provisioning'],,,,,,,,,
42,Anomaly Detection from System Tracing Data Using Multimodal Deep Learning,"['S. Nedelkoski', ' J. Cardoso', ' O. Kao']",2019,"The concept of Artificial Intelligence for IT Operations (AIOps) combines big data and machine learning methods to replace a broad range of IT operations including availability and performance monitoring of services. Such platforms typically use separate models for each modality of monitoring data (e.g., textual properties and real-valued response time in logs and traces) to detect faults and upcoming anomalies in cloud services, which do not capture the existing correlation between the modalities. This paper extends the range of utilized data types for creation of a single model to improve the anomaly detection. We use a bimodal distributed tracing data from large cloud infrastructures in order to detect an anomaly in the execution of system components. We propose an anomaly detection method, which utilizes a single modality of the data with information about the trace structure. In the next step, we extend the single-modality neural architecture to a multimodal neural network with long short-term memory (LSTM) to enable the learning from the sequential nature of both modalities in the tracing data. Furthermore, we demonstrate an approach to detect dependent and concurrent events using the ability of the model to reconstruct the execution path. The implemented prototype is experimentally evaluated with data from a large-scale production cloud. The results demonstrate that the novel approaches outperform other deep-learning methods based on traditional architectures.",https://ieeexplore.ieee.org/document/8814585,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 341}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 692}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 62}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 278}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 976}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 61}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('IT operations')"", 'index': 1}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 33}, {'database': 'IEEE', 'search_string': ""'AIOps'"", 'index': 6}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('IT operations')"", 'index': 2}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 5}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 118}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 934}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 495}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1119}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 339}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 70}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 832}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('IT operations')"", 'index': 4}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 1582}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('IT operations')"", 'index': 2}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 537}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 624}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 55}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 464}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 972}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('fault detection' OR 'failure detection')"", 'index': 89}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 115}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 316}]",0.0,['rnn'],"['events', 'traces']",['novel-use'],['failure-detection'],,2019 IEEE 12th International Conference on Cloud Computing (CLOUD),True,['failure-management'],,,,,,,,,
43,Cloud Client Prediction Models Using Machine Learning Techniques,"['S. A. Ajila', ' A. A. Bankole']",2013,"One way to proactively provision resources and meet Service Level Agreements (SLA) is by predicting future resource demands a few minutes ahead because of Virtual Machine (VM) boot time. In this research, we have developed and evaluated cloud client prediction models for TPC-W benchmark web application using three machine learning techniques: Support Vector Machine (SVM), Neural Networks (NN) and Linear Regression (LR). We have included two SLA metrics -- Response Time and Throughput with the aim of providing the client with a more robust scaling decision choice. As an improvement to our previous work, we implemented our model on a public cloud infrastructure: Amazon EC2. Furthermore, we extended the experimentation time by over 200%. Finally, we have employed random workload pattern to reflect a more realistic simulation. Our results show that Support Vector Machine provides the best prediction model.",https://ieeexplore.ieee.org/document/6649808,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 350}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 160}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1630}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 525}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1957}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 182}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1573}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 191}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 432}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 162}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 224}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 399}]",16.0,,,,,,2013 IEEE 37th Annual Computer Software and Applications Conference,True,['resource-provisioning'],,True,,,,,,,
44,Improving the Smartness of Cloud Management via Machine Learning Based Workload Prediction,"['Y. Yu', ' V. Jindal', ' F. Bastani', ' F. Li', ' I. Yen']",2018,"Cloud computing has been widely adopted by many companies and government entities. To ensure high quality computing resource provisioning, cloud platforms should offer smart resource management solutions. An important step toward better resource management is to accurately predict the workloads of the applications running on the cloud. Many existing workload prediction methods are regression based, which require the workloads of the applications show clear seasonality and trend. However, it is difficult to use these methods for tasks which may not have such recurring workload patterns. From careful analysis of the workloads in a real-world cloud, we found that many tasks have busty workloads that are very difficult to predict using regression-based prediction. Instead, we consider a job-pool based approach, where the knowledge about the workloads of a large pool of tasks is used to help predict the workloads of new tasks. In particular, we develop a clustering-based learning approach to realize the job-pool based concept. The pool of jobs are clustered based on their workloads, and a neuralnet is used to learn the characteristics of the workloads in each cluster. When a new job arrives, we use its initial workload pattern and submission parameters to find the cluster it belongs to. Then, the corresponding neuralnet is used to predict the workload of the new job far into the future. Based on this predicted long-term workload, smart resource management decisions can be made to reduce the potential overhead in scaling and migration. We also consider a non-clustering based learning solution and compare it with the clustering-based learning solution. Experimental results show that the clustering-based learning approach can predict the workload more accurately.",https://ieeexplore.ieee.org/document/8377827,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 359}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 150}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 620}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 480}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1171}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 528}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1080}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 174}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 410}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 460}]",1.0,,,,,,2018 IEEE 42nd Annual Computer Software and Applications Conference (COMPSAC),True,['resource-provisioning'],,,,,,,,,
45,Cloud Workload Prediction and Generation Models,"['G. M. Wamba', ' Y. Li', ' A. Orgerie', ' N. Beldiceanu', ' J. Menaud']",2017,"Cloud computing allows for elasticity as users can dynamically benefit from new virtual resources when their workload increases. Such a feature requires highly reactive resource provisioning mechanisms. In this paper, we propose two new workload prediction models, based on constraint programming and neural networks, that can be used for dynamic resource provisioning in Cloud environments. We also present two workload trace generators that can help to extend an experimental dataset in order to test more widely resource optimization heuristics. Our models are validated using real traces from a small Cloud provider. Both approaches are shown to be complimentary as neural networks give better prediction results, while constraint programming is more suitable for trace generation.",https://ieeexplore.ieee.org/document/8102182,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 364}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 369}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 545}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 343}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 819}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 439}]",5.0,,,,,,2017 29th International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD),True,['resource-provisioning'],,,,,,,,,
46,Learning-Based Anomaly Cause Tracing with Synthetic Analysis of Logs from Multiple Cloud Service Components,"['Y. Yuan', ' H. Anu', ' W. Shi', ' B. Liang', ' B. Qin']",2019,"It is critical for a reliable cloud to effectively find out root causes of the cloud service anomalies for efficacious treatment. System logs are widely used for anomaly detection and analysis. Many efforts have been made to handle massive cloud logs automatically. However, existing work can still not effectually make comprehensive use of logs from multiple cloud service components to locate the causes of cloud service anomalies automatically. In this paper, we propose a learning-based approach for fine-grained deep cloud service anomaly cause tracing by synthetically utilizing logs from multiple service components of a cloud. We focus on uncovering root causes of anomalies corresponding to system executions of each user operation rather than roughly taking various system tasks as a whole. Log patterns are learned from past experience of system runs with anomalies occurred before, where the mined log event sequences to represent system behaviors related to each user operation are treated as natural language sequences. When an anomaly is to be diagnosed, the corresponding log patterns can be recognized for root cause identification. We implemented and evaluated our approach in OpenStack. Experimental results show that our approach can effectively trace the root causes for anomalies in cloud environments.",https://ieeexplore.ieee.org/document/8754032,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 367}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 231}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 25}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 590}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 24}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1102}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 167}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1512}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1189}]",0.0,"['clustering', 'logistic-regression']",['logs'],,['root-cause-analysis'],['openstack'],2019 IEEE 43rd Annual Computer Software and Applications Conference (COMPSAC),True,['failure-management'],,True,,,,,,,
47,Modeling and Forecasting of Time-Aware Dynamic QoS Attributes for Cloud Services,"['Y. Syu', ' C. Wang', ' Y. Fanjiang']",2019,"Currently, statistical time-series methods have primarily been employed to predict time-aware dynamic quality of service (QoS) attributes for Web services. In this paper, we propose the application of genetic programming (GP) for such predictions. Our experimental results indicate that the GP-based approach is more accurate than the other approaches presented for comparison. However, for the efficient management of such attributes for cloud services, including their modeling and forecasting, the current research is insufficient because a set of research questions remains unanswered. In this paper, we first clearly define these research questions and then design and perform a set of empirical experiments to address the questions. Finally, the experimental results are exhaustively discussed to answer the studied research questions. The empirical study and analysis presented in this paper could be informative for the management (modeling and forecasting) of the time-aware dynamic QoS attributes of cloud services. For example, we verify that machine-learning approaches are generally superior to the widely used statistical time-series methods in terms of both modeling accuracy and forecasting accuracy. Furthermore, after considering a variety of situations and cases, the GP-based approach is still the best option for the studied problem. In addition, except for the technical approaches, this paper also exhaustively studies the influence of the properties of the cloud dynamic QoS attributes, including their size and time granularity.",https://ieeexplore.ieee.org/document/8558533,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 368}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1254}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 861}]",0.0,,,,['workload-prediction'],,IEEE Transactions on Network and Service Management,True,['resource-provisioning'],,,,,,,,,
48,Dynamic fault diagnosis framework for virtual machine rolling upgrade operation in google cloud platform,"['A. Cauveri', ' R. Kalpana']",2017,"Now a day's failure of system outages in cloud application is the main drawbacks of cloud environment. It reduces the economic losses for the business environment. Anomaly detection in the cloud system will reduce the loss. Anomaly detection at the user end is difficult, particularly Rolling Upgrade operation. Due to the enormous anomaly the operations are indistinguishable. It is very difficult to find the faults or anomalies. Current system fails to provide reliable assurance of successful execution to the system. Proposed anomaly detection using Decision tree J48 classifier giving the reliable assurance during running virtual machine upgrade operation. The proposed method was evaluated on the Google Cloud Platform. It will give the high precision, recall, accuracy to the cloud environment.",https://ieeexplore.ieee.org/document/8081093,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 389}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 675}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 903}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 332}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 707}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 224}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 350}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 784}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1235}]",1.0,,,,['root-cause-analysis'],"['vm', 'gcp']",2017 International Conference on Power and Embedded Drive Control (ICPEDC),True,['failure-management'],,,,,,,,,
49,Intelligent Resource Scheduling at Scale: A Machine Learning Perspective,"['R. Yang', ' X. Ouyang', ' Y. Chen', ' P. Townend', ' J. Xu']",2018,"Resource scheduling in a computing system addresses the problem of packing tasks with multi-dimensional resource requirements and non-functional constraints. The exhibited heterogeneity of workload and server characteristics in Cloud-scale or Internet-scale systems is adding further complexity and new challenges to the problem. Compared with,,,, existing solutions based on ad-hoc heuristics, Machine Learning (ML) has the potential to improve further the efficiency of resource management in large-scale systems. In this paper we,,,, will describe and discuss how ML could be used to understand automatically both workloads and environments, and to help to cope with scheduling-related challenges such as consolidating co-located workloads, handling resource requests, guaranteeing application's QoSs, and mitigating tailed stragglers. We will introduce a generalized ML-based solution to large-scale resource scheduling and demonstrate its effectiveness through a case study that deals with performance-centric node classification and straggler mitigation. We believe that an MLbased method will help to achieve architectural optimization and efficiency improvement.",https://ieeexplore.ieee.org/document/8359158,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 392}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 1580}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 199}]",1.0,,,,,,2018 IEEE Symposium on Service-Oriented System Engineering (SOSE),True,['resource-provisioning'],,,,,,,,,
50,TE-Based Machine Learning Techniques for Link Fault Localization in Complex Networks,"['S. Madapuzi Srinivasan', ' T. Truong-Huu', ' M. Gurusamy']",2018,"Communication networks such as wireless sensor networks, Internet of Things and vehicular ad-hoc networks are becoming more complex and increasing in size. This leads to high overhead (network and computation) and difficulty in determining the accurate network topology, which is an important information for traffic engineering and network management. Localization of link failures in such networks is a challenging problem and requires a novel approach to achieve the goal without any prior information about the network topology. In this paper, we present a traffic engineering (TE)-based machine learning approach to detect and localize link failures. Instead of using topology information and actively injecting additional packets to localize a failed link, the proposed machine learning model adopts a passive mechanism to learn the network traffic behavior from propagation delay, number of flows and average packet loss at every node in the network under normal working conditions and failure scenarios. We train the learning model with machine learning algorithms such as naive Bayes, logistic regression, support vector machine, multi-layer perceptron, decision tree and random forest. We implement the proposed approach and carry out extensive experiments using the Mininet platform. The performance study shows that our proposed approach localizes link failures with at least 90% accuracy using random forest algorithm while requiring less time-to-localization of a link failure compared to other existing works.",https://ieeexplore.ieee.org/document/8457989,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 402}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault localization' OR 'failure localization')"", 'index': 18}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 762}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault localization' OR 'failure localization')"", 'index': 14}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 493}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault localization' OR 'failure localization')"", 'index': 5}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 25}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('cloud')"", 'index': 91}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 2}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 1941}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('fault localization' OR 'failure localization')"", 'index': 2}]",3.0,"['random-forest', 'naive-bayes', 'logistic-regression', 'support-vector-machine', 'multilayer-perceptron', 'decision-tree']",,,['root-cause-analysis'],['network'],2018 IEEE 6th International Conference on Future Internet of Things and Cloud (FiCloud),True,['failure-management'],,True,,,,,,,
51,Semantic Aware Online Detection of Resource Anomalies on the Cloud,"['A. Bhattacharyya', ' S. A. J. Jandaghi', ' S. Sotiriadis', ' C. Amza']",2016,"As cloud based platforms become more popular, it becomes an essential task for the cloud administrator to efficiently manage the costly hardware resources in the cloud environment. Prompt action should be taken whenever hardware resources are faulty, or configured and utilized in a way that causes application performance degradation, hence poor quality of service. In this paper, we propose a semantic aware technique based on neural network learning and pattern recognition in order to provide automated, real-time support for resource anomaly detection. We incorporate application semantics to narrow down the scope of the learning and detection phase, thus enabling our machine learning technique to work at a very low overhead when executed online. As our method runs ""life-long"" on monitored resource usage on the cloud, in case of wrong prediction, we can leverage administrator feedback to improve prediction on future runs. This feedback directed scheme with the attached context helps us to achieve an anomaly detection accuracy of as high as 98.3% in our experimental evaluation, and can be easily used in conjunction with other anomaly detection techniques for the cloud.",https://ieeexplore.ieee.org/document/7830676,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 417}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 530}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 726}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 498}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 662}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 977}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1196}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 502}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 621}]",13.0,['rnn'],"['kpis', 'host-metrics']",['novel-use'],['failure-detection'],"['casssandra', 'mongodb', 'apache']",2016 IEEE International Conference on Cloud Computing Technology and Science (CloudCom),True,['failure-management'],,True,['anomaly-detection'],,,,,,
52,An Optimal Strategy for Resource Utilization in Cloud Data Centers,"['W. Ni', ' Y. Zhang', ' W. W. Li']",2019,"Cloud computing has emerged in recent years as one of the most interesting developments in technology. With the gaining popularity of cloud-based solutions, more and more applications are migrating into the Cloud and thus have highly demanding critical requirements for networking resources. Virtual technology associated with a Data Center consists of a set of servers, storage and network devices, power systems, cooling systems, etc., and makes it possible for the resource management of physical machines to be more finely tuned and thus support multiple virtual machines well. The growing challenge, however, is how to efficiently provision these resources to meet the requirements of the different qualities of service levels. This paper offers and investigates a general situation wherein a datacenter can determine the cost of using resources and a Cloud service user can decide whether it will pay the price for the resource or not for an incoming task. By establishing a Continuous-Time Markov Decision Process model for both an average reward model and a discounted expected reward model, the optimal policy of each model for admitting tasks can be verified to be a State-related control limit (threshold) policy, respectively. Further, a detailed statement and verification of the upper boundaries for such an optimal policy, and a comprehensive set of experiments on the various cases to validate this proposed solution are provided. Particularly, the machine learning method is implemented to obtain the optimal threshold values by using a feed-forward neural network model. Several numerical examples are also provided on how to derive optimal threshold values. The results offered in this paper can be easily utilized to help datacenter operate in an economically optimal way when providing different needed application services to Cloud service users.",https://ieeexplore.ieee.org/document/8887281,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 419}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 602}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 841}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 589}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 668}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 50}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 85}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 52}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 584}]",0.0,"['multilayer-perceptron', 'reinforcement-learning']",,,['resource-consolidation'],,,True,['resource-provisioning'],,True,,,,,,,
53,An Efficient Deep Learning Model to Predict Cloud Workload for Industry Informatics,"['Q. Zhang', ' L. T. Yang', ' Z. Yan', ' Z. Chen', ' P. Li']",2018,"Deep learning, as the most important architecture of current computational intelligence, achieves super performance to predict the cloud workload for industry informatics. However, it is a nontrivial task to train a deep learning model efficiently since the deep learning model often includes a great number of parameters. In this paper, an efficient deep learning model based on the canonical polyadic decomposition is proposed to predict the cloud workload for industry informatics. In the proposed model, the parameters are compressed significantly by converting the weight matrices to the canonical polyadic format. Furthermore, an efficient learning algorithm is designed to train the parameters. Finally, the proposed efficient deep learning model is applied to the workload prediction of virtual machines on cloud. Experiments are conducted on the datasets collected from PlanetLab to validate the performance of the proposed model by comparing with other machine-learning-based approaches for workload prediction of virtual machines. Results indicate that the proposed model achieves a higher training efficiency and workload prediction accuracy than state-of-the-art machine-learning-based approaches, proving the potential of the proposed model to provide predictive services for industry informatics.",https://ieeexplore.ieee.org/document/8301555,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 458}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 62}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 125}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1124}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 787}]",16.0,['autoencoder'],['host-metrics'],['novel-use'],['workload-prediction'],,IEEE Transactions on Industrial Informatics,True,['resource-provisioning'],,,,,,,,,
54,An adaptive machine learning on Map-Reduce framework for improving performance of large-scale data analysis on EC2,"['W. Romsaiyud', ' W. Premchaiswadi']",2013,"Map-Reduce is a programming for writing applications that rapidly process vast amounts of data in parallel on large cluster of computer nodes and can be deployed on cloud computing. However, to run a Map-Reduce job, many configuration parameters are required for tuning and improving the performance to set up such as number of running mappers and maximum number of reduce slots in the cluster in order to minimize the data transferred between map and reduce tasks. To say simple, the main emphasis is on reducing the job execution time as well as shuffling tweaks to tune parameters for memory management. In this paper, we introduce a machine learning model on top of Map-Reduce for automate setting of tuning parameters for Map-Reduce programs. Our model consists of three main steps; 1) describe the plan baseline marked for verification. 2) Propose a ML algorithm for learning and predicting the model, and 3) develop our automated method to run the program automatically at a specific time. In our experiments, we run Hadoop on 20-nodes cluster on EC2.",https://ieeexplore.ieee.org/document/6756290,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 461}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 153}]",5.0,,,,,,2013 Eleventh International Conference on ICT and Knowledge Engineering,True,['resource-provisioning'],,,,,,,,,
55,On the Use of Machine Learning to Predict the Time and Resources Consumed by Applications,"['A. Matsunaga', ' J. A. B. Fortes']",2010,"Most data centers, clouds and grids consist of multiple generations of computing systems, each with different performance profiles, posing a challenge to job schedulers in achieving the best usage of the infrastructure. A useful piece of information for scheduling jobs, typically not available, is the extent to which applications will use available resources once they are executed. This paper comparatively assesses the suitability of several machine learning techniques for predicting spatio temporal utilization of resources by applications. Modern machine learning techniques able to handle large number of attributes are used, taking into account application- and system-specific attributes (e.g., CPU micro architecture, size and speed of memory and storage, input data characteristics and input parameters). The work also extends an existing classification tree algorithm, called Predicting Query Runtime (PQR), to the regression problem by allowing the leaves of the tree to select the best regression method for each collection of data on leaves. The new method (PQR2) yields the best average percentage error, predicting execution time, memory and disk consumption for two bioinformatics applications, BLAST and RAxML, deployed on scenarios that differ in system and usage. In specific scenarios where usage is a non-linear function of system and application attributes, certain configurations of two other machine learning algorithms, Support Vector Machine and k-nearest neighbors, also yield competitive results. In addition, experiments show that the inclusion of system performance and application-specific attributes also improves the performance of machine learning algorithms investigated.",https://ieeexplore.ieee.org/document/5493447,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 462}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 400}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 590}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 521}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 1231}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 343}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 1744}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 341}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 949}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 327}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 93}]",73.0,,,,['workload-prediction'],,"2010 10th IEEE/ACM International Conference on Cluster, Cloud and Grid Computing",True,['resource-provisioning'],,,,,,10.1109/CCGRID.2010.98,Conference Paper,"['classifier tree', 'regression', 'application resource usage', 'machine learning']",
56,Localizing Faults in Cloud Systems,"['L. Mariani', ' C. Monni', ' M. Pezzé', ' O. Riganelli', ' R. Xin']",2018,"By leveraging large clusters of commodity hardware, the Cloud offers great opportunities to optimize the operative costs of software systems, but impacts significantly on the reliability of software applications. The lack of control of applications over Cloud execution environments largely limits the applicability of state-of-the-art approaches that address reliability issues by relying on heavyweight training with injected faults. In this paper, we propose LOUD, a lightweight fault localization approach that relies on positive training only, and can thus operate within the constraints of Cloud systems. LOUD relies on machine learning and graph theory. It trains machine learning models with correct executions only, and compensates the inaccuracy that derives from training with positive samples, by elaborating the outcome of machine learning techniques with graph theory algorithms. The experimental results reported in this paper confirm that LOUD can localize faults with high precision, by relying only on a lightweight positive training.",https://ieeexplore.ieee.org/document/8367054,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 488}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 901}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 42}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 30}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1043}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud')"", 'index': 825}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('fault localization' OR 'failure localization')"", 'index': 3}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 85}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 6}]",12.0,['graph-mining'],['kpis'],['new-method'],['root-cause-analysis'],,"2018 IEEE 11th International Conference on Software Testing, Verification and Validation (ICST)",True,['failure-management'],,True,['fault-localization'],,,,,,['online']
57,Categorizing hardware failure in large scale cloud computing environment,"['M. H. Khalil', ' W. M. Sheta', ' A. S. Elmaghraby']",2016,"Cloud computing environments are growing in complexity creating more challenges for improved resilience and availability. Cloud computing research can benefit from machine learning and data mining by using data from actual operational cloud systems. One aspect that needs in-depth analysis is the failure characteristics of cloud environments. Failure is the main contributor to reduced resiliency of applications and services in cloud computing. This work presents a categorizing method to identify machines removed from the system based on failure or due to maintenance. Our experiments are targeting large scale cloud computing environments and experimental data consists of 25 million submitted tasks on 12500 severs over a 29 day period. The parameters of categorizing are CPU and memory utilization. Also, this work developed a support vector machine (SVM) model for learning and prediction of machine failure. The devolved model achieved 99.04 % accuracy. Precision and Recall curves demonstrate that the model is consistent with varying data size. The model has very good consistency with max difference from theoretical data by only 0.008%.",https://ieeexplore.ieee.org/document/7886058,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 491}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 284}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 84}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 251}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 796}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 196}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 64}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 35}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 195}]",0.0,['support-vector-machine'],['host-metrics'],['novel-use'],['root-cause-analysis'],['hardware'],2016 IEEE International Symposium on Signal Processing and Information Technology (ISSPIT),True,['failure-management'],,,['root-cause-diagnosis'],,,,,,
58,Service Clustering for Autonomic Clouds Using Random Forest,"['R. B. Uriarte', ' S. Tsaftaris', ' F. Tiezzi']",2015,"Managing and optimising cloud services is one of the main challenges faced by industry and academia. A possible solution is resorting to self-management, as fostered by autonomic computing. However, the abstraction layer provided by cloud computing obfuscates several details of the provided services, which, in turn, hinders the effectiveness of autonomic managers. Data-driven approaches, particularly those relying on service clustering based on machine learning techniques, can assist the autonomic management and support decisions concerning, for example, the scheduling and deployment of services. One aspect that complicates this approach is that the information provided by the monitoring contains both continuous (e.g. CPU load) and categorical (e.g. VM instance type) data. Current approaches treat this problem in a heuristic fashion. This paper, instead, proposes an approach, which uses all kinds of data and learns in a data-driven fashion the similarities and resource usage patterns among the services. In particular, we use an unsupervised formulation of the Random Forest algorithm to calculate similarities and provide them as input to a clustering algorithm. For the sake of efficiency and meeting the dynamism requirement of autonomic clouds, our methodology consists of two steps: (i) off-line clustering and (ii) on-line prediction. Using datasets from real-world clouds, we demonstrate the superiority of our solution with respect to others and validate the accuracy of the on-line prediction. Moreover, to show the applicability of our approach, we devise a service scheduler that uses the notion of similarity among services and evaluate it in a cloud test-bed.",https://ieeexplore.ieee.org/document/7152517,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 512}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 265}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 308}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 310}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1738}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 210}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1168}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 281}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 753}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 181}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 38}, {'database': 'ACM', 'search_string': ""'clustering' AND ('cloud')"", 'index': 64}]",2.0,,,,,,"2015 15th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing",True,['resource-provisioning'],,,,,,10.1109/CCGrid.2015.41,Conference Paper,,
59,Prediction Of Cloud Computing Resource Utilization,"['T. Mehmood', ' S. Latif', ' S. Malik']",2018,"Efficient resource utilization leads cloud provider to low cost and high performance. Cloud Computing is a dynamic environment that provides on-demand services over the internet on pay as you go model. Cloud platform has a dynamic resource usage as it is shared among large number of users. Resource allocator provisions resources to dynamic demands of user from finite set of resources. There should be no over and under provisioning of resources. Underutilized resources causes resource wastage and more cost whereas over utilized resource can lead to service degradation. If Resource allocators can presume future resource usage they can take resource provisioning decision efficiently. A resource utilization prediction mechanism is required to assist resource allocator for optimum resource provisioning. Accurate prediction is a challenge in such a dynamic resource usage. Machine learning techniques can help in creating a model that yields accurate prediction results. In machine learning, Ensemble mechanisms are renowned for improving the prediction accuracy which uses a combination of learners rather than a single learner. In this study, an “Ensemble based workload prediction mechanism” is proposed that is based on stack generalization. Experiments are conducted in order to compare the proposed model with the individual and baseline prediction models. For comparison with baseline model, we have used Root Mean Square Error(RMSE) as results of baseline model were given in RMSE. Proposed mechanism has shown 6% and 17% reduction in RMSE in CPU usage and in Memory usage prediction respectively. For comparing our proposed ensemble with independent learner(K Nearest Neighbor, Neural Network, Decision Tree, Support Vector Machine and Naïve Bayes), we have used accuracy as evaluation parameter. The proposed ensemble has improved the prediction accuracy by ≈2%.",https://ieeexplore.ieee.org/document/8551339,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 518}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 1026}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 794}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 436}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 372}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 356}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 543}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 250}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 189}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 900}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 874}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 121}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 823}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 613}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 494}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 224}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 792}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 847}]",1.0,,,,,,2018 15th International Conference on Smart Cities: Improving Quality of Life Using ICT & IoT (HONET-ICT),True,['resource-provisioning'],,,,,,,,,
60,Self-adaptive and sensitivity-aware QoS modeling for the cloud,"['T. Chen', ' R. Bahsoon']",2013,"Given the elasticity, dynamicity and on-demand nature of the cloud, cloud-based applications require dynamic models for Quality of Service (QoS), especially when the sensitivity of QoS tends to fluctuate at runtime. These models can be autonomically used by the cloud-based application to correctly self-adapt its QoS provision. We present a novel dynamic and self-adaptive sensitivity-aware QoS modeling approach, which is fine-grained and grounded on sound machine learning techniques. In particular, we combine symmetric uncertainty with two training techniques: Auto-Regressive Moving Average with eXogenous inputs model (ARMAX) and Artificial Neural Network (ANN) to reach two formulations of the model. We describe a middleware for implementing the approach. We experimentally evaluate the effectiveness of our models using the RUBiS benchmark and the FIFA 1998 workload trends. The results show that our modeling approach is effective and the resulting models produce better accuracy when compared with conventional models.",https://ieeexplore.ieee.org/document/6595491,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 551}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 823}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1582}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 696}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1257}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 641}]",10.0,,,,,,2013 8th International Symposium on Software Engineering for Adaptive and Self-Managing Systems (SEAMS),True,['resource-provisioning'],,,,,,,,,
61,Workload management for cloud databases via machine learning,"['R. Marcus', ' O. Papaemmanouil']",2016,"As elastic IaaS clouds continue to become more cost efficient than on-site datacenters, a wide range of data management applications are migrating to pay-as-you-go cloud computing resources. These diverse applications come with an equally diverse set of performance goals, resource demands, and budget constraints. While existing research has tackled individual tasks such as query placement, scheduling, and resource provisioning to meet these goals and constraints, these techniques fail to provide end-to-end customizable workload management solutions, leading application developers to hand-craft custom heuristics that fit their workload specifications and performance goals. In this vision paper, we argue that workload management challenges can be addressed by leveraging machine learning algorithms. These algorithms can be trained on application-specific properties and performance metrics to automatically learn how to provision resources as well as distribute and schedule the execution of incoming query workloads. Towards this goal, we sketch our vision of WiSeDB, a learning-based service that relies on supervised and reinforcement learning to generate workload management strategies for both static and dynamic workloads.",https://ieeexplore.ieee.org/document/7495611,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 573}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 114}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1524}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 169}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 830}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1208}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 115}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1269}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 852}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 1048}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 550}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1592}]",4.0,,,,,,2016 IEEE 32nd International Conference on Data Engineering Workshops (ICDEW),True,['resource-provisioning'],,,,,,10.1109/ICDEW.2016.7495611,,,
62,Autonomic scaling of Cloud Computing resources using BN-based prediction models,['A. Bashar'],2013,"The recent surge in the popularity and usage of Cloud Computing services by both the enterprise and individual consumers has necessitated efficient and proactive management of data center resources which host services having varied characteristics. One of the major issues concerning both the cloud service providers and consumers is the automatic scalability of resources (i.e., compute, storage and bandwidth) in response to the highly unpredictable demands. To this end, an opportunity exists to harness the predictive and diagnostic capabilities of machine learning approaches to incorporate dynamic scaling up and scaling down of resources without violating the Service Level Agreements (SLA) and simultaneously ensuring adequate revenue to the providers. This paper proposes, implements and evaluates a Bayesian Networks based predictive modeling framework to provide for an autonomic scaling of utility computing resources in the Cloud Computing scenario. In essence, the BN-based model captures the historical behavior of the system involving various performance metrics (indicators) and predicts the desired unknown metric (e.g. SLA parameter). Initial simulated experiments involving random demand scenarios provide insights into the feasibility and applicability of the proposed approach for improving the management of present data center facilities.",https://ieeexplore.ieee.org/document/6710578,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 575}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1787}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 546}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 925}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1333}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 67}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 115}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 375}]",12.0,,,,,,2013 IEEE 2nd International Conference on Cloud Networking (CloudNet),True,['resource-provisioning'],,,,,,,,,
63,Intelligent management of virtualized resources for database systems in cloud environment,"['P. Xiong', ' Y. Chi', ' S. Zhu', ' H. J. Moon', ' C. Pu', ' H. Hacigümüş']",2011,"In a cloud computing environment, resources are shared among different clients. Intelligently managing and allocating resources among various clients is important for system providers, whose business model relies on managing the infrastructure resources in a cost-effective manner while satisfying the client service level agreements (SLAs). In this paper, we address the issue of how to intelligently manage the resources in a shared cloud database system and present SmartSLA, a cost-aware resource management system. SmartSLA consists of two main components: the system modeling module and the resource allocation decision module. The system modeling module uses machine learning techniques to learn a model that describes the potential profit margins for each client under different resource allocations. Based on the learned model, the resource allocation decision module dynamically adjusts the resource allocations in order to achieve the optimum profits. We evaluate SmartSLA by using the TPC-W benchmark with workload characteristics derived from real-life systems. The performance results indicate that SmartSLA can successfully compute predictive models under different hardware resource allocations, such as CPU and memory, as well as database specific resources, such as the number of replicas in the database systems. The experimental results also show that SmartSLA can provide intelligent service differentiation according to factors such as variable workloads, SLA levels, resource costs, and deliver improved profit margins.",https://ieeexplore.ieee.org/document/5767928,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 581}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 184}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 659}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1280}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1314}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 193}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 856}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 578}]",64.0,,,,['resource-consolidation'],,2011 IEEE 27th International Conference on Data Engineering,True,['resource-provisioning'],,,,,,,,,
64,AIOps for a Cloud Object Storage Service,"['A. Levin', ' S. Garion', ' E. K. Kolodner', ' D. H. Lorenz', ' K. Barabash', ' M. Kugler', ' N. McShane']",2019,"With the growing reliance on the ubiquitous availability of IT systems and services, these systems become more global, scaled, and complex to operate. To maintain business viability, IT service providers must put in place reliable and cost efficient operations support. Artificial Intelligence for IT Operations (AIOps) is a promising technology for alleviating operational complexity of IT systems and services. AIOps platforms utilize big data, machine learning and other advanced analytics technologies to enhance IT operations with proactive actionable dynamic insight. In this paper we share our experience applying the AIOps approach to a production cloud object storage service to get actionable insights into system's behavior and health. We describe a real-life production cloud scale service and its operational data, present the AIOps platform we have created, and show how it has helped us resolving operational pain points.",https://ieeexplore.ieee.org/document/8818188,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 586}, {'database': 'IEEE', 'search_string': ""'AIOps'"", 'index': 2}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1116}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 335}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('IT operations')"", 'index': 3}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('IT operations')"", 'index': 1}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1654}]",1.0,,,['discussion'],,['filesystem'],2019 IEEE International Congress on Big Data (BigDataCongress),True,['aiops-general'],,True,,,,,,,
65,The Beginning of a Cognitive Software Engineering Era with Self-Managing Applications,['G. Morana'],2018,"The recent explosion of data and the resurgence of AI, Machine Learning and Deep Learning, and the emergence of unbounded cloud computing resources are stretching current software engineering practices to meet business application development, deployment and management requirements. As consumers demand communication, collaboration and commerce almost at the speed of light without interruption, businesses are looking for information technologies that keep up the pace in delivering faster time to market and real-time data processing to meet rapid fluctuations in both workload demands and available computing resources. While the performance of server, network and storage resources have dramatically improved by orders of magnitude in the past decade, software engineering practices and IT operations are evolving at a slow pace. This paper explores a new approach that will provide a path to self-managing software systems with fluctuation tolerance to both workload demands and available resource pools. The infusion of a cognitive control overlay enables an advanced management of application workloads in a distributed multi-cloud computing infrastructure. Resulting architecture provides a uniform framework for managing workload non-functional requirements such as availability, performance, security, data compliance and cost independent of the execution venue for functional requirement workflows.",https://ieeexplore.ieee.org/document/8452780,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 589}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('IT operations')"", 'index': 4}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 229}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 75}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 223}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 794}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('IT operations')"", 'index': 7}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('IT operations')"", 'index': 7}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 176}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('IT operations')"", 'index': 17}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('IT operations')"", 'index': 7}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('IT operations')"", 'index': 3}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('IT operations')"", 'index': 5}]",0.0,,,,,,2018 IEEE/ACM 1st International Workshop on Software Engineering for Cognitive Services (SE4COG),True,['resource-provisioning'],,,,,,10.1145/3195555.3195557,Conference Paper,"['DIME network architecture', 'semantic network', 'modularity', 'dev-ops', 'turing O-machine', 'connectivity', 'cloud computing']",
66,Towards AI-Powered Multiple Cloud Management,"['B. D. Martino', ' A. Esposito', ' E. Damiani']",2019,"Cloud users worldwide are looking at the next generation of artificial intelligence (AI) powered cloud management tools to automate cloud performance tuning and anomaly detection. To be effective across clouds, AI tools need a common representation of cloud services and support for machine learning optimization targeting multiple objectives. We put forward the notion that ontology-based models can support both.",https://ieeexplore.ieee.org/document/8662800,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 600}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 127}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 43}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1139}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 106}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1289}]",2.0,,,,"['resource-consolidation', 'service-composition']",,IEEE Internet Computing,True,['resource-provisioning'],,,,,,,,,
67,Adaptive Scheduling on Power-Aware Managed Data-Centers Using Machine Learning,"['J. L. Berral', ' R. Gavalda', ' J. Torres']",2011,"Energy-related costs have become one of the major economic factors in IT data-centers, and companies and the research community are currently working on new efficient power-aware resource management strategies, also known as ""Green IT"". Here we propose a framework for autonomic scheduling of tasks and web-services on cloud environments, optimizing the profit taking into account revenue for task execution minus penalties for service-level agreement violations, minus power consumption cost. The principal contribution is the combination of consolidation and virtualization technologies, mathematical optimization methods, and machine learning techniques. The data-center infrastructure, tasks to execute, and desired profit are casted as a mathematical programming model, which can then be solved in different ways to find good task scheduling. We use an exact solver based on mixed linear programming as a proof of concept but, since it is an NP-complete problem, we show that approximate solvers provide valid alternatives for finding approximately optimal schedules. The machine learning is used to estimate the initially unknown parameters of the mathematical model. In particular, we need to predict a priori resource usage (such as CPU consumption) by different tasks under current workloads, and estimate task service-level-agreement (such as response time) given workload features, host characteristics, and contention among tasks in the same host. Experiments show that machine learning algorithms can predict system behavior with acceptable accuracy, and that their combination with the exact or approximate schedulers manages to allocate tasks to hosts striking a balance between revenue for executed tasks, quality of service, and power consumption.",https://ieeexplore.ieee.org/document/6076500,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 626}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 43}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 168}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 345}]",26.0,,,,,,2011 IEEE/ACM 12th International Conference on Grid Computing,True,['resource-provisioning'],,,,,,,,,
68,Empirical study based on machine learning approach to assess the QoS/QoE correlation,"['M. S. Mushtaq', ' B. Augustin', ' A. Mellouk']",2012,"The appearance of new emerging multimedia services have created new challenges for cloud service providers, which have to react quickly to end-users experience and offer a better Quality of Service (QoS). Cloud service providers should use such an intelligent system that can classify, analyze, and adapt to the collected information in an efficient way to satisfy end-users' experience. This paper investigates how different factors contributing the Quality of Experience (QoE), in the context of video streaming delivery over cloud networks. Important parameters which influence the QoE are: network parameters, characteristics of videos, terminal characteristics and types of users' profiles. We describe different methods that are often used to collect QoE datasets in the form of a Mean Opinion Score (MOS). Machine Learning (ML) methods are then used to classify a preliminary QoE dataset collected using these methods. We evaluate six classifiers and determine the most suitable one for the task of QoS/QoE correlation.",https://ieeexplore.ieee.org/document/6249939,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 631}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 1023}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 1348}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 540}]",38.0,,,,,,2012 17th European Conference on Networks and Optical Communications,True,['resource-provisioning'],,,,,,,,,
69,Online approach to performance fault localization for cloud and datacenter services,"['J. Ahmed', ' A. Johnsson', ' F. Moradi', ' R. Pasquini', ' C. Flinta', ' R. Stadler']",2017,Automated detection and diagnosis of the performance faults in cloud and datacenter environments is a crucial task to maintain smooth operation of different services and minimize downtime. We demonstrate an effective machine learning approach based on detecting metric correlation stability violations (CSV) for automated localization of performance faults for datacenter services running under dynamic load conditions.,https://ieeexplore.ieee.org/document/7987390,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 658}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 1018}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1871}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1088}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 14}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 16}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 15}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 14}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 436}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1144}]",2.0,,['kpis'],,"['root-cause-analysis', 'failure-detection']",,2017 IFIP/IEEE Symposium on Integrated Network and Service Management (IM),True,['failure-management'],,True,,,,,,,
70,Automated diagnostic of virtualized service performance degradation,"['J. Ahmed', ' T. Josefsson', ' A. Johnsson', ' C. Flinta', ' F. Moradi', ' R. Pasquini', ' R. Stadler']",2018,"Service assurance for cloud applications is a challenging task and is an active area of research for academia and industry. One promising approach is to utilize machine learning for service quality prediction and fault detection so that suitable mitigation actions can be executed. In our previous work, we have shown how to predict service-level metrics in real-time just from operational data gathered at the server side. This gives the service provider early indications on whether the platform can support the current load demand. This paper provides the logical next step where we extend our work by proposing an automated detection and diagnostic capability for the performance faults manifesting themselves in cloud and datacenter environments. This is a crucial task to maintain the smooth operation of running services and minimizing downtime. We demonstrate the effectiveness of our approach which exploits the interpretative capabilities of Self- Organizing Maps (SOMs) to automatically detect and localize different performance faults for cloud services.",https://ieeexplore.ieee.org/document/8406234,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 673}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 781}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 739}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 337}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 56}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 45}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 119}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 29}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 218}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 618}]",2.0,['som'],,,"['root-cause-analysis', 'failure-detection']",,NOMS 2018 - 2018 IEEE/IFIP Network Operations and Management Symposium,True,['failure-management'],,,"['anomaly-detection', 'fault-localization']",,,,,,
71,Scalable Prediction of Service-Level Events in Datacenter Infrastructure Using Deep Neural Networks,"['A. Mozo', ' I. Segall', ' U. Margolin', ' S. Gómez-Canaval']",2019,"The complexity of cloud datacenter scenarios poses new challenges in infrastructure management processes such as the impracticality of collecting specific service level events from inside the datacenter infrastructure, and the scalability issues that can appear during event monitoring when thousands of virtual machines have to be polled at a granularity of seconds. Therefore, it would be desirable to provide mechanisms for obtaining these types of events without incurring in the previously described problems. To this end, we propose a generic and scalable method based on the application of deep neural network architectures for predicting service level events using only a reduced number of generic datacenter infrastructure statistics that can be monitored in a scalable way. We demonstrate in a controlled scenario of a real datacenter and using only three variables from a physical machine that it is possible to predict events in real-time and with decent accuracy, without needing to deploy any meter in the end-user equipment. Specifically, we demonstrate this over two service-level events: i) the so-called Noisy Neighbors effect, a harmful situation that appears in physical machines due to the interferences created by the interaction of virtual machines running on them; and ii) the jitter values of a multimedia call running in a virtual machine. We set up a testbed in a real datacenter deploying physical and virtual machines, running a large amount of different experiments for 1000 hours and collecting samples at a 10 seconds granularity in a dataset of 260,000 records. Two different scenarios, in which training and testing data sets contain significant statistical differences, are deployed to demonstrate a better generalization ability of deep models in changing scenarios when compared with traditional Machine Learning techniques. A set of different deep architectures are proposed for both use cases and approximately 4,000 deep models were trained and tested. In both use cases, the best deep models show a good performance when predicting service level events, even if the inputs do not exactly follow the statistical patterns of the data used during training.",https://ieeexplore.ieee.org/document/8915840,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 677}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 420}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 633}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 500}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 11}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 317}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1829}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 342}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 891}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 307}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 1610}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 18}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 47}, {'database': 'IEEE', 'search_string': ""'classification' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 26}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 11}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 402}, {'database': 'IEEE', 'search_string': ""'regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 12}]",0.0,['multilayer-perceptron'],,"['comparison', 'novel-use']",['failure-prediction'],,,True,['failure-management'],,True,['system-failure-prediction'],,,,,,
72,LASER: A Deep Learning Approach for Speculative Execution and Replication of Deadline-Critical Jobs in Cloud,"['M. Xu', ' S. Alamro', ' T. Lan', ' S. Subramaniam']",2017,"Meeting desired application deadlines is crucial as the nature of cloud applications is becoming increasingly mission-critical and deadline-sensitive. Empirical studies on large-scale clusters reveal that a few slow tasks, known as stragglers, could significantly stretch job execution times. A number of strategies are proposed to mitigate stragglers by launching speculative or clone (task) attempts. These strategies often rely on a model-based approach to optimize key operating parameters and are prone to inaccuracy/incompleteness in the underlying models. In this paper, we present LASER, a deep learning approach for speculative execution and replication of deadline-critical jobs. Machine learning has been successfully used to solve a large variety of classification and prediction problems. In particular, the deep neural network (DNN), consisting of multiple hidden layers of units between input and output layers, can provide more accurate regression (prediction) than traditional machine learning algorithms. We compare LASER with SRQuant, a speculative- resume strategy that is based on quantitative analysis. Both these scheduling algorithms aim to improve Probability of Completion before Deadlines (PoCD), i.e., the probability that MapReduce jobs meet their desired deadlines, and reduce the cost of speculative execution, measured by the total (virtual) machine time. We evaluate and compare the two strategies through testbed experiments. The results show that our two strategies outperform Hadoop without speculation (Hadoop-NS) and Hadoop with speculation (Hadoop-S) by up to 89% in PoCD and 13% in cost.",https://ieeexplore.ieee.org/document/8038373,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 682}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 425}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 857}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 221}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 229}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 697}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 936}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1648}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 432}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 1020}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 779}]",3.0,,,,,,2017 26th International Conference on Computer Communication and Networks (ICCCN),True,['resource-provisioning'],,,,,,,,,
73,Unsupervised Learning of Dynamic Resource Provisioning Policies for Cloud-Hosted Multitier Web Applications,"['W. Iqbal', ' M. N. Dailey', ' D. Carrera']",2016,"Dynamic resource provisioning for Web applications allows for low operational costs while meeting service-level objectives (SLOs). However, the complexity of multitier Web applications makes it difficult to automatically provision resources for each tier without human supervision. In this paper, we introduce unsupervised machine learning methods to dynamically provision multitier Web applications, while observing user-defined performance goals. The proposed technique operates in real time and uses learning techniques to identify workload patterns from access logs, reactively identifies bottlenecks for specific workload patterns, and dynamically builds resource allocation policies for each particular workload. We demonstrate the effectiveness of the proposed approach in several experiments using synthetic workloads on the Amazon Elastic Compute Cloud (EC2) and compare it with industry-standard rule-based autoscale strategies. Our results show that the proposed techniques would enable cloud infrastructure providers or application owners to build systems that automatically manage multitier Web applications, while meeting SLOs, without any prior knowledge of the applications' resource utilization or workload patterns.",https://ieeexplore.ieee.org/document/7111215,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 690}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 151}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 98}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 821}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 80}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1168}]",14.0,,,,['resource-consolidation'],,IEEE Systems Journal,True,['resource-provisioning'],,,,,,,,,
74,Workload prediction model based on supervised learning for energy efficiency in cloud,"['N. Verma', ' A. Sharma']",2017,"With the rapid emergence of cloud and its magnificent features such as high availability, reliability and security, a large number of clients are moving to the cloud platform. This migration of clients has led to increased burden on the resources such as CPU, network and memory. Hence, the problem of high power consumption, increased carbon footprints and need for higher cooling effects arose. However, resource scheduling and provisioning has become a milestone to handle such diverse issues related to cloud and get them under centralized control system. Workload prediction is an utmost requirement to dynamically predict the incoming workload and schedule the resources according to client needs. Dynamic workload prediction based on historical data can prove useful for pattern matching of current scenario with the past ones and allocate the resources in the most efficient way. This paper aims to propose a prototype model for workload prediction using machine learning models to handle the dynamic nature of the cloud infrastructure. This prediction is then used for resource provisioning. Diverse scheduling scenarios over the cloud are also covered in this paper and alternate remedies have been suggested for low power consumption and temperature control.",https://ieeexplore.ieee.org/document/8066526,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 696}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 257}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 184}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1166}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1285}]",3.0,,,,"['power-management', 'resource-consolidation']",,"2017 2nd International Conference on Communication Systems, Computing and IT Applications (CSCITA)",True,['resource-provisioning'],,,,,,,,,
75,Failure Prediction Model for Predictive Maintenance,"['K. Mishra', ' S. K. Manjhi']",2018,"As financial organizations strive to deliver superior omnichannel customer experiences, they are transforming their branch environments with latest digital technologies for ATMs, Branch platforms, self-service devices and other branch technologies. Simultaneously, mixing new with older, installed technologies from multiple vendors can create complex maintenance challenges. One could opt for each individual vendor's solution, but this can add complexity and may not put the crucial needs of the customer first. To maintain a customer-centric approach that leads to a high-quality brand image, improved customer satisfaction and ultimately a better bottom line, there is a need for service-oriented, vendor focused approach on delivering an integrated maintenance and technical support strategy, so that concentration on customers can be accomplished. In this direction, predictive maintenance plays a very vital role in enabling financial organizations to drive their ATM and branch business effectively to create maximum impact through predictive maintenance leveraging predictive analytics and machine learning technologies. We propose a method and Machine Learning model that takes various input data and determines likelihood of failure at a device and its component level within a stipulated future time-period with certain accuracy and precision for financial clients.",https://ieeexplore.ieee.org/document/8648634,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 699}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 658}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1497}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 5}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 11}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 3}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 1386}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 268}]",309.0,['random-forest'],"['logs', 'tickets']",,['failure-prediction'],,2018 IEEE International Conference on Cloud Computing in Emerging Markets (CCEM),True,['failure-management'],,,['system-failure-prediction'],,,,,,
76,Analyzing the Effectiveness of Machine Learning Algorithms for Determining Faulty Classes: A Comparative Analysis,"['P. Singh', ' R. Malhotra', ' S. Bansal']",2019,"The quality of the software can be improved by determining its faulty portions in the initial phases of the lifecycle of a software product. There are various machine learning algorithms proposed in literature studies that can be used to predict faulty classes. The machine learning algorithms determine faulty classes by using object oriented metrics as predictors. These models will allow the developers to predict faulty classes and concentrate the constraint resources in testing these weaker portions of the software. This study evaluates and compares the predictive capability of six machine learning algorithms amongst themselves and with logistic regression, a statistical algorithm for determining faulty portions of a software. The results are validated using seven open source software.",https://ieeexplore.ieee.org/document/8776946,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 706}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 536}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 405}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 319}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 23}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 202}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 327}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 76}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 104}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('cloud computing')"", 'index': 48}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('cloud')"", 'index': 63}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 50}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 72}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 279}]",0.0,"['random-forest', 'naive-bayes', 'logistic-regression', 'decision-tree']",['code-metrics'],"['comparison', 'novel-use']",['failure-prevention'],['source-code'],"2019 9th International Conference on Cloud Computing, Data Science & Engineering (Confluence)",True,['failure-management'],,True,['software-defect-prediction'],,,,,,
77,Towards systems level prognostics in the Cloud,"['B. Deb', ' M. Shah', ' S. Evans', ' M. Mehta', ' A. Gargulak', ' T. Lasky']",2013,"Many application systems are transforming from device centric architectures to cloud based systems that leverage shared compute resources to reduce cost and maximize reach. These systems require new paradigms to assure availability and quality of service. In this paper, we discuss the challenges in assuring Availability and Quality of Service in a Cloud Based Application System. We propose machine learning techniques for monitoring systems logs to assess the health of the system. A web services data set is employed to show that variety of services can be clustered to different service classes using a k-means clustering scheme. Reliability, Availability, and Serviceability (RAS) logs and Job logs dataset from high performance computing system is employed to show that impending fatal errors in the system can be predicted from the logs using an SVM classifier. These approaches illustrate the feasibility of methods to monitor the systems health and performance of compute resources and hence can be used to manage these systems for high availability and quality of service for critical tasks such as health care monitoring in the cloud.",https://ieeexplore.ieee.org/document/6621449,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 745}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 305}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 1130}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 142}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 272}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 261}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 1503}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1258}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 968}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 259}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 442}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 320}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1650}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 658}]",4.0,"['dimensionality-reduction', 'support-vector-machine', 'clustering']",['logs'],['novel-use'],['failure-prediction'],,2013 IEEE Conference on Prognostics and Health Management (PHM),True,['failure-management'],,True,['system-failure-prediction'],,,,,,
78,Using Deep Learning to Predict and Optimize Hadoop Data Analytic Service in a Cloud Platform,"['C. Chen', ' Y. Hasio', ' C. Lin', ' S. Lu', ' H. Lu', ' J. Chou']",2017,"Hadoop is a popular computing framework to deliver timely and cost-effective data processing on a large cluster of commodity machines. It relieves the burden of the programmers dealing with distributed programming, and an ecosystem of Big Data solutions have developed around it. However, Hadoop job execution time can be greatly depending on the its runtime configurations and resource selections. Hence, optimizing Hadoop execution still requires a substantial amount of expertise and experiences. To address this challenge, this paper aims to develop a learning-based technique to predict Hadoop job time based on historical execution data, and built a system to optimize its performance in a shared resource cloud environment. While deep learning has shown successes in many application domains, little attention has been paid to apply such technique in job time prediction. In this work, we conducted extensive experimental studies to compare our deep learning prediction method with three other state-of-art regression-based prediction methods in a cloud platform built by OpenStack. The results showed that our prediction method out-performed traditional approaches in most test cases, and improved job performance and cost significantly.",https://ieeexplore.ieee.org/document/8328497,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 763}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 203}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 90}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 138}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1321}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 293}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1239}]",0.0,['multilayer-perceptron'],,"['comparison', 'novel-use']",['workload-prediction'],"['hadoop', 'openstack']","2017 IEEE 15th Intl Conf on Dependable, Autonomic and Secure Computing, 15th Intl Conf on Pervasive Intelligence and Computing, 3rd Intl Conf on Big Data Intelligence and Computing and Cyber Science and Technology Congress(DASC/PiCom/DataCom/CyberSciTech)",True,['resource-provisioning'],,,,,,,,,
79,Task Runtime Prediction in Scientific Workflows Using an Online Incremental Learning Approach,"['M. H. Hilman', ' M. A. Rodriguez', ' R. Buyya']",2018,"Many algorithms in workflow scheduling and resource provisioning rely on the performance estimation of tasks to produce a scheduling plan. A profiler that is capable of modeling the execution of tasks and predicting their runtime accurately, therefore, becomes an essential part of any Workflow Management System (WMS). With the emergence of multi-tenant Workflow as a Service (WaaS) platforms that use clouds for deploying scientific workflows, task runtime prediction becomes more challenging because it requires the processing of a significant amount of data in a near real-time scenario while dealing with the performance variability of cloud resources. Hence, relying on methods such as profiling tasks' execution data using basic statistical description (e.g., mean, standard deviation) or batch offline regression techniques to estimate the runtime may not be suitable for such environments. In this paper, we propose an online incremental learning approach to predict the runtime of tasks in scientific workflows in clouds. To improve the performance of the predictions, we harness fine-grained resources monitoring data in the form of time-series records of CPU utilization, memory usage, and I/O activities that are reflecting the unique characteristics of a task's execution. We compare our solution to a state-of-the-art approach that exploits the resources monitoring data based on regression machine learning technique. From our experiments, the proposed strategy improves the performance, in terms of the error, up to 29.89%, compared to the state-of-the-art solutions.",https://ieeexplore.ieee.org/document/8603156,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 778}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 254}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 999}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1746}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 154}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 474}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 708}, {'database': 'arxiv', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 4}, {'database': 'arxiv', 'search_string': ""'regression' AND ('cloud')"", 'index': 20}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 98}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 224}]",1.0,,['kpis'],,,,2018 IEEE/ACM 11th International Conference on Utility and Cloud Computing (UCC),True,['resource-provisioning'],,,,,,10.1109/UCC.2018.00018,,,
80,Online Density Grid Pattern Analysis to Classify Anomalies in Cloud and NFV Systems,"['A. Acker', ' F. Schmidt', ' A. Gulenko', ' O. Kao']",2018,"Technologies like machine-to-machine communication, autonomous driving or virtual reality applications form an increasingly diverse service landscape. This entails individual and dynamic requirements regarding scalability, availability, latency or throughput from the underlying IT infrastructure. To meet those, telecommunication and network providers started a transformation process towards virtualized technologies like network function virtualization (NFV). However, this drastically increases the infrastructure complexity to a point where more autonomous management is required. In order to meet the reliability of dedicated hardware, virtualized solutions are in demand of autonomous recovery and remediation systems. For critical network systems, actions must be selected very cautiously to not disrupt the operational process. To enable a precise handling, anomaly situations need to be accurately identified based on monitoring data streams. Therefore, we present a supervised machine learning method for an online classification of anomaly states based on similarities between anomaly type-specific density grid patterns. For evaluation, we created an extensive NFV testbed running a virtual implementation of the IP multimedia subsystem. Applying our method to classify various synthetically injected anomaly situations, the results reveal an average overall accuracy of 0.94. Further results also show that the classification model is applicable for identifying previously unknown anomaly situations. Thus, our approach provides a valuable step towards autonomous maintenance of virtualized IT infrastructures.",https://ieeexplore.ieee.org/document/8591032,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 839}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 387}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 537}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 990}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 476}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 348}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1398}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 541}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 560}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 824}, {'database': 'IEEE', 'search_string': ""'classification' AND ('remediation' OR 'recovery')"", 'index': 387}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('remediation' OR 'recovery')"", 'index': 246}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1049}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1295}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 690}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 1047}]",0.0,,,,['failure-detection'],['network'],2018 IEEE International Conference on Cloud Computing Technology and Science (CloudCom),True,['failure-management'],,,,,,,,,
81,Neural Network Based Classification of Virtual Machines in IaaS,"['E. Patel', ' A. Mohan', ' D. S. Kushwaha']",2018,"Machine Learning algorithms are being widely used to make intelligent systems. In Cloud Computing, issues like server consolidation, load balancing and virtual machine to physical machine mapping have been addressed using clustering analysis which is unsupervised learning. It is important that Cloud Service Providers address these challenges in order to provide enhanced user experience and superlative Quality of Service. Neural Networks, which fall under supervised learning algorithms, have emerged tremendously and are used in different areas such as Bio-Informatics, Pattern Recognition and Classification. Using Neural Networks and Back Propagation Algorithm, this work classifies different types of virtual machines that are providing heterogeneous services. Further, the technique is compared to K-means Clustering Algorithm which groups similar types of virtual machines into several clusters. The objective of this work is to form labeled clusters of virtual machines with the help of Neural Networks based on their types and loads and perform server consolidation. The proposed work achieves higher performance in terms of number of virtual machine samples correctly classified, as compared to K-Means Clustering algorithm. Migrating dissimilar virtual machines onto a single physical host improved machine utilization with the trade-off of slight increase in response time of the machines.",https://ieeexplore.ieee.org/document/8596875,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 866}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 212}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 357}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 380}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 200}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 305}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 155}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 336}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 528}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 644}]",0.0,,,,,,"2018 5th IEEE Uttar Pradesh Section International Conference on Electrical, Electronics and Computer Engineering (UPCON)",True,['resource-provisioning'],,True,,,,,,,
82,Fitness-Aware Containerization Service Leveraging Machine Learning,"['S. Venkateswaran', ' S. Sarkar']",2019,"Containerized deployment of microservices has gained immense traction across industries. To meet demand, traditional cloud providers offer container-as-a-service, where selection of the container and containerization of workloads remain developer's responsibility. This task is arduous for a developer since the choice of containers across different cloud providers is many. Furthermore, there does not exist any mechanism using which one can compare and contrast the capabilities of containers across different providers. In this scenario, we envisage the need for a smart cloud broker that can automatically deploy a chosen IT service into the best-fit container environment mapped to performance requirements, from among the set of available underpinning brokered container hosting systems spread across multiple cloud providers. We propose a novel fitness-aware containerization-as-a-service to achieve this. We show why a best-fit container selection process is operationally complex and time consuming, and how we heuristically prune the associated decision tree in two phases so that it becomes viable to implement this as an on-demand service. We propose a new metric called fitness quotient (FQ) to evaluate containers obtained from heterogeneous providers. We leverage machine learning techniques to inject automation into these two phases: unsupervised K-Means clustering in the first-level build-time phase to accurately classify IaaS cost and performance data, and polynomial regression during the second-level provisioning-time phase to discover relationships between SaaS performance and container strength.",https://ieeexplore.ieee.org/document/8638583,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 876}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 494}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 397}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 1192}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 1015}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 671}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 339}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 1380}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 710}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 1488}]",0.0,"['linear-regression', 'clustering']",,,"['configuration', 'workload-prediction']",['container'],IEEE Transactions on Services Computing,True,['resource-provisioning'],,True,,,,,,,
83,Identifying resources for cloud garbage collection,"['Z. Shen', ' C. C. Young', ' S. Zeng', ' K. Murthy', ' K. Bai']",2016,"Infrastructure as a Service (IaaS) clouds provide users with the ability to easily and quickly provision servers. A recent study found that one in three data center servers continues to consume resources without producing any useful work. A number of techniques have been proposed to identify such unproductive instances. However, those approaches adopt the strategy to identify idle cloud instances based on resource utilization. Resource utilization as indicator alone could be misleading, which is especially true for enterprise cloud environment. In this paper, we present Pleco, a tool that detects unproductive instances in IaaS clouds. Pleco captures dependency information between users and cloud instances by constructing a weighted reference model based on application knowledge. To handle cases of insufficient application knowledge, Pleco also supplements its dependency results with a machine learning model trained on resource utilization data. Pleco gives a confidence level and justification for each identified unproductive instances. Cloud administrators can then take different actions according to the information provided by Pleco. Pleco is lightweight and requires no modification to existing applications.",https://ieeexplore.ieee.org/document/7818426,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 894}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 906}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 72}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 825}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 149}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1521}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 26}, {'database': 'ACM', 'search_string': ""'classification' AND ('cloud')"", 'index': 24}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 71}]",6.0,['decision-tree'],,['novel-use'],"['failure-detection', 'anomaly-detection']",['vm'],2016 12th International Conference on Network and Service Management (CNSM),True,"['resource-provisioning', 'failure-management']",,,,,,,Conference Paper,"['Pleco', 'Zombie Servers', 'Unproductive Instances', 'Infrastructure as a Service', 'Resource Utilization', 'Cloud']",
84,Supporting Autonomic Management of Clouds: Service Clustering With Random Forest,"['R. B. Uriarte', ' F. Tiezzi', ' S. A. Tsaftaris']",2016,"A promising solution for the management of services in clouds, as fostered by autonomic computing, is to resort to self-management. However, the obfuscation of underlying details of services in cloud computing, also due to privacy requirements, affects the effectiveness of autonomic managers. Data-driven approaches, in particular those relying on service clustering based on machine learning techniques, can assist the autonomic management and support decisions concerning, e.g., the scheduling and deployment of services. Unfortunately, applying such approaches is further complicated by the coexistence of different types of data within the information provided by the monitoring of cloud systems: both continuous (e.g., CPU load) and categorical (e.g., VM instance type) data are available. Current approaches deal with this problem in a heuristic fashion. In this paper, instead, we propose an approach that uses all types of data, and learns in a data-driven fashion the similarities and patterns among the services. More specifically, we design an unsupervised formulation of random forest to calculate service similarities and provide them as input to a clustering algorithm. For the sake of efficiency and to meet the dynamism requirement of autonomic clouds, our methodology consists of two steps: 1) off-line clustering and 2) on-line prediction. Using datasets from real-world clouds, we demonstrate the superiority of our solution with respect to others and validate the accuracy of the on-line prediction. Moreover, to show applicability of our approach, we devise a service scheduler that uses similarity among services, and evaluate its performance in a cloud test-bed using realistic data.",https://ieeexplore.ieee.org/document/7469788,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 920}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 230}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 209}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 168}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 324}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 377}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 588}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 375}]",7.0,,,,,,IEEE Transactions on Network and Service Management,True,['resource-provisioning'],,True,,,,,,,
85,Modeling Application Performance in Docker Containers Using Machine Learning Techniques,"['K. Ye', ' Y. Kou', ' C. Lu', ' Y. Wang', ' C. Xu']",2018,"Docker container is experiencing a rapid development with the support from industry like Google and is being widely used in large scale production cloud environments. However the performance of applications running in Docker containers is still not clear due to the complex relationship between container resource allocation and application performance. In this paper, we first study the impact of key parameters in container resource allocation that affect the performance of containerized applications. Then, we present modeling techniques over CPU, memory and I/O resources to characterize the performance of applications running in containers. To address this multi-dimensional modeling problem, we propose three machine learning techniques, i.e. Linear Regression (LR), Support Vector Machine (SVM) and Artificial Neural Network (ANN). We implement and evaluate the modeling techniques for four complex benchmark workloads from Spark. Experimental results demonstrate the proposed models can achieve as low as 2.27% prediction error, with an average of 10.13% for most applications. Furthermore, the prediction accuracy of SVM and ANN models are substantially better than LR based approaches, with 48.13% and 29.30% improvement.",https://ieeexplore.ieee.org/document/8644581,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 928}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 404}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 1182}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1967}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 689}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 351}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 280}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 223}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 603}]",1.0,"['multilayer-perceptron', 'linear-regression', 'support-vector-machine']",,,,['container'],2018 IEEE 24th International Conference on Parallel and Distributed Systems (ICPADS),True,['resource-provisioning'],,,,,,,,,
86,A deep learning approach for VM workload prediction in the cloud,"['F. Qiu', ' B. Zhang', ' J. Guo']",2016,"In order to manage the resources in cloud efficiently, ensure the performance of cloud services and reduce the power consumption, it is critical to predict the workload of virtual machines (VM) accurately. In this paper, a new approach for VM workload prediction based on deep learning was proposed. A deep learning prediction model was designed with a deep belief network (DBN) composed of multiple-layered restricted Boltzmann machines (RBMs) and a regression layer. The DBN is used to extract the high level features from all VMs workload data and the regression layer is used to predict the workload of the VMs in the future. With little prior knowledge, DBN could learn the features efficiently for the VM workload prediction in an unsupervised fashion. Experimental results show that the proposed approach improves the workload prediction performance compared with other widely used workload prediction approaches.",https://ieeexplore.ieee.org/document/7515919,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 936}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 279}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 224}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 123}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 107}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 745}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 317}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 975}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 220}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 939}]",20.0,,,,['workload-prediction'],,"2016 17th IEEE/ACIS International Conference on Software Engineering, Artificial Intelligence, Networking and Parallel/Distributed Computing (SNPD)",True,['resource-provisioning'],,True,,,,,,,
87,Black-box approach to capacity identification for multi-tier applications hosted on virtualized platforms,"['W. Iqbal', ' M. N. Dailey', ' D. Carrera']",2011,"In cloud-based Web application hosting environments, virtualization offers the potential to exploit dynamic resource provisioning and scaling to maintain service level agreements while minimizing resource utilization for a given workload. However, optimal proactive resource provisioning and scaling for a specific Web application require, at the least, a profile of the application's current workload and a model of the application's capacity under various resource configurations. Here we focus on multi-tier Web applications. The capacity of a multi-tier Web application varies substantially as the pattern of requests in the workload changes. In this paper, we propose and evaluate a black-box method for capacity prediction that first identifies workload patterns for a multi-tier Web application from access logs using unsupervised machine learning and then, based on those patterns, builds a model capable of predicting the application's capacity for any specific workload pattern. In an experimental evaluation, we compare a baseline method that predicts capacity without a model of the application-specific workload patterns to several regression models using the proposed workload identification method. All of the models based on workload pattern identification outperform the baseline method. The best model, a Gaussian process regression model, gives only 6.42% error. Cloud providers utilizing our method can proactively perform dynamic allocation of resources to multi-tier Web applications, meeting service level agreements at minimal cost.",https://ieeexplore.ieee.org/document/6138506,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 964}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 400}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 317}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 410}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 842}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 353}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 238}, {'database': 'IEEE', 'search_string': ""'regression' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 442}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1256}]",10.0,,,,,,2011 International Conference on Cloud and Service Computing,True,['resource-provisioning'],,,,,,,,,
88,Job Aware Scheduling Algorithm for MapReduce Framework,"['R. Nanduri', ' N. Maheshwari', ' A. Reddyraja', ' V. Varma']",2011,"MapReduce framework has received a wide acclaim over the past few years for large scale computing. It has become a standard paradigm for batch oriented workloads. As the adoption of this paradigm has increased rapidly, scheduling of these MapReduce jobs has become a problem of great interest in research community. We propose an approach which tries to maintain harmony among the jobs running on the cluster, and in turn decrease their runtime. In our model, the scheduler is made aware of different types of jobs running on the cluster. The scheduler tries to allocate a task on a node if the incoming task does not affect the tasks already running on that node. From the list of available pending tasks, our algorithm selects the one that is most compatible with the tasks already running on that node. We bring up heuristic and machine learning based solutions to our approach and try to maintain a resource balance on the cluster by not overloading any of the nodes, thereby reducing the overall runtime of the jobs. The results show a saving of runtime of around 21% in the case of heuristic based approach and around 27% in the case of machine learning based approach when compared to Yahoo's Capacity scheduler.",https://ieeexplore.ieee.org/document/6133221,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 995}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 1196}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1665}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 437}]",27.0,,,,,,2011 IEEE Third International Conference on Cloud Computing Technology and Science,True,['resource-provisioning'],,,,,,,,,
89,Root cause analysis of software bugs using machine learning techniques,"['H. Lal', ' G. Pahwa']",2017,"Root cause analysis (RCA) is a systematic process for identifying “root causes” of problems or events and an approach for responding to them. The factor that caused a problem or defect should be permanently eliminated through process improvement. In the context of Software development process it may be used to refer to a specific module or a category of bug which in turn can be useful for tackling the problem at its root. In this paper we propose a machine learning approach for finding root cause of a newly filed software bugs which in turn would help in the faster and cleaner resolution of software bugs. This proposed approach is evaluated for feasibility study on an open source system eclipse. [7], [6]",https://ieeexplore.ieee.org/document/7943132,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1025}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 305}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 3}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 1536}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 1}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 508}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 4}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 8}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 490}]",2.0,"['naive-bayes', 'logistic-regression', 'support-vector-machine', 'decision-tree', 'dimensionality-reduction']","['reports', 'source-code']","['comparison', 'novel-use']",['root-cause-analysis'],['software'],"2017 7th International Conference on Cloud Computing, Data Science & Engineering - Confluence",True,['failure-management'],,True,['root-cause-diagnosis'],,,,,,
90,Datacenter Workload Classification and Characterization: An Empirical Approach,"['V. S. Shekhawat', ' A. Gautam', ' A. Thakrar']",2018,"Datacenter traffic has increased significantly due to rising number of web applications on Internet. These applications have diverse Quality of Service (QoS) requirements making datacenter management a complex task. For a datacenter the amount of resources required for a given resource type (computing, memory, network and storage) is termed as workload. In cloud datacenters, workload classification and characterization is used for resource management, application performance management, capacity sizing, and for estimating the future resource demand. An accurate estimation of future resource demand helps in meeting QoS requirements and ensure efficient resource utilization. Thus modeling and characterization of datacenter workloads becomes necessary to meet performance requirements of applications in a cost-efficient manner. In this paper, a methodology to classify datacenter workloads and characterize them based on resource usage is proposed. Two different workloads have been used, one is Google Cluster Trace (GCT) dataset and other is Bit Brains Trace (BBT) dataset. Seven different machine learning algorithms for workload classification have been used. Workload distribution is estimated in a mix of heterogeneous applications for both GCT and BBT. The seven machine learning algorithms have been compared on the basis of their classification accuracy. Finally, an algorithm to estimate the importance of different attributes for classification is proposed in this paper.",https://ieeexplore.ieee.org/document/8721402,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1030}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 539}, {'database': 'IEEE', 'search_string': ""'classification' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 172}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 386}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 775}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 4}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 5}, {'database': 'IEEE', 'search_string': ""'classification' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 1}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 235}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 736}]",0.0,"['random-forest', 'logistic-regression', 'support-vector-machine', 'multilayer-perceptron', 'decision-tree', 'clustering']",['host-metrics'],['comparison'],['workload-prediction'],,2018 IEEE 13th International Conference on Industrial and Information Systems (ICIIS),True,['resource-provisioning'],,True,,,,,,,
91,Predicting of Job Failure in Compute Cloud Based on Online Extreme Learning Machine: A Comparative Study,"['C. Liu', ' J. Han', ' Y. Shang', ' C. Liu', ' B. Cheng', ' J. Chen']",2017,"Early prediction of job failures and specific disposal steps in advance could significantly improve the efficiency of resource utilization in large-scale data center. The existing machine learning-based prediction methods commonly adopt offline working pattern, which cannot be used for online prediction in practical operations, in which data arrive sequentially. To solve this problem, a new method based on online sequential extreme learning machine (OS-ELM) is proposed in this paper to predict online job termination status. With this method, real-time data are collected according to the sequence of job arriving, the job status could be predicted and the operation model is thus updated based on these data. The method with online incremental learning strategy has fast learning speed and good generalization. Comparative study using Google trace data shows that prediction accuracy of the proposed method is 93% with updating model in 0.01 s. Compared with some state-of-the-art methods, such, as support vector machine (SVM), ELM, and OS-SVM, the method developed in this paper has many advantages, such as less time-consuming in establishing and updating the model, higher prediction accuracy and precision, and better false negative performance.",https://ieeexplore.ieee.org/document/7932064,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1048}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 434}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 505}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 353}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 230}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 118}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('remediation' OR 'recovery')"", 'index': 167}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1761}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 191}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 75}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 323}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 339}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1612}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 913}]",8.0,['multilayer-perceptron'],"['sla', 'host-metrics']","['comparison', 'novel-use']",['failure-prediction'],['job'],,True,['failure-management'],,True,['system-failure-prediction'],,,,,,['online']
92,Non-Intrusive Anomaly Detection With Streaming Performance Metrics and Logs for DevOps in Public Clouds: A Case Study in AWS,"['D. Sun', ' M. Fu', ' L. Zhu', ' G. Li', ' Q. Lu']",2016,"Public clouds are a style of computing platforms, where scalable and elastic Information Technology-enabled capabilities are provided as a service to external customers using Internet technologies. Using public cloud services can reduce costs and increase the choices of technologies, but it also implies limited system information for users. Thus, anomaly detection at user end has to be non-intrusive and hence difficult, particularly during DevOps operations because the impacts from both anomalies and these operations are often indistinguishable, and hence, it is hard to detect the anomalies. In this paper, our work is specific to a successful public cloud, Amazon Web Service, and a representative DevOps operation, rolling upgrade, on which we report our anomaly detection that can effectively detect anomalies. Our anomaly detection requires only metrics data and logs supplied by most public clouds officially. We use support vector machine to train multiple classifiers from monitored data for different system environments, on which the log information can indicate the best suitable classifier. Moreover, our detection aims at finding anomalies over every time interval, called window, such that the features include not only some indicative performance metrics but also the entropy and the moving average of metrics data in each window. Our experimental evaluation systematically demonstrates the effectiveness of our approach.",https://ieeexplore.ieee.org/document/7389388,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1108}, {'database': 'IEEE', 'search_string': ""'classification' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 7}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 59}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 375}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 7}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 51}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 1203}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 71}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 1908}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 299}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('anomaly detection' OR 'outlier detection')"", 'index': 149}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 133}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 252}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 57}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 179}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 30}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 130}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 826}]",23.0,"['support-vector-machine', 'entropy-selection']","['software-metrics', 'host-metrics', 'logs']",['novel-use'],['failure-detection'],"['streaming', 'aws']",IEEE Transactions on Emerging Topics in Computing,True,['failure-management'],,True,['anomaly-detection'],,,,,,['non-intrusive']
93,DRL-Scheduling: An Intelligent QoS-Aware Job Scheduling Framework for Applications in Clouds,"['Y. Wei', ' L. Pan', ' S. Liu', ' L. Wu', ' X. Meng']",2018,"As an increasing number of traditional applications migrated to the cloud, achieving resource management and performance optimization in such a dynamic and uncertain environment becomes a big challenge for cloud-based application providers. In particular, job scheduling is a non-trivial task, which is responsible for allocating massive job requests submitted by users to the most suitable resources and satisfying user QoS requirements as much as possible. Inspired by recent success of using deep reinforcement learning techniques to solve AI control problems, in this paper, we propose an intelligent QoS-aware job scheduling framework for application providers. A deep reinforcement learning-based job scheduler is the key component of the framework. It is able to learn to make appropriate online job-to-VM decisions for continuous job requests directly from its experiences without any prior knowledge. Experimental results using synthetic workloads and real-world NASA workload traces show that compared with other baseline solutions, our proposed job scheduling approach can efficiently reduce average job response time (e.g., reduced by 40.4% compared with the best baseline for NASA traces), guarantee the QoS at a high level (e.g., job success rate is higher than 93% for all simulated changing workload scenarios), and adapt to different workload conditions.",https://ieeexplore.ieee.org/document/8476582,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1131}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 185}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 197}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1166}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1151}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 245}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1760}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 698}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 829}]",5.0,,,,,,,True,['resource-provisioning'],,True,,,,,,,
94,Toward a Smart Cloud: A Review of Fault-tolerance Methods in Cloud Systems,"['M. A. Mukwevho', ' T. Celik']",2018,"This paper presents a comprehensive survey of the state-of-the-art work on fault tolerance methods proposed for cloud computing. The survey classifies fault-tolerance methods into three categories: 1) ReActive Methods (RAMs); 2) PRoactive Methods (PRMs); and 3) ReSilient Methods (RSMs). RAMs allow the system to enter into a fault status and then try to recover the system. PRMs tend to prevent the system from entering a fault status by implementing mechanisms that enable them to avoid errors before they affect the system. On the other hand, recently emerging RSMs aim to minimize the amount of time it takes for a system to recover from a fault. Machine Learning and Artificial Intelligence have played an active role in RSM domain in such a way that the recovery time is mapped to a function to be optimized (i.e by converging the recovery time to a fraction of milliseconds). As the system learns to deal with new faults, the recovery time will become shorter. In addition, current issues and challenges in cloud fault tolerance are also discussed to identify promising areas for future research.",https://ieeexplore.ieee.org/document/8318693,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1182}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 919}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1788}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 1900}]",24.0,,,['survey'],"['failure-prediction', 'remediation', 'failure-prevention']",,IEEE Transactions on Services Computing,True,['failure-management'],True,True,,,,,,,
95,A Global Manufacturing Big Data Ecosystem for Fault Detection in Predictive Maintenance,"['W. Yu', ' T. Dillon', ' F. Mostafa', ' W. Rahayu', ' Y. Liu']",2019,"Artificial intelligence, big data, machine learning, cloud computing, and Internet of Things (IoT) are terms which have driven the fourth industrial revolution. The digital revolution has transformed the manufacturing industry into smart manufacturing through the development of intelligent systems. In this paper, a big data ecosystem is presented for the implementation of fault detection and diagnosis in predictive maintenance with real industrial big data gathered directly from large-scale global manufacturing plants, aiming to provide a complete architecture which could be used in industrial IoT-based smart manufacturing in an industrial 4.0 system. The proposed architecture overcomes multiple challenges including big data ingestion, integration, transformation, storage, analytics, and visualization in a real-time environment using various technologies such as the data lake, NoSQL database, Apache Spark, Apache Drill, Apache Hive, OPC Collector, and other techniques. Transformation protocols, authentication, and data encryption methods are also utilized to address data and network security issues. A MapReduce-based distributed PCA model is designed for fault detection and diagnosis. In a large-scale manufacturing system, not all kinds of failure data are accessible, and the absence of labels precludes all the supervised methods in the predictive phase. Furthermore, the proposed framework takes advantage of some of the characteristics of PCA such as its ease of implementation on Spark, its simple algorithmic structure, and its real-time processing ability. All these elements are essential for smart manufacturing in the evolution to Industry 4.0. The proposed detection system has been implemented into the real-time industrial production system in a cooperated company, running for several years, and the results successfully provide an alarm warning several days before the fault happens. A test case involving several outages in 2014 is reported and analyzed in detail during the experiment section.",https://ieeexplore.ieee.org/document/8710319,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1201}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 406}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 229}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 524}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 133}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 653}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1495}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 191}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 581}]",4.0,"['dimensionality-reduction', 'clustering']",,['new-method'],"['root-cause-analysis', 'failure-detection']","['spark', 'mapreduce', 'nosql']",IEEE Transactions on Industrial Informatics,True,['failure-management'],,,['anomaly-detection'],,,,,,['data-processing']
96,Anomaly Detection using Graph Neural Networks,"['A. Chaudhary', ' H. Mittal', ' A. Arora']",2019,"Conventional methods for anomaly detection include techniques based on clustering, proximity or classification. With the rapidly growing social networks, outliers or anomalies find ingenious ways to obscure themselves in the network and making the conventional techniques inefficient. In this paper, we utilize the ability of Deep Learning over topological characteristics of a social network to detect anomalies in email network and twitter network. We present a model, Graph Neural Network, which is applied on social connection graphs to detect anomalies. The combinations of various social network statistical measures are taken into account to study the graph structure and functioning of the anomalous nodes by employing deep neural networks on it. The hidden layer of the neural network plays an important role in finding the impact of statistical measure combination in anomaly detection.",https://ieeexplore.ieee.org/document/8862186,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1276}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 19}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 98}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 670}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 47}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 1531}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 22}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 29}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 30}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 1037}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 2}]",0.0,['multilayer-perceptron'],,['novel-use'],['failure-detection'],,"2019 International Conference on Machine Learning, Big Data, Cloud and Parallel Computing (COMITCon)",True,['failure-management'],,True,['anomaly-detection'],,,,,,
97,Self-Aware Workload Forecasting in Data Center Power Prediction,"['Y. Hsu', ' K. Matsuda', ' M. Matsuoka']",2018,"The number and scale of data centers are rapidly increasing, due to the growing demand for cloud computing services. Cloud computing infrastructure relies on a massive amount of information and communication technology (ICT) equipment, which consume an enormous amount of power. Power saving and energy optimization have therefore become essential goals for data centers. An enhanced data center energy management system (DEMS) provides a solution for data center power consumption based on its coordinative control of ICT equipment. An efficient power prediction model is essential for such a DEMS because it facilitates the proactive control of ICT equipment and reduces the total power consumption. In this paper, we propose a novel self-aware workload forecasting (SAWF) framework for total power consumption prediction in data centers. It includes three major components. First, there is a feature selection module, which evaluates the importance of variables from all ICT equipment in a data center and dynamically selects the most relevant variables for data input. Second, we propose an accurate and efficient neural network model to forecast future total power consumption. Third, we provide an online error monitoring and model updating module that continuously monitors prediction errors and updates the model when necessary.",https://ieeexplore.ieee.org/document/8411036,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1297}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 872}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 563}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 20}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 10}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1037}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 6}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 6}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 49}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 2}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 17}, {'database': 'ACM', 'search_string': ""'regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 2}]",5.0,,,,,,"2018 18th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID)",True,['resource-provisioning'],,True,,,,10.1109/CCGRID.2018.00047,Conference Paper,"['data center', 'machine learning', 'power consumption prediction', 'neural network', 'self-aware computing', 'workload forecasting']",
98,Fault and Performance Management in Multi-Cloud Based NFV Using Shallow and Deep Predictive Structures,"['L. Gupta', ' M. Samaka', ' R. Jain', ' A. Erbad', ' D. Bhamare', ' H. A. Chan']",2017,"Deployment of Network Function Virtualization (NFV) over multiple clouds accentuates its advantages like flexibility of virtualization, proximity to customers and lower total cost of operation. However, NFV over multiple clouds has not yet attained the level of performance to be a viable replacement for traditional networks. One of the reasons is the absence of a standard based Fault, Configuration, Accounting, Performance and Security (FCAPS) framework for the virtual network services. In NFV, faults and performance issues can have complex geneses within virtual resources as well as virtual networks and cannot be effectively handled by traditional rule-based systems. To tackle the above problem, we propose a fault detection and localization model based on a combination of shallow and deep learning structures. Relatively simpler detection has been effectively shown to be handled by shallow machine learning structures like Support Vector Machine (SVM). Deeper structure, i.e., the stacked autoencoder has been found to be useful for a more complex localization function where a large amount of information needs to be worked through to get to the root cause of the problem. We provide evaluation results using a dataset adapted from fault datasets available on Kaggle and another based on multivariate kernel density estimation and Markov sampling.",https://ieeexplore.ieee.org/document/8038530,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1316}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('fault localization' OR 'failure localization')"", 'index': 10}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault detection' OR 'failure detection')"", 'index': 364}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 827}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 530}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 406}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 312}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault localization' OR 'failure localization')"", 'index': 9}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 64}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 45}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 236}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 380}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1274}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('fault detection' OR 'failure detection')"", 'index': 113}, {'database': 'arxiv', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 251}, {'database': 'arxiv', 'search_string': ""('DL' OR 'deep learning') AND ('fault detection' OR 'failure detection')"", 'index': 11}, {'database': 'arxiv', 'search_string': ""('DL' OR 'deep learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 155}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 563}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 54}, {'database': 'arxiv', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 614}, {'database': 'arxiv', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault detection' OR 'failure detection')"", 'index': 37}, {'database': 'arxiv', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1518}, {'database': 'arxiv', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 14}, {'database': 'arxiv', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault detection' OR 'failure detection')"", 'index': 2}, {'database': 'arxiv', 'search_string': ""('support vector machine' OR 'SVM') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 18}]",12.0,"['random-forest', 'autoencoder', 'multilayer-perceptron', 'support-vector-machine', 'decision-tree']",['network-metrics'],['novel-use'],"['root-cause-analysis', 'failure-detection']",['network'],2017 26th International Conference on Computer Communication and Networks (ICCCN),True,['failure-management'],,True,"['anomaly-detection', 'root-cause-diagnosis', 'fault-localization']",,,10.1007/s40860-017-0053-y,,,
99,Anomaly Detection and Classification using Distributed Tracing and Deep Learning,"['S. Nedelkoski', ' J. Cardoso', ' O. Kao']",2019,"Artificial Intelligence for IT Operations (AIOps) combines big data and machine learning to replace a broad range of IT Operations tasks including availability, performance, and monitoring of services. By exploiting log, tracing, metric, and network data, AIOps enable detection of faults and issues of services. The focus of this work is on detecting anomalies based on distributed tracing records that contain detailed information for the availability and the response time of the services. In large-scale distributed systems, where a service is deployed on heterogeneous hardware and has multiple scenarios of normal operation, it becomes challenging to detect such anomalous cases. We address the problem by proposing unsupervised, response time anomaly detection based on deep learning data modeling techniques; unsupervised dynamic error threshold approach; tolerance module for false positive reduction; and descriptive classification of the anomalies. The evaluation shows that the approach achieves high accuracy and solid performance in both, experimental testbed and large-scale production cloud.",https://ieeexplore.ieee.org/document/8752866,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1379}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 306}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 73}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 63}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 176}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('IT operations')"", 'index': 2}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 307}, {'database': 'IEEE', 'search_string': ""'AIOps'"", 'index': 3}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 21}, {'database': 'IEEE', 'search_string': ""'classification' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 97}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 322}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 308}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1164}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 939}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 191}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 338}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 1080}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 290}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 106}, {'database': 'IEEE', 'search_string': ""'classification' AND ('IT operations')"", 'index': 1}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1460}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('IT operations')"", 'index': 5}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('IT operations')"", 'index': 3}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 971}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 532}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 134}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('IT operations')"", 'index': 1}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 530}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1663}]",2.0,"['autoencoder', 'rnn', 'cnn']","['kpis', 'traces', 'logs', 'network-metrics']",['novel-use'],"['root-cause-analysis', 'failure-detection']",,"2019 19th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID)",True,['failure-management'],,,,,,,,,
100,Predicting service metrics for cluster-based services using real-time analytics,"['R. Yanggratoke', ' J. Ahmed', ' J. Ardelius', ' C. Flinta', ' A. Johnsson', ' D. Gillblad', ' R. Stadler']",2015,"Predicting the performance of cloud services is intrinsically hard. In this work, we pursue an approach based upon statistical learning, whereby the behaviour of a system is learned from observations. Specifically, our testbed implementation collects device statistics from a server cluster and uses a regression method that accurately predicts, in real-time, client-side service metrics for a video streaming service running on the cluster. The method is service-agnostic in the sense that it takes as input operating-systems statistics instead of service-level metrics. We show that feature set reduction significantly improves prediction accuracy in our case, while simultaneously reducing model computation time. We also discuss design and implementation of a real-time analytics engine, which processes streams of device statistics and service metrics from testbed sensors and produces model predictions through online learning.",https://ieeexplore.ieee.org/document/7367349,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1534}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 405}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 272}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 281}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 374}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1112}]",15.0,"['random-forest', 'linear-regression']",['kpis'],,,,2015 11th International Conference on Network and Service Management (CNSM),True,['resource-provisioning'],,,,,,,,,
101,Predicting real-time service-level metrics from device statistics,"['R. Yanggratoke', ' J. Ahmed', ' J. Ardelius', ' C. Flinta', ' A. Johnsson', ' D. Gillblad', ' R. Stadler']",2015,"While real-time service assurance is critical for emerging telecom cloud services, understanding and predicting performance metrics for such services is hard. In this paper, we pursue an approach based upon statistical learning whereby the behavior of the target system is learned from observations. We use methods that learn from device statistics and predict metrics for services running on these devices. Specifically, we collect statistics from a Linux kernel of a server machine and predict client-side metrics for a video-streaming service (VLC). The fact that we collect thousands of kernel variables, while omitting service instrumentation, makes our approach service-independent and unique. While our current lab configuration is simple, our results, gained through extensive experimentation, prove the feasibility of accurately predicting client-side metrics, such as video frame rates and RTP packet rates, often within 10-15% error (NMAE), also under high computational load and across traces from different scenarios.",https://ieeexplore.ieee.org/document/7140318,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1536}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 297}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 755}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1115}]",4.0,"['dimensionality-reduction', 'random-forest', 'linear-regression']",['kpis'],,,,2015 IFIP/IEEE International Symposium on Integrated Network Management (IM),True,['resource-provisioning'],,,,,,,,,
102,Learning to Diagnose Stragglers in Distributed Computing,"['C. Li', ' H. Shen', ' T. Huang']",2016,"In cloud computing and high performance computing, a large job is typically divided into many small tasks for parallel execution in a distributed environment. Due to different reasons, some tasks (so-called `stragglers') are considerably slower than the others, delaying the completion of the job. We propose a new machine learning approach to automatically identify and diagnose the stragglers. To first identify stragglers, an unsupervised clustering method is employed to group the tasks based on their execution time. We then use a supervised rule learning algorithm to learn diagnosis rules inferring the stragglers with their resource assignment and performance counter data. Preliminary experiments from the trace of a Google's Borg cluster demonstrate that our method is able to generate simple and easy-to-read rules with decent performance in predicting stragglers.",https://ieeexplore.ieee.org/document/7836548,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1566}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 42}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 55}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1603}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 329}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1398}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 100}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 299}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 844}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 867}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 771}]",6.0,"['clustering', 'decision-tree']","['kpis', 'host-metrics']",['novel-use'],['failure-detection'],['straggler'],"2016 9th Workshop on Many-Task Computing on Clouds, Grids, and Supercomputers (MTAGS)",True,['failure-management'],,,['anomaly-detection'],,,,,,
103,Graph-based diagnosis in software-defined infrastructure,"['J. Wahba', ' H. Soliman', ' H. Bannazadeh', ' A. Leon-Garcia']",2016,"Performing system diagnosis is a critical task in modern datacenters. Investigating individual resource behavior may not be efficient in detecting abnormal behavior in large and complex datacenters. In this paper, we propose a scalable graph based diagnosis framework to detect system anomalies in Software-Defined Infrastructure running in SAVI testbed. We have leveraged Graph Mining and Machine Learning techniques in our approach in order to detect different kinds of anomalies. We have experimentally tested our framework on several use cases: Webserver-Database workload pattern, bandwidth throttling between a pair of VMs, denial-of-service (DoS) attack on a webserver and Spark Job failure. Our framework was able to detect the aforementioned anomalies accurately.",https://ieeexplore.ieee.org/document/7818425,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1609}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 757}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1138}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1219}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 31}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 71}]",188.0,,,,['failure-detection'],,2016 12th International Conference on Network and Service Management (CNSM),True,['failure-management'],,,['anomaly-detection'],,,,Conference Paper,"['Software-Defined Infrastructure', 'System Diagnosis', 'Anomaly Detection', 'Graph Mining', 'Machine Learning']",
104,Host Hypervisor Trace Mining for Virtual Machine Workload Characterization,"['H. Nemati', ' S. V. Azhari', ' M. R. Dagenais']",2019,"The efficient operation and resource management of multi-tenant data centers hosting thousands of services is a demanding task, that requires precise and detailed information regarding the behaviour of each and every virtual machine (VM). Often, coarse measures such as CPU, memory, disk and network usage by VMs are considered in grouping them onto the same physical server, as detailed measures would require access to the guest operating system (OS), which is not feasible in a multi-tenant setting. In this paper, we propose host-level hypervisor tracing as a non-intrusive means to extract useful features, that can provide for fine grain characterization of VM behaviour. In particular, we extract VM blocking periods as well as virtual interrupt injection rates to detect multiple levels of resource intensiveness. In addition, we consider the resource contention rate due to other VMs and the host, along with reasons for exit from non-root to root privileged mode, revealing useful information about the nature of the underlying VM workload. We also use tracing to get information about the rate of process and thread preemption in each VM, extracting process and thread contention as another feature set. We then employ various feature selection strategies and assess the quality of the resulting workload clustering. Notably, we adopt a two-stage feature selection approach in addition to a one shot clustering scheme. Moreover, we consider inter-cluster and intra-cluster similarity metrics, such as the silhouette score, to discover distinct groups of workloads as well as workload groups with significant overlap. This information can be used by 1) data center administrators to gain deeper visibility into the nature of various VMs running on their infrastructure, 2) performance engineers to assist root cause analysis of VM issues and 3) IaaS providers to help in resource management based on VM behavior.",https://ieeexplore.ieee.org/document/8790013,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1655}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 19}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 898}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 965}, {'database': 'IEEE', 'search_string': ""'clustering' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 42}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1655}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 70}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 33}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 101}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 120}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1562}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 1260}]",1.0,['clustering'],"['host-metrics', 'network-metrics']",['new-method'],['workload-prediction'],['vm'],2019 IEEE International Conference on Cloud Engineering (IC2E),True,['resource-provisioning'],,,,,,,,,
105,Real-Time Scheduling Policy Selection from Queue and Machine States,"[""L. Sant'ana"", ' D. Carastan-Santos', ' D. Cordeiro', ' R. De Camargo']",2019,"Task Scheduling in large-scale HPC platforms is normally accomplished with simple heuristics combined with a backfilling algorithm. Some strategies, such as the First-Come-First-Serve (FCFS) with backfilling, provide reasonable results in a variety of scenarios, including different HPC platforms and task set characteristics. But for each scenario, a different strategy might be the most appropriate for minimizing some metric, such as the average task waiting time or turnaround time. In this work, we present a real-time scheduling policy selection algorithm, which takes as input the running queue job characteristics and machine states. We evaluated the use of logistic regression and support-vector machines to perform the mapping from queue and machine state to selected scheduling policy. The machine learning algorithms are trained and evaluated using simulations configured using HPC platform traces. When selecting among 8 (eight) scheduling policies, we obtained an accuracy above 80%, when compared to the best selection. When simulating the online real-time selection of policies for a period of one year, we obtained a reduction in the mean queue waiting time of tasks of up to 40% over using FCFS and 10% over randomly selecting policies. Moreover, the method performed close the best possible selection of policies, with a maximum of 9% increase in the mean queue waiting time.",https://ieeexplore.ieee.org/document/8752858,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1729}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 558}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 18}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 893}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('cloud')"", 'index': 72}, {'database': 'IEEE', 'search_string': ""'regression' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 185}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 399}]",0.0,"['support-vector-machine', 'logistic-regression']",,"['comparison', 'novel-use']",['scheduling'],['hpc'],"2019 19th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID)",True,['resource-provisioning'],,True,,,,,,,
106,Towards Data-Driven Autonomics in Data Centers,"['A. Sîrbu', ' O. Babaoglu']",2015,"Continued reliance on human operators for managing data centers is a major impediment for them from ever reaching extreme dimensions. Large computer systems in general, and data centers in particular, will ultimately be managed using predictive computational and executable models obtained through data-science tools, and at that point, the intervention of humans will be limited to setting high-level goals and policies rather than performing low-level operations. Data-driven autonomics, where management and control are based on holistic predictive models that are built and updated using generated data, opens one possible path towards limiting the role of operators in data centers. In this paper, we present a data-science study of a public Google dataset collected in a 12K-node cluster with the goal of building and evaluating a predictive model for node failures. We use BigQuery, the big data SQL platform from the Google Cloud suite, to process massive amounts of data and generate a rich feature set characterizing machine state over time. We describe how an ensemble classifier can be built out of many Random Forest classifiers each trained on these features, to predict if machines will fail in a future 24-hour window. Our evaluation reveals that if we limit false positive rates to 5%, we can achieve true positive rates between 27% and 88% with precision varying between 50% and 72%. We discuss the practicality of including our predictive model as the central component of a data-driven autonomic manager and operating it on-line with live data streams (rather than off-line on data logs). All of the scripts used for BigQuery and classification analyses are publicly available from the authors' website.",https://ieeexplore.ieee.org/document/7312140,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1758}, {'database': 'IEEE', 'search_string': ""'classification' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 448}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 111}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 524}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 551}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 1493}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 77}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 87}, {'database': 'IEEE', 'search_string': ""'classification' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 60}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 172}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 460}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 1248}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 524}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 161}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 764}, {'database': 'arxiv', 'search_string': ""'classification' AND ('cloud')"", 'index': 649}, {'database': 'arxiv', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 655}, {'database': 'arxiv', 'search_string': ""'classification' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 3}, {'database': 'arxiv', 'search_string': ""'clustering' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 53}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 288}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 398}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 2}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 922}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 5}, {'database': 'arxiv', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 374}, {'database': 'arxiv', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1393}, {'database': 'arxiv', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 7}]",15.0,['random-forest'],['logs'],['novel-use'],['failure-prediction'],['bigquery'],2015 International Conference on Cloud and Autonomic Computing,True,['failure-management'],,,['system-failure-prediction'],,,,,,
107,Load Prediction for Data Centers Based on Database Service,"['R. Cao', ' Z. Yu', ' T. Marbach', ' J. Li', ' G. Wang', ' X. Liu']",2018,"In the era of cloud computing, the over-occupancy of data center resources (CPU, memory, disk) and subsequent machine failure have resulted in great loss to users and enterprises. So it makes sense to anticipate the server workload in advance. Previous research on server workloads has focused on trend analysis and time series fitting. We propose an approach to forecast the workloads of servers based on machine learning. And our data comes from a database-based data center that is real, large-scale, and enterprise-class. We use the servers' historical monitoring data for our models to predict future workloads and hence provide the ability to automatically warn overload and reallocate resources. We calculate the failure detection rate and false alarm rate of our overload detection models, as well as put forward an evaluation based on the overload processing cost. Experimental results show that machine learning methods especially Random Forest can better predict the server load than traditional time series analysis method. We use the forecast results to propose some scheduling strategies to prevent server overload, achieve intelligent operation and maintenance, and failure prediction. Compared with the traditional time series analysis method, our method uses less data and lower dimensions, and yields more accurate predictions.",https://ieeexplore.ieee.org/document/8377734,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1833}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 1060}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 1052}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 220}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault detection' OR 'failure detection')"", 'index': 630}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 101}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 171}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 1938}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 441}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1067}]",3.0,,,"['new-method', 'novel-use']",['workload-prediction'],['database'],2018 IEEE 42nd Annual Computer Software and Applications Conference (COMPSAC),True,['resource-provisioning'],,True,,,,,,,
108,Predicting SLA conformance for cluster-based services using distributed analytics,"['J. Ahmed', ' A. Johnsson', ' R. Yanggratoke', ' J. Ardelius', ' C. Flinta', ' R. Stadler']",2016,"Service assurance for the telecom cloud is a challenging task and is continuously being addressed by academics and industry. One promising approach is to utilize machine learning to predict service quality in order to take early mitigation actions. In previous work we have shown how to predict service-level metrics, such as frame rate for a video application on the client side, from operational data gathered at the server side. This gives the service provider early indications on whether the platform can support the current load demand. This paper extends previous work by addressing scalability issues for cluster-based services. Operational data being generated in large volumes, from several sources, and at high velocity puts strain on computational and communication resources. We propose and evaluate a distributed machine learning system based on the Winnow algorithm to tackle scalability issues, and then compare the new distributed solution with the previously proposed centralized solution. We show that network overhead and computational execution time is substantially reduced while maintaining high prediction accuracy making it possible to achieve real-time service quality predictions in large systems.",https://ieeexplore.ieee.org/document/7502913,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1836}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1888}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1183}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1528}]",1.0,['random-forest'],,,,['server'],NOMS 2016 - 2016 IEEE/IFIP Network Operations and Management Symposium,True,['resource-provisioning'],,,,,,,,,
109,Crying Wolf and Meaning It: Reducing False Alarms in Monitoring of Sporadic Operations through POD-Monitor,"['X. Xu', ' L. Zhu', ' M. Fu', ' D. Sun', ' A. B. Tran', ' P. Rimba', ' S. Dwarakanathan', ' L. Bass']",2015,"When monitoring complex applications in cloud systems, a difficult problem for operators is receiving false positive alarms. This becomes worse when the system is sporadically being changed and upgraded due to the emerging continuous deployment practice. Other legitimate but sporadic maintenance operations, such as log compression, garbage collection and data reconstruction in distributed systems can also trigger false alarms. Consequently, traditional baseline-based anomaly detection and monitoring is less effective. A normal but dangerous practice is to turn off normal monitoring during sporadic operations such as upgrade and maintenance. In this paper, we report on the use of the process context information of sporadic operations to suppress false positive alarms. We use the context information both directly and in machine learning. Our experimental evaluation shows that 1) using process context directly improves the alarm precision up to 0.226 (36.1% improvement), 2) using process-context trained machine learning models improves the precision rate up to 0.421 (84.7% improvement).",https://ieeexplore.ieee.org/document/7181485,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1838}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 834}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1244}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1022}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1542}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1532}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 62}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 19}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 85}]",10.0,['support-vector-machine'],"['host-metrics', 'logs']",['novel-use'],['failure-detection'],,2015 IEEE/ACM 1st International Workshop on Complex Faults and Failures in Large Software Systems (COUFLESS),True,['failure-management'],,True,,,,,Conference Paper,"['alarm', 'monitoring', 'operation']",['precision']
110,Layerwise Perturbation-Based Adversarial Training for Hard Drive Health Degree Prediction,"['J. Zhang', ' J. Wang', ' L. He', ' Z. Li', ' P. S. Yu']",2018,"With the development of cloud computing and big data, the reliability of data storage systems becomes increasingly important. Previous researchers have shown that machine learning algorithms based on SMART attributes are effective methods to predict hard drive failures. In this paper, we use SMART attributes to predict hard drive health degrees which are helpful for taking different fault tolerant actions in advance. Given the highly imbalanced SMART datasets, it is a nontrivial work to predict the health degree precisely. The proposed model would encounter overfitting and biased fitting problems if it is trained by the traditional methods. In order to resolve this problem, we propose two strategies to better utilize imbalanced data and improve performance. Firstly, we design a layerwise perturbation-based adversarial training method which can add perturbations to any layers of a neural network to improve the generalization of the network. Secondly, we extend the training method to the semi-supervised settings. Then, it is possible to utilize unlabeled data that have a potential of failure to further improve the performance of the model. Our extensive experiments on two real-world hard drive datasets demonstrate the superiority of the proposed schemes for both supervised and semi-supervised classification. The model trained by the proposed method can correctly predict the hard drive health status 5 and 15 days in advance.",https://ieeexplore.ieee.org/document/8595006,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1928}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 422}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 1023}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 794}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 393}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 479}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 668}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 925}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1195}, {'database': 'arxiv', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 220}, {'database': 'arxiv', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 60}, {'database': 'arxiv', 'search_string': ""'classification' AND ('cloud')"", 'index': 661}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 671}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 108}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 648}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 267}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 25}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 196}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 209}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 45}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 631}]",1.0,['multilayer-perceptron'],[],,['failure-prediction'],['hard-drive'],2018 IEEE International Conference on Data Mining (ICDM),True,['failure-management'],,,,,,,,,
111,Syslog processing for switch failure diagnosis and prediction in datacenter networks,"['Shenglin Zhang', ' Weibin Meng', ' Jiahao Bu', ' Sen Yang', ' Ying Liu', ' D. Pei', ' J. Xu', ' Yu Chen', ' Hui Dong', ' Xianping Qu', ' Lei Song']",2017,"Syslogs on switches are a rich source of information for both post-mortem diagnosis and proactive prediction of switch failures in a datacenter network. However, such information can be effectively extracted only through proper processing of syslogs, e.g., using suitable machine learning techniques. A common approach to syslog processing is to extract (i.e., build) templates from historical syslog messages and then match syslog messages to these templates. However, existing template extraction techniques either have low accuracies in learning the “correct” set of templates, or does not support incremental learning in the sense the entire set of templates has to be rebuilt (from processing all historical syslog messages again) when a new template is to be added, which is prohibitively expensive computationally if used for a large datacenter network. To address these two problems, we propose a frequent template tree (FT-tree) model in which frequent combinations of (syslog) words are identified and then used as message templates. FTtree empirically extracts message templates more accurately than existing approaches, and naturally supports incremental learning. To compare the performance of FT-tree and three other template learning techniques, we experimented them on two-years' worth of failure tickets and syslogs collected from switches deployed across 10+ datacenters of a tier-1 cloud service provider. The experiments demonstrated that FT-tree improved the estimation/prediction accuracy (as measured by F1) by 155% to 188%, and the computational efficiency by 117 to 730 times.",https://ieeexplore.ieee.org/document/7969130,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1929}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 889}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 32}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 36}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 1403}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 70}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1607}]",14.0,['markov-model'],['logs'],"['new-method', 'comparison']","['failure-prediction', 'failure-detection']","['network', 'switch']",2017 IEEE/ACM 25th International Symposium on Quality of Service (IWQoS),True,['failure-management'],True,True,"['hardware-failure-prediction', 'anomaly-detection']",True,40.0,,,,
112,Deep Convolutional Neural Networks for Log Event Classification on Distributed Cluster Systems,"['R. Ren', ' J. Cheng', ' Y. Yin', ' J. Zhan', ' L. Wang', ' J. Li', ' C. Luo']",2018,"With the widespread development of cloud computing, cluster systems are becoming increasingly complex, system logs is an universal and effective approach for automatic system management and troubleshooting. Log event classification as an effective preprocessing method for log analysis, which is helpful for system administrators to locate or predict components' have errors or failures. In this paper, we design and implement an automatic log classification system based on deep CNN (Convolutional Neural Network) models, and take advantage of the feature engineering and learning algorithm to improve classification performance. First, in the feature engineering step, to address the problem of that the original unstructured event logs are unsuitable for numerical calculation in deep CNN models, we propose a novel and effective log preprocessing method, which include building categories dictionary libraries, filtering abundant information, generating numerical semantic feature vectors by calculating and combining the semantic similarity values for filtered log events. Additionally, in the learning step, we measure a series of deep CNN algorithms with varied hyper-parameter combinations by using standard evaluation metrics, and the results of our study reveal the advantages and potential capabilities of the proposed deep CNN models for log classification tasks on cluster systems. The optimal classification precision of our approach is 98.14%, which surpasses the popular traditional machine learning methods, and it can also be applied to other large-scale system logs with good accuracy. Just like the experiment results, different choices of learning algorithm do result in performance numbers varying, and subsequently careful feature engineering enables promoting performances, thus both of approaches contribute to best learning model finding.",https://ieeexplore.ieee.org/document/8622611,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1963}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 8}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 30}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1537}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 933}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 1269}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 803}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 657}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 399}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 423}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 11}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 251}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1854}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 737}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 4}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 10}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 3}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 560}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 535}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1212}]",2.0,['cnn'],"['events', 'logs']","['comparison', 'novel-use']",['root-cause-analysis'],['cluster'],2018 IEEE International Conference on Big Data (Big Data),True,['failure-management'],,,['root-cause-diagnosis'],,,,,,
113,A pragmatic approach to predict hardware failures in storage systems using MPP database and big data technologies,"['R. Kumar', ' S. Vijayakumar', ' S. A. Ahamed']",2014,"A storage system in a data center consists of various components such as Disk Array Enclosure (DAE), disks, processors, servers, hosts running different applications, and so on. Hard disk and server failures are not frequent but are often very costly. Such failures can have a very adverse effect on the business of an organization. The ability to accurately predict an impending disk or server failure can add an essential functionality for designing a reliable, fault tolerant and continuously available storage system. This paper explains a novel approach to predict hardware failures using spectrum-kernel Parallel Support Vector Machine (Parallel SVM) method by analyzing the system events logged in the system log files. These log files not only records the events processed by the system but it also holds the messages as the system state changes. A single message in the system log file is insufficient for any prediction and such prediction is bound to be less accurate. The approach introduced in the paper uses a sequence or pattern of messages from the system log file using a Sliding Window of messages with window size of 5 message sequence to predict the likelihood of a failure. These Sliding Windows of message sequences acts as inputs to the Parallel SVM. The Parallel SVM further tags the messages to a failure or non-failure system. Data Mining techniques are used in extracting useful information from the raw dataset. A solutioning model is developed using the structured dataset and Machine Learning algorithms. This environment when implemented using actual system logs from Linux-based storage system have shown to predict a hardware failure with accuracy of 90-92 percent.",https://ieeexplore.ieee.org/document/6779422,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1985}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 359}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 106}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 352}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 155}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 132}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 438}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 1558}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('remediation' OR 'recovery')"", 'index': 63}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 733}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 120}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1062}, {'database': 'IEEE', 'search_string': ""'classification' AND ('remediation' OR 'recovery')"", 'index': 780}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 42}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 172}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1305}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 975}]",6.0,['support-vector-machine'],['logs'],['novel-use'],['failure-prediction'],"['server', 'hard-drive']",2014 IEEE International Advance Computing Conference (IACC),True,['failure-management'],,True,['hardware-failure-prediction'],,,,,,
114,CARVE: A Cognitive Agent for Resource Value Estimation,"['J. Wildstrom', ' P. Stone', ' E. Witchel']",2008,"Recently, industry has begun investigating and moving towards utility computing, where computational resources (processing, memory and I/O) are availably on demand at a market cost. On-demand access to computational resources enables fine-grained resource allocation for web-based applications, e.g., the possibility of provisioning for a minimum workload while allowing the rental of additional resources for unexpected workload changes. However, renting additional resources relies on the ability to quickly and accurately estimate the value of the resource. This paper introduces CARVE: a cognitive agent for resource value estimation. CARVE is a machine-learning based approach that learns to predict the change in system value of having more or less system resources. Using only low-level statistics and with no custom instrumentation of the operating system or middleware, CARVE is able to make informed decisions about the return on investment of physical memory when implemented and evaluated on a partitioned system running a multi-partition, multi-process distributed benchmark. We show that CARVE is competitive with static choices of computing resources over a variety of test workloads and also has the ability to outperform all static configurations.",https://ieeexplore.ieee.org/document/4550839,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 1995}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 441}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 691}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1334}]",12.0,,,,,,2008 International Conference on Autonomic Computing,True,['resource-provisioning'],,,,,,,,,
115,Proactive Failure Management by Integrated Unsupervised and Semi-Supervised Learning for Dependable Cloud Systems,"['Q. Guan', ' Z. Zhang', ' S. Fu']",2011,"Cloud computing systems continue to grow in their scale and complexity. They are changing dynamically as well due to the addition and removal of system components, changing execution environments, frequent updates and upgrades, online repairs and more. In such large-scale complex and dynamic systems, failures are common. In this paper, we present a failure prediction mechanism exploiting both unsupervised and semi-supervised learning techniques for building dependable cloud computing systems. The unsupervised failure detection method uses an ensemble of Bayesian models. It characterizes normal execution states of the system and detects anomalous behaviors. After the anomalies are verified by system administrators, labeled data are available. Then, we apply supervised learning based on decision tree classier to predict future failure occurrences in the cloud. Experimental results in an institute-wide cloud computing system show that our proposed method can forecast failure dynamics with high accuracy.",https://ieeexplore.ieee.org/document/6045942,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 8}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 232}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 7}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 169}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 6}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 759}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 174}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 852}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault detection' OR 'failure detection')"", 'index': 1490}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault detection' OR 'failure detection')"", 'index': 251}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 54}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('remediation' OR 'recovery')"", 'index': 445}, {'database': 'IEEE', 'search_string': ""'classification' AND ('remediation' OR 'recovery')"", 'index': 764}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('remediation' OR 'recovery')"", 'index': 5}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 280}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 4}]",36.0,"['dimensionality-reduction', 'naive-bayes', 'decision-tree']","['host-metrics', 'network-metrics']",['novel-use'],"['failure-prediction', 'failure-detection']",,"2011 Sixth International Conference on Availability, Reliability and Security",True,['failure-management'],,,"['anomaly-detection', 'system-failure-prediction']",,,,,,
116,Research on Reinforcement Learning-Based Dynamic Power Management for Edge Data Center,"['Q. Guo', ' R. Huo', ' H. Meng', ' E. Xinhua', ' J. Liu', ' T. Huang']",2018,"Mobile Edge Computing (MEC) is a supplement to traditional cloud computing. Its characteristics are low latency and high reliability, and it will be widely used in the future. However, their dense deployment pattern raises a big concern on the system-wide energy consumption. Dynamic power management (DPM) method is an important method to solve energy consumption problems, it saves energy by shutting down servers in the EDC that are idle or have low utilization. In this paper, a DPM method based on reinforcement learning was proposed, it achieves the trade-off between EDC service performance and energy consumption by learning the global optimal dynamic timeout threshold power management strategy by trial and error. Experiments have shown that the proposed method saves no less than 6.35% energy consumption compared to the expert-based method.",https://ieeexplore.ieee.org/document/8663880,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 13}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 1}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1023}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 30}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 4}]",0.0,['reinforcement-learning'],,,['power-management'],,2018 IEEE 9th International Conference on Software Engineering and Service Science (ICSESS),True,['resource-provisioning'],,,,,,,,,
117,Self-Learning Cloud Controllers: Fuzzy Q-Learning for Knowledge Evolution,"['P. Jamshidi', ' A. M. Sharifloo', ' C. Pahl', ' A. Metzger', ' G. Estrada']",2015,"Auto-scaling features enable cloud applications to maintain enough resources to satisfy demand spikes, reduce costs and keep performance in check. Most auto-scaling strategies rely on a predefined set of rules to scale up/down the required resources depending on the application usage. Those rules are however difficult to devise and generalize, and users are often left alone tuning auto-scale parameters of essentially blackbox applications. In this paper, we propose a novel fuzzy reinforcement learning controller, FQL4KE, which automatically scales up or down resources to meet performance requirements. The Q-Learning technique, a model-free reinforcement learning strategy, frees users of most tuning parameters. FQL4KE has been successfully applied and we therefore think that a fuzzy controller with Q-Learning is indeed a promising combination for auto-scaling resources.",https://ieeexplore.ieee.org/document/7312157,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 21}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1294}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 648}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 14}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 716}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 1104}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud')"", 'index': 259}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 60}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 96}]",23.0,,,,['resource-consolidation'],,2015 International Conference on Cloud and Autonomic Computing,True,['resource-provisioning'],,True,,,,,,,
118,Coordinated Self-Configuration of Virtual Machines and Appliances Using a Model-Free Learning Approach,"['X. Bu', ' J. Rao', ' C. Xu']",2013,"Cloud computing has a key requirement for resource configuration in a real-time manner. In such virtualized environments, both virtual machines (VMs) and hosted applications need to be configured on-the-fly to adapt to system dynamics. The interplay between the layers of VMs and applications further complicates the problem of cloud configuration. Independent tuning of each aspect may not lead to optimal system wide performance. In this paper, we propose a framework, namely CoTuner, for coordinated configuration of VMs and resident applications. At the heart of the framework is a model-free hybrid reinforcement learning (RL) approach, which combines the advantages of Simplex method and RL method and is further enhanced by the use of system knowledge guided exploration policies. Experimental results on Xen-based virtualized environments with TPC-W and TPC-C benchmarks demonstrate that CoTuner is able to drive a virtual server cluster into an optimal or near-optimal configuration state on the fly, in response to the change of workload. It improves the systems throughput by more than 30 percent over independent tuning strategies. In comparison with the coordinated tuning strategies based on basic RL or Simplex algorithm, the hybrid RL algorithm gains 25 to 40 percent throughput improvement.",https://ieeexplore.ieee.org/document/6216363,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 28}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 582}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 506}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 70}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1169}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 1062}]",24.0,,,,,,IEEE Transactions on Parallel and Distributed Systems,True,['resource-provisioning'],,,,,,,,,
119,DERP: A Deep Reinforcement Learning Cloud System for Elastic Resource Provisioning,"['C. Bitsakos', ' I. Konstantinou', ' N. Koziris']",2018,"Modern large scale computer clusters benefit significantly from elasticity. Elasticity allows a cluster to dynamically allocate computer resources, based on the user's fluctuating workload demands. Many cloud providers use threshold-based approaches, which have been proven to be difficult to configure and optimise, while others use reinforcement learning and decision-tree approaches, which struggle when having to handle large multidimensional cluster states. In this work we use Deep Reinforcement learning techniques to achieve automatic elasticity. We use three different approaches of a Deep Reinforcement learning agent, called DERP (Deep Elastic Resource Provisioning), that takes as input the current multi-dimensional state of a cluster and manages to train and converge to the optimal elasticity behaviour after a finite amount of training steps. The system automatically decides and proceeds on requesting/releasing VM resources from the provider and orchestrating them inside a NoSQL cluster according to user-defined policies/rewards. We compare our agent to state-of-the-art, Reinforcement learning and decision-tree based, approaches in demanding simulation environments and show that it gains rewards up to 1.6 times better on its lifetime. We then test our approach in a real life cluster environment and show that the system resizes clusters in real-time and adapts its performance through a variety of demanding optimisation strategies, input and training loads.",https://ieeexplore.ieee.org/document/8590989,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 35}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 631}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 265}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 650}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 37}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 457}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 234}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 507}]",1.0,"['multilayer-perceptron', 'reinforcement-learning']",['metrics'],,,['vm'],2018 IEEE International Conference on Cloud Computing Technology and Science (CloudCom),True,['resource-provisioning'],,,,,,,,,
120,DRL-cloud: Deep reinforcement learning-based resource provisioning and task scheduling for cloud service providers,"['M. Cheng', ' J. Li', ' S. Nazarian']",2018,"Cloud computing has become an attractive computing paradigm in both academia and industry. Through virtualization technology, Cloud Service Providers (CSPs) that own data centers can structure physical servers into Virtual Machines (VMs) to provide services, resources, and infrastructures to users. Profit-driven CSPs charge users for service access and VM rental, and reduce power consumption and electric bills so as to increase profit margin. The key challenge faced by CSPs is data center energy cost minimization. Prior works proposed various algorithms to reduce energy cost through Resource Provisioning (RP) and/or Task Scheduling (TS). However, they have scalability issues or do not consider TS with task dependencies, which is a crucial factor that ensures correct parallel execution of tasks. This paper presents DRL-Cloud, a novel Deep Reinforcement Learning (DRL)-based RP and TS system, to minimize energy cost for large-scale CSPs with very large number of servers that receive enormous numbers of user requests per day. A deep Q-learning-based two-stage RP-TS processor is designed to automatically generate the best long-term decisions by learning from the changing environment such as user request patterns and realistic electric price. With training techniques such as target network, experience replay, and exploration and exploitation, the proposed DRL-Cloud achieves remarkably high energy cost efficiency, low reject rate as well as low runtime with fast convergence. Compared with one of the state-of-the-art energy efficient algorithms, the proposed DRL-Cloud achieves up to 320% energy cost efficiency improvement while maintaining lower reject rate on average. For an example CSP setup with 5,000 servers and 200,000 tasks, compared to a fast round-robin baseline, the proposed DRL-Cloud achieves up to 144% runtime reduction.",https://ieeexplore.ieee.org/document/8297294,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 51}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 14}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 833}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 26}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 220}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 140}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 41}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 79}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 65}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 82}]",7.0,,,,,,2018 23rd Asia and South Pacific Design Automation Conference (ASP-DAC),True,['resource-provisioning'],,True,,,,,Conference Paper,"['cloud computing', 'cloud resource management', 'resource provisioning', 'deep Q-learning', 'task scheduling', 'deep reinforcement learning']",
121,Horizontal and Vertical Scaling of Container-Based Applications Using Reinforcement Learning,"['F. Rossi', ' M. Nardelli', ' V. Cardellini']",2019,"Software containers are changing the way distributed applications are executed and managed on cloud computing resources. Interestingly, containers offer the possibility of handling workload fluctuations by exploiting both horizontal and vertical elasticity ""on the fly"". However, most of the existing control policies consider horizontal and vertical scaling as two disjointed control knobs. In this paper, we propose Reinforcement Learning (RL) solutions for controlling the horizontal and vertical elasticity of container-based applications with the goal to increase the flexibility to cope with varying workloads. Although RL represents an interesting approach, it may suffer from a possible long learning phase, especially when nothing about the system is known a-priori. To speed up the learning process and identify better adaptation policies, we propose RL solutions that exploit different degrees of knowledge about the system dynamics (i.e., Q-learning, Dyna-Q, and Model-based). We integrate the proposed policies in Elastic Docker Swarm, our extension that introduces self-adaptation capabilities in the container orchestration tool Docker Swarm. We demonstrate the effectiveness and flexibility of model-based RL policies through simulations and prototype-based experiments.",https://ieeexplore.ieee.org/document/8814555,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 63}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 959}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 48}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 615}]",7.0,['reinforcement-learning'],,,['resource-consolidation'],"['container', 'docker']",2019 IEEE 12th International Conference on Cloud Computing (CLOUD),True,['resource-provisioning'],,,,,,,,,
122,A Hierarchical Framework of Cloud Resource Allocation and Power Management Using Deep Reinforcement Learning,"['N. Liu', ' Z. Li', ' J. Xu', ' Z. Xu', ' S. Lin', ' Q. Qiu', ' J. Tang', ' Y. Wang']",2017,"Automatic decision-making approaches, such as reinforcement learning (RL), have been applied to (partially) solve the resource allocation problem adaptively in the cloud computing system. However, a complete cloud resource allocation framework exhibits high dimensions in state and action spaces, which prohibit the usefulness of traditional RL techniques. In addition, high power consumption has become one of the critical concerns in design and control of cloud computing systems, which degrades system reliability and increases cooling cost. An effective dynamic power management (DPM) policy should minimize power consumption while maintaining performance degradation within an acceptable level. Thus, a joint virtual machine (VM) resource allocation and power management framework is critical to the overall cloud computing system. Moreover, novel solution framework is necessary to address the even higher dimensions in state and action spaces. In this paper, we propose a novel hierarchical framework for solving the overall resource allocation and power management problem in cloud computing systems. The proposed hierarchical framework comprises a global tier for VM resource allocation to the servers and a local tier for distributed power management of local servers. The emerging deep reinforcement learning (DRL) technique, which can deal with complicated control problems with large state space, is adopted to solve the global tier problem. Furthermore, an autoencoder and a novel weight sharing structure are adopted to handle the high-dimensional state space and accelerate the convergence speed. On the other hand, the local tier of distributed server power managements comprises an LSTM based workload predictor and a model-free RL based power manager, operating in a distributed manner. Experiment results using actual Google cluster traces show that our proposed hierarchical framework significantly saves the power consumption and energy usage than the baseline while achieving no severe latency degradation. Meanwhile, the proposed framework can achieve the best trade-off between latency and power/energy consumption in a server cluster.",https://ieeexplore.ieee.org/document/7979983,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 69}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 113}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 831}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 702}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 77}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 820}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 194}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 37}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 156}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 3}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 82}]",26.0,"['multilayer-perceptron', 'reinforcement-learning']",,,"['resource-consolidation', 'power-management']",,2017 IEEE 37th International Conference on Distributed Computing Systems (ICDCS),True,['resource-provisioning'],,True,,,,,,,
123,Learn-as-You-Go with Megh: Efficient Live Migration of Virtual Machines,"['D. Basu', ' X. Wang', ' Y. Hong', ' H. Chen', ' S. Bressan']",2017,"We propose a reinforcement learning algorithm, Megh, for live migration of virtual machines that simultaneously reduces the cost of energy consumption and enhances the performance. Megh learns the uncertain dynamics of workloads as-it-goes. Megh uses a dimensionality reduction scheme to projectthe combinatorially explosive state-action space to a polynomial dimensional space. These schemes enable Megh to be scalable and to work in real-time. We experimentally validate that Megh is more cost-effective and time-efficient than the MadVM and MMT algorithms.",https://ieeexplore.ieee.org/document/7980251,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 70}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 262}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 1083}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 287}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1332}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 150}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 351}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 938}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1932}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 1509}]",4.0,['reinforcement-learning'],['metrics'],,,['vm'],2017 IEEE 37th International Conference on Distributed Computing Systems (ICDCS),True,['resource-provisioning'],,True,,,,,,,
124,Unsupervised Neural Predictor to Auto-administrate the Cloud Infrastructure,"['H. Chihi', ' W. Chainbi', ' K. Ghedira']",2012,"Due to all the pollutants generated by it and the steady increases in its rates, energy consumption has become a key issue. Cloud computing is an emerging model for distributed utility computing and is being considered as an attractive opportunity for saving energy through central management of computational resources. Obviously, a substantial reduction in energy consumption can be made by powering down servers when they are not in use. This work presents a resources provisioning approach based on an unsupervised predictor model in the form of an unsupervised, recurrent neural network based on a self-organizing map. Unsupervised learning in computers has for long been considered as the desired ambition of computer problems. Unlike conventional prediction-learning methods which assign credit by means of the difference between predicted and actual outcomes, the proposed study assigns credit by means of the difference between temporally successive predictions. We have shown that the proposed approach gives promising results.",https://ieeexplore.ieee.org/document/6424970,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 84}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 142}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 132}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 90}]",2.0,,,,,,2012 IEEE Fifth International Conference on Utility and Cloud Computing,True,['resource-provisioning'],,,,,,,,,
125,Fuzzy Reinforcement Learning based Microservice Allocation in Cloud Computing Environments,"['C. T. Joseph', ' J. P. Martin', ' K. Chandrasekaran', ' A. Kandasamy']",2019,"Nowadays the Cloud Computing paradigm has become the defacto platform for deploying and managing user applications. Monolithic Cloud applications pose several challenges in terms of scalability and flexibility. Hence, Cloud applications are designed as microservices. Application scheduling and energy efficiency are key concerns in Cloud computing research. Allocating the microservice containers to the hosts in the datacenter is an NP-hard problem. There is a need for efficient allocation strategies to determine the placement of the microservice containers in Cloud datacenters to minimize Service Level Agreement violations and energy consumption. In this paper, we design a Reinforcement Learning-based Microservice Allocation (RL-MA) approach. The approach is implemented in the ContainerCloudSim simulator. The evaluation is conducted using the real-world Google cluster trace. Results indicate that the proposed method reduces both the SLA violation and energy consumption when compared to the existing policies.",https://ieeexplore.ieee.org/document/8929586,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 88}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1715}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 237}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1631}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 30}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 319}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 507}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 134}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 725}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 114}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 163}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 212}]",0.0,['reinforcement-learning'],,,['resource-consolidation'],['microservice'],TENCON 2019 - 2019 IEEE Region 10 Conference (TENCON),True,['resource-provisioning'],,True,,,,,,,
126,A Comparison of Reinforcement Learning Techniques for Fuzzy Cloud Auto-Scaling,"['H. Arabnejad', ' C. Pahl', ' P. Jamshidi', ' G. Estrada']",2017,"A goal of cloud service management is to design self-adaptable auto-scaler to react to workload fluctuations and changing the resources assigned. The key problem is how and when to add/remove resources in order to meet agreed service-level agreements. Reducing application cost and guaranteeing service-level agreements (SLAs) are two critical factors of dynamic controller design. In this paper, we compare two dynamic learning strategies based on a fuzzy logic system, which learns and modifies fuzzy scaling rules at runtime. A self-adaptive fuzzy logic controller is combined with two reinforcement learning (RL) approaches: (i) Fuzzy SARSA learning FSL and (ii) Fuzzy Q-learning FQL. As an off-policy approach, Q-learning learns independent of the policy currently followed, whereas SARSA as an on-policy always incorporates the actual agent's behavior and leads to faster learning. Both approaches are implemented and compared in their advantages and disadvantages, here in the OpenStack cloud platform. We demonstrate that both auto-scaling approaches can handle various load traffic situations, sudden and periodic, and delivering resources on demand while reducing operating costs and preventing SLA violations. The experimental results demonstrate that FSL and FQL have acceptable performance in terms of adjusted number of virtual machine targeted to optimize SLA compliance and response time.",https://ieeexplore.ieee.org/document/7973689,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 90}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 879}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 88}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 40}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 266}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 221}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 21}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 74}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 47}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 161}, {'database': 'arxiv', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 523}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 61}]",2.0,,,,,,"2017 17th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID)",True,['resource-provisioning'],,True,,,,10.1109/CCGRID.2017.15,Conference Paper,"['OpenStack', 'Q-Learning', 'Cloud Computing', 'Controller', 'SARSA', 'Fuzzy Logic', 'Orchestration']",
127,Coordinated Power and Performance-Efficient Virtual Machines Scheduling in the Cloud,"['S. Wang', ' X. Zhou', ' M. Shang', ' X. Shi']",2018,"Cloud computing with live migration technique is considered as one of the most promising ways to cope with power consumption and performance management of a data center. Most prior works on performance and power management of the whole server farm are achieved in a separate way. To address this issue, in this paper we propose an efficient method for the whole server farm, which aims to dynamically consolidate virtual machines in a coordinated way that optimizes the energy and performance trade-off. Firstly, we focus on the virtual machine (VM) selection step. Then we consider the VM selection task as a Dynamic Decision-Making (DDM) problem and construct a coordinated cost function with power and performance. In this study, the Q-Learning strategy of Reinforcement Learning (RL) is adopted to solve this DDM problem. The proposed algorithm is simulated in CloudSim toolkit using real-world workload traces. Finally, experimental results indicate that our approach outperforms other algorithms in terms of energy consumption, the number of VM migrations, average SLA violation and the number of host shutdowns.",https://ieeexplore.ieee.org/document/8768909,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 110}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 176}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1534}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 19}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 870}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 135}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1291}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 117}]",0.0,['reinforcement-learning'],,,"['scheduling', 'power-management']",['vm'],"2018 10th International Conference on Communications, Circuits and Systems (ICCCAS)",True,['resource-provisioning'],,True,,,,,,,
128,Automated Enforcement of SLA for Cloud Services,"['S. Vakilinia', ' C. Truchan', ' J. Kempf', ' H. Elbiaze']",2018,"Orchestration and management of cloud computing entities necessitate measuring and analysis of real-time monitored performance metrics. However, decision making in current management platforms are addressed separately in different cloud stack layers. These isolated active management decisions may degrade the total performance of the cloud system. Since, cloud computing platforms lack an integrated analytics and management capability, in this paper, we propose an integrated platform to detect and predict situations where corrective actions are required. First, a Dynamic Bayesian Network (DBN) is trained and updated by collected data to calculate the causal dependencies among various entities in different cloud service layers. The correlation values are then fed into a Long Short-Term Memory (LSTM) neural network to predict the future states. States that violate the Service Level Agreement(SLA) of cloud services are learned with training data, and if the forecasted states threaten the SLA of cloud services, associated events are generated to trigger management actions. Next, management actions are assigned a different set of events using a reinforcement learning approach. A set of experiments based on collected data from a real cloud service environment is conducted to validate the proposed approach. Experimental results indicate that the proposed method outperforms the current management solutions and improves web request response time by up to 7% and decreases SLA violation by 79% in the context of web application auto-scaling.",https://ieeexplore.ieee.org/document/8457782,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 155}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 19}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 11}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 478}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 673}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 112}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 388}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 141}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 39}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 40}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 569}]",0.0,"['rnn', 'bayesian-network', 'reinforcement-learning']",['sla'],['novel-use'],['failure-prediction'],,2018 IEEE 11th International Conference on Cloud Computing (CLOUD),True,['failure-management'],,True,,,,,,,
129,Energy-Efficient Virtual Machines Consolidation in Cloud Data Centers Using Reinforcement Learning,"['F. Farahnakian', ' P. Liljeberg', ' J. Plosila']",2014,"Dynamic consolidation techniques optimize resource utilization and reduce energy consumption in Cloud data centers. They should consider the variability of the workload to decide when idle or underutilized hosts switch to sleep mode in order to minimize energy consumption. In this paper, we propose a Reinforcement Learning-based Dynamic Consolidation method (RL-DC) to minimize the number of active hosts according to the current resources requirement. The RL-DC utilizes an agent to learn the optimal policy for determining the host power mode by using a popular reinforcement learning method. The agent learns from past knowledge to decide when a host should be switched to the sleep or active mode and improves itself as the workload changes. Therefore, RL-DC does not require any prior information about workload and it dynamically adapts to the environment to achieve online energy and performance management. Experimental results on the real workload traces from more than a thousand PlanetLab virtual machines show that RL-DC minimizes energy consumption and maintains required performance levels.",https://ieeexplore.ieee.org/document/6787321,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 190}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 109}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 879}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 136}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1663}]",59.0,,,,['resource-consolidation'],,"2014 22nd Euromicro International Conference on Parallel, Distributed, and Network-Based Processing",True,['resource-provisioning'],,,,,,,,,
130,Towards QoS-Aware Cloud Live Transcoding: A Deep Reinforcement Learning Approach,"['Z. Pang', ' L. Sun', ' T. Huang', ' Z. Wang', ' S. Yang']",2019,"Video transcoding is widely adopted in live streaming services to bridge the format and resolution gap between content producers and consumers (i.e., broadcasters and viewers). Meanwhile, the cloud has been recognized as one of the most reliable and cost-effective ways for video transcoding. However, due to the dynamic and uncertainty of the transcoding workloads in live streaming, it is very challenging for cloud service providers to provision computing resources and schedule transcoding tasks while guaranteeing the Service Level Agreement (SLA). To this end, we propose a joint resource provisioning and task scheduling approach for transcoding live streams in the cloud. We adopt Deep Reinforcement Learning (DRL) to train a neural network model for resource provisioning under dynamic workloads. Moreover, we design a QoS-aware task scheduling algorithm that maps transcoding tasks to Virtual Machines (VMs) by considering the real-time QoS requirement. We evaluate our approach with trace-driven experiments and the results demonstrate that our approach outperforms heuristic baselines by up to 89% improvements on average QoS with 4% extra resource overhead at most.",https://ieeexplore.ieee.org/document/8785022,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 193}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 845}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 178}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1637}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 896}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 108}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 908}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 863}]",0.0,"['multilayer-perceptron', 'reinforcement-learning']",,,['resource-consolidation'],['vm'],2019 IEEE International Conference on Multimedia and Expo (ICME),True,['resource-provisioning'],,True,,,,,,,
131,Adaptive Dispatching of Tasks in the Cloud,"['L. Wang', ' E. Gelenbe']",2018,"The increasingly wide application of Cloud Computing enables the consolidation of tens of thousands of applications in shared infrastructures. Thus, meeting the QoS requirements of so many diverse applications in such shared resource environments has become a real challenge, especially since the characteristics and workload of applications differ widely and may change over time. This paper presents an experimental system that can exploit a variety of online QoS aware adaptive task allocation schemes, and three such schemes are designed and compared. These are a measurement driven algorithm that uses reinforcement learning, secondly a “sensible” allocation algorithm that assigns tasks to sub-systems that are observed to provide a lower response time, and then an algorithm that splits the task arrival stream into sub-streams at rates computed from the hosts' processing capabilities. All of these schemes are compared via measurements among themselves and with a simple round-robin scheduler, on two experimental test-beds with homogenous and heterogenous hosts having different processing capacities.",https://ieeexplore.ieee.org/document/7229325,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 197}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 710}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 613}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 455}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 252}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1029}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 482}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 20}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 150}]",82.0,,,['comparison'],['scheduling'],,IEEE Transactions on Cloud Computing,True,['resource-provisioning'],,True,,,,,,,
132,Elastic management of cloud applications using adaptive reinforcement learning,"['K. Lolos', ' I. Konstantinou', ' V. Kantere', ' N. Koziris']",2017,"Modern large-scale computing deployments consist of complex applications running over machine clusters. An important issue in these is the offering of elasticity, i.e., the dynamic allocation of resources to applications to meet fluctuating workload demands. Threshold based approaches are typically employed, yet they are difficult to calibrate and optimize. Approaches based on reinforcement learning (RL) have been proposed, but they require a large number of states in order to model complex application behavior. Methods that adaptively partition the state space have been proposed, but their partitioning criteria and strategies are sub-optimal. In this work we present MDP_DT, a novel full-model based reinforcement learning algorithm for elastic resource management that employs adaptive state space partitioning. We propose two novel statistical criteria and three strategies and we experimentally prove that they correctly decide both where and when to partition, outperforming existing approaches. We experimentally evaluate MDP_DT in a real large scale cluster over variable not-encountered workloads and we show that it takes more informed decisions compared to static, model-free and threshold approaches, while requiring a minimal amount of training data. We experimentally show that this adaptation enabled MDP_DT to optimize the achieved profit while being 40% cheaper than calibrated RL and threshold approaches.",https://ieeexplore.ieee.org/document/8257928,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 222}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 217}]",1.0,['reinforcement-learning'],,,['resource-consolidation'],,2017 IEEE International Conference on Big Data (Big Data),True,['resource-provisioning'],,,,,,,,,
133,DeepEE: Joint Optimization of Job Scheduling and Cooling Control for Data Center Energy Efficiency Using Deep Reinforcement Learning,"['Y. Ran', ' H. Hu', ' X. Zhou', ' Y. Wen']",2019,"The past decade witnessed the tremendous growth of power consumption in data centers due to the rapid development of cloud computing, big data analytics, and machine learning, etc. The prior approaches that optimize the power consumption of the information technology (IT) system and/or the cooling system always fail to capture the system dynamics or suffer from the complexity of system states and action spaces. In this paper, we propose a Deep Reinforcement Learning (DRL) based optimization framework, named DeepEE, to improve the energy efficiency for data centers by considering the IT and cooling systems concurrently. In DeepEE, we first propose a PArameterized action space based Deep Q-Network (PADQN) algorithm to solve the hybrid action space problem and jointly optimize the job scheduling for the IT system and the airflow rate adjustment for the cooling system. Then, a two-time-scale control mechanism is applied in PADQN to coordinate the IT and cooling systems more accurately and efficiently. In addition, to train and evaluate the proposed PADQN in a safe and quick way, we build a simulation platform to model the dynamics of IT workload and cooling systems simultaneously. Through extensive real-trace based simulations, we demonstrate that: 1) our algorithm can save up to 15% and 10% energy consumption in comparison with the baseline siloed and joint optimization approaches respectively; 2) our algorithm achieves more stable performance gain in terms of power consumption by adopting the parameterized action space; and 3) our algorithm leads to a better tradeoff between energy saving and service quality.",https://ieeexplore.ieee.org/document/8885255,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 253}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 1885}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 124}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1010}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 975}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 421}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 735}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1842}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 767}]",1.0,"['multilayer-perceptron', 'reinforcement-learning']",['traces'],,"['scheduling', 'power-management']",,2019 IEEE 39th International Conference on Distributed Computing Systems (ICDCS),True,['resource-provisioning'],,True,,,,,,,
134,Self-Adaptive Resource Management System in IaaS Clouds,"['F. Farahnakian', ' R. Bahsoon', ' P. Liljeberg', ' T. Pahikkala']",2016,"Resource management in cloud infrastructures is one of the most challenging problems due to the heterogeneity of resources, variability of the workload and scale of data centers. Efficient management of physical and virtual resources can be achieved considering performance requirements of hosted applications and infrastructure costs. In this paper, we present a self-adaptive resource management system based on a hierarchical multi-agent based architecture. The system uses novel adaptive utilization threshold mechanism and benefits from reinforcement learning technique to dynamically adjust CPU and memory thresholds for each Physical Machine (PM). It periodically runs a Virtual Machine (VM) placement optimization algorithm to keep the total resource utilization of each PM within given thresholds for improving Service Level Agreement (SLA) compliance. More-over, the algorithm consolidates VMs into the minimum number of active PMs in order to reduce the energy consumption. Experimental results on real workload traces show that our recourse management system can provide substantial improvement over other approaches in terms of performance requirements, energy consumption and the number of VM migrations.",https://ieeexplore.ieee.org/document/7820316,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 260}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 316}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1491}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 996}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 142}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 399}]",6.0,,,,,,2016 IEEE 9th International Conference on Cloud Computing (CLOUD),True,['resource-provisioning'],,True,,,,,,,
135,A Supervised Deep Learning Framework for Proactive Anomaly Detection in Cloud Workloads,"['S. Gupta', ' N. Muthiyan', ' S. Kumar', ' A. Nigam', ' D. A. Dinesh']",2017,"Cloud environment is highly prone to failures due to its distributed nature and inherent complexity. Proactive identification of failures aids the service providers to avert these failures by taking corrective actions before they actually happen. In this paper, we analyze the resource usage patterns to identify failures due to resource contention in cloud. The resource usage and performance metrics of the working system are analyzed at regular time instants to model the normal and anomalous working behaviors. A two stage framework has been implemented where a hybrid of long short term memory (LSTM) and bidirectional long short term memory (BLSTM) is used to predict the future resource usage and performance metric values in the first stage. In the second stage, the hybrid model is used to classify the expected state as either normal or abnormal. We evaluate the proposed anomaly detection model in a virtual environment set up using Docker containers. The experimental results show that the proposed algorithm outperforms state-of-the-art algorithms.",https://ieeexplore.ieee.org/document/8488109,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 297}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 53}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1693}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 182}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 344}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 248}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 234}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1319}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 334}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 173}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 397}]",1.0,['rnn'],"['kpis', 'metrics']",,['failure-detection'],['container'],2017 14th IEEE India Council International Conference (INDICON),True,['failure-management'],,,,,,,,,
136,VScaler: Autonomic Virtual Machine Scaling,"['L. Yazdanov', ' C. Fetzer']",2013,"Recent research results in cloud community found that cloud users increasingly force providers to shift from fixed bundle instance types(e.g. Amazon instances) to flexible bundles and shrinked billing cycles. This means that cloud applications can dynamically provision the used amount of resources in a more fine-grained fashion. This observation calls for approaches which are able to automatically implement fine granular VM resource allocation with respect to user-provided SLAs. In this work we propose VScaler, a framework which implements autonomic resource allocation using a novel approach to reinforcement learning.",https://ieeexplore.ieee.org/document/6676697,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 327}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 674}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 256}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 441}]",7.0,,,,,,2013 IEEE Sixth International Conference on Cloud Computing,True,['resource-provisioning'],,,,,,,,,
137,Adaptive Anomaly Detection in Performance Metric Streams,"['O. Ibidunmoye', ' A. Rezaie', ' E. Elmroth']",2018,"Continuous detection of performance anomalies such as service degradations has become critical in cloud and Internet services due to impact on quality of service and end-user experience. However, the volume and fast changing behavior of metric streams have rendered it a challenging task. Many diagnosis frameworks often rely on thresholding with stationarity or normality assumption, or on complex models requiring extensive offline training. Such techniques are known to be prone to spurious false-alarms in online settings as metric streams undergo rapid contextual changes from known baselines. Hence, we propose two unsupervised incremental techniques following a two-step strategy. First, we estimate an underlying temporal property of the stream via adaptive learning and, then we apply statistically robust control charts to recognize deviations. We evaluated our techniques by replaying over 40 time-series streams from the Yahoo! Webscope S5 datasets as well as four other traces of real Web service QoS and ISP traffic measurements. Our methods achieve high detection accuracy and few false-alarms, and better performance in general compared to an open-source package for time-series anomaly detection.",https://ieeexplore.ieee.org/document/8031053,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 343}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 238}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 461}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 184}]",14.0,['similarity-matching'],['kpis'],['novel-use'],['failure-detection'],,IEEE Transactions on Network and Service Management,True,['failure-management'],,,['anomaly-detection'],,,,,,
138,Automated anomaly detection and root cause analysis in virtualized cloud infrastructures,"['J. Lin', ' Q. Zhang', ' H. Bannazadeh', ' A. Leon-Garcia']",2016,"Cloud data centers today use visualization technologies to facilitate allocation of physical resources to multiple applications. As cloud data centers continue to grow in scale and complexity, effectively monitoring and identifying system anomalies is becoming a critical problem. Furthermore, due to complex dependencies between system components in a virtualized data center, a single cause of anomaly can typically trigger multiple alarms. Therefore there is also a need to efficiently analyze and identify the causes of the anomalies in a scalable and effective manner, in order to reduce the overhead of diagnosis and troubleshooting performed by the cloud operator. Motivated by these observations, we present a mechanism for automatic anomaly detection and root cause analysis in virtualized cloud data centers. We first use unsupervised learning techniques to identify abnormal system behaviors, and then propose a technique for root cause analysis with consideration to anomaly propagation among system components. Using a real virtualized cloud testbed, we show that our mechanism efficiently identifies system anomalies and accurately determines their causes.",https://ieeexplore.ieee.org/document/7502857,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 396}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 4}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 33}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 319}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 514}]",10.0,"['graph-mining', 'clustering']",['traces'],['novel-use'],"['failure-detection', 'root-cause-analysis']",['vm'],NOMS 2016 - 2016 IEEE/IFIP Network Operations and Management Symposium,True,['failure-management'],,,"['anomaly-detection', 'fault-localization']",,,,,,
139,Resource Provisioning with Budget Constraints for Adaptive Applications in Cloud Environments,"['Q. Zhu', ' G. Agrawal']",2012,"The recent emergence of clouds is making the vision of utility computing realizable, i.e., computing resources and services can be delivered, utilized, and paid for as utilities such as water or electricity. This, however, creates new resource provisioning problems. Because of the pay-as-you-go model, resource provisioning should be performed in a way to keep resource costs to a minimum, while meeting an application's needs. In this work, we focus on the use of cloud resources for a class of adaptive applications, where there could be application-specific flexibility in the computation that may be desired. Furthermore, there may be a fixed time-limit as well as a resource budget. Within these constraints, such adaptive applications need to maximize their Quality of Service (QoS), more precisely, the value of an application-specific benefit function, by dynamically changing adaptive parameters. We present the design, implementation, and evaluation of a framework that can support such dynamic adaptation for applications in a cloud computing environment. The key component of our framework is a multi-input-multi-output feedback control model-based dynamic resource provisioning algorithm which adopts reinforcement learning to adjust adaptive parameters to guarantee the optimal application benefit within the time constraint. Then a trained resource model changes resource allocation accordingly to satisfy the budget. We have evaluated our framework with two real-world adaptive applications, and have demonstrated that our approach is effective and causes a very low overhead.",https://ieeexplore.ieee.org/document/6122011,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 408}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1222}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 419}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1141}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 54}]",73.0,,,,['resource-consolidation'],,IEEE Transactions on Services Computing,True,['resource-provisioning'],,,,,,10.1145/1851476.1851516,Conference Paper,"['adaptive applications', 'cloud computing', 'resource provisioning']",
140,IFTM - Unsupervised Anomaly Detection for Virtualized Network Function Services,"['F. Schmidt', ' A. Gulenko', ' M. Wallschläger', ' A. Acker', ' V. Hennig', ' F. Liu', ' O. Kao']",2018,"Telecommunication system providers move their IP multimedia subsystems to virtualized services in the cloud. For such systems, dedicated hardware solutions provided a reliability of 99.999% in the past. Although virtualization offers more cost efficient usage of such services, it comes with higher complexity for providing reliable running software components due to the fragile computation stack. In order to hide the impact of such problematic behaviors, automatic mechanisms may help to detect degraded state anomalies in order to execute remediation actions. This work introduces IFTM as a framework for unsupervised anomaly detection in a distributed environment based on real-time monitoring data. The proposed approach consists of two key concepts using an automatic identity function and threshold learning to distinguish between normal and abnormal system behaviors. The evaluation is performed on a testbed running an open source implementation of the IP multimedia subsystem (Clearwater) executed on a replicated Openstack cloud environment. Results show the applicability of IFTM with high detection rates (98%) and low number of false alarms.",https://ieeexplore.ieee.org/document/8456348,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 445}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 1245}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 763}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 1889}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 597}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 360}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 1701}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 124}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('remediation' OR 'recovery')"", 'index': 255}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('remediation' OR 'recovery')"", 'index': 831}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 284}]",5.0,"['autoencoder', 'rnn']","['host-metrics', 'network-metrics']",['new-method'],['failure-detection'],['openstack'],2018 IEEE International Conference on Web Services (ICWS),True,['failure-management'],,True,['anomaly-detection'],,,,,,
141,An Anomaly Detection Framework for Autonomic Management of Compute Cloud Systems,"['D. Smith', ' Q. Guan', ' S. Fu']",2010,"In large-scale compute cloud systems, component failures become norms instead of exceptions. Failure occurrence as well as its impact on system performance and operation costs are becoming an increasingly important concern to system designers and administrators. When a system fails to function properly, health-related data are valuable for troubleshooting. However, it is challenging to effectively detect anomalies from the voluminous amount of noisy, high-dimensional data. The traditional manual approach is time-consuming, error-prone, and not scalable. In this paper, we present an autonomic mechanism for anomaly detection in compute cloud systems. A set of techniques is presented to automatically analyze collected data: data transformation to construct a uniform data format for data analysis, feature extraction to reduce data size, and unsupervised learning to detect the nodes acting differently from others. We evaluate our prototype implementation on an institute-wide compute cloud environment. The results show that our mechanism can effectively detect faulty nodes with high accuracy and low computation overhead.",https://ieeexplore.ieee.org/document/5615245,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 467}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 501}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 510}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('remediation' OR 'recovery')"", 'index': 192}]",40.0,"['dimensionality-reduction', 'similarity-matching', 'bayesian-network']",['host-metrics'],['new-method'],['failure-detection'],,2010 IEEE 34th Annual Computer Software and Applications Conference Workshops,True,['failure-management'],,True,['anomaly-detection'],,,,,,
142,Network Fault Prediction Based on CNN-LSTM Hybrid Neural Network,"['Z. Tan', ' P. Pan']",2019,"The network brings convenience and efficiency to people's life and work, and at the same time, the network will also cause loss to human beings because of failures, so it is particularly important to predict faults before the faults occur. The fault prediction technology can prepare the staff to repair the faults in advance, reduce the repair time of the fault, and thus reduce the loss caused by the faults. Therefore, this paper proposes a network log-based CNN-LSTM hybrid prediction model for wireless network faults: firstly, the network log is preprocessed, the sample is extracted by two-level time window, then the sample features are extracted by CNN, and finally, the extracted features are input into LSTM for prediction. To demonstrate the superiority of the CNN-LSTM hybrid neural network prediction model, in this experiment, it is compared with CNN and Random Forest. The result shows that the prediction performance of CNN-LSTM is better than the other two models.",https://ieeexplore.ieee.org/document/8805962,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 4}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prediction' OR 'failure prediction')"", 'index': 12}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 22}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 333}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 746}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 1}]",0.0,"['rnn', 'cnn']",['logs'],,['failure-prediction'],['network'],"2019 International Conference on Communications, Information System and Computer Engineering (CISCE)",True,['failure-management'],,True,,,,,,,
143,Detecting Anomaly in Big Data System Logs Using Convolutional Neural Network,"['S. Lu', ' X. Wei', ' Y. Li', ' L. Wang']",2018,"Nowadays, big data systems are being widely adopted by many domains for offering effective data solutions, such as manufacturing, healthcare, education, and media. Big data systems produce tons of unstructured logs that contain buried valuable information. However, it is a daunting task to manually unearth the information and detect system anomalies. A few automatic methods have been developed, where the cutting-edge machine learning technique is one of the most promising ways. In this paper, we propose a novel approach for anomaly detection from big data system logs by leveraging Convolutional Neural Networks (CNN). Different from other existing statistical methods or traditional rule-based machine learning approaches, our CNN-based model can automatically learn event relationships in system logs and detect anomaly with high accuracy. Our deep neural network consists of logkey2vec embeddings, three 1D convolutional layers, dropout layer, and max-pooling. According to our experiment, our CNN-based approach has better accuracy(reaches to 99%) compared to other approaches using Long Short term memory (LSTM) and Multilayer Perceptron (MLP) on detecting anomaly in Hadoop Distributed File System (HDFS) logs.",https://ieeexplore.ieee.org/document/8511880,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 31}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 40}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 590}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 195}]",8.0,['cnn'],['logs'],"['comparison', 'novel-use']",['failure-detection'],,"2018 IEEE 16th Intl Conf on Dependable, Autonomic and Secure Computing, 16th Intl Conf on Pervasive Intelligence and Computing, 4th Intl Conf on Big Data Intelligence and Computing and Cyber Science and Technology Congress(DASC/PiCom/DataCom/CyberSciTech)",True,['failure-management'],,,['anomaly-detection'],,,,,,
144,Automated IT system failure prediction: A deep learning approach,"['K. Zhang', ' J. Xu', ' M. R. Min', ' G. Jiang', ' K. Pelechrinis', ' H. Zhang']",2016,"In mission critical IT services, system failure prediction becomes increasingly important; it prevents unexpected system downtime, and assures service reliability for end users. While operational console logs record rich and descriptive information on the health status of those IT systems, existing system management technologies mostly use them in a labor-intensive forensics approach, i.e., identifying what went wrong after the fact. Recent efforts on log-based system management take an automation approach with text mining techniques, such as term frequency - inverse document frequency (TF-IDF). However, those techniques lead to a high-dimensional feature space, and are not easily generalizable to heterogeneous log formats. In this paper, we present a novel system that automatically parses streamed console logs and detects early warning signals for IT system failure prediction. In particular, our solution includes a log pattern extraction method by clustering together logs with similar format and content. We then resemble the TF-IDF idea by considering each pattern as a word and the set of patterns in each discretized epoch as a document. This leads to a feature space with significantly lower dimensionality that can provide robust signals for the status of the system. As system failures tend to occur very rare, we apply a recurrent neural network, namely, Long Short-Term Memory (LSTM), to deal with the “rarity” of labeled data in the training process. LSTM is able to capture the long-range dependency across sequences, therefore outperforms traditional supervised learning methods in our application domain. We evaluated and compared our proposed technology with state-of-the-art machine learning approaches using real log traces from two large enterprise systems. The results showed the advantage and potentials of our system in prediction of complex IT failures. To our knowledge, our work is the first that employs LSTM for log-based system failure prediction.",https://ieeexplore.ieee.org/document/7840733,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 43}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 148}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 133}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1531}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 362}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 77}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('remediation' OR 'recovery')"", 'index': 731}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 19}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 2}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 21}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 7}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 5}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 31}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 21}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('remediation' OR 'recovery')"", 'index': 61}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('remediation' OR 'recovery')"", 'index': 59}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('remediation' OR 'recovery')"", 'index': 462}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 34}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1030}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 580}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 7}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 983}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 531}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 56}]",50.0,"['rnn', 'pattern-matching', 'clustering']",['logs'],"['comparison', 'novel-use']",['failure-prediction'],,2016 IEEE International Conference on Big Data (Big Data),True,['failure-management'],True,True,['system-failure-prediction'],True,48.0,10.1109/BigData.2016.7840733,,,
145,Prediction of failure occurrence time based on system log message pattern learning,"['M. Sonoda', ' Y. Watanabe', ' Y. Matsumoto']",2012,"In order to avoid failures or diminish the impact of them, it is important to deal with them before its occurrence. Some existing approaches for online failure prediction are insufficient to handle the upcoming failures beforehand, because they cannot predict the failures early enough to execute workaround operations for failure. To solve this problem, we have developed a method to estimate the prediction period (the time period when a failure is expected to occur). Our method identifies the message patterns showing predictive signs of a certain failure through Bayesian learning from log messages and past failure reports. Using these patterns it predicts the occurrence of failures and their prediction period with sufficient interval. We conducted the evaluation of our approach with log data obtained from an actual system. The results shows that our method predicted the occurrence of failure with sufficient interval (60 minutes before the occurrence of failures) and sufficient accuracy (precision: over 0.7, recall: over 0.8).",https://ieeexplore.ieee.org/document/6211960,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 76}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 297}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 128}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 60}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 61}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 123}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('remediation' OR 'recovery')"", 'index': 347}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 541}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 111}]",9.0,['naive-bayes'],['logs'],['novel-use'],['failure-prediction'],,2012 IEEE Network Operations and Management Symposium,True,['failure-management'],,True,['system-failure-prediction'],,,,,,
146,Towards MapReduce based Bayesian deep learning network for monitoring big data applications,"['M. O. Shafiq', ' E. Torunski']",2017,"One of the most commonly used ways to monitor execution of software applications is by analyzing logs. Logs are execution foot-print of software applications that are produced and stored for real-time or post-execution analysis of execution. With the software applications becoming large, complex, distributed, web-scale, also called as big data applications, logs produced by such software applications are also large-scale. That means, such logs are large in volume, velocity and variety. That makes it crucial to have such logs analyzed in an automated, scalable and effective manner to ensure high veracity and have analytics with high value. In this paper, we present our proposed solution of a formal model for organizing and structuring logs. We then present a Bayesian deep learning network based analysis approach that utilizes the formal model for logs to detect and predict any possible faults and consequences of such faults. Moreover, we also present our MapReduce based distributed, parallel, single-pass and incremental approach to build, train and execute the proposed Bayesian deep learning framework. This helps in effective processing of logs on cloud platforms and therefore efficient handling of logs that are produced at the scale of big data by big data applications.",https://ieeexplore.ieee.org/document/8258159,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 97}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 765}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 1131}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 329}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 155}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 363}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 1534}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 6}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault detection' OR 'failure detection')"", 'index': 1548}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 196}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 962}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 340}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('fault detection' OR 'failure detection')"", 'index': 60}]",0.0,['multilayer-perceptron'],['logs'],,"['failure-prediction', 'failure-detection']",['mapreduce'],2017 IEEE International Conference on Big Data (Big Data),True,['failure-management'],,True,"['log-enhancement', 'system-failure-prediction']",,,,,,
147,Learning Latent Events From Network Message Logs,"['S. Satpathi', ' S. Deb', ' R. Srikant', ' H. Yan']",2019,"We consider the problem of separating error messages generated in large distributed data center networks into error events. In such networks, each error event leads to a stream of messages generated by hardware and software components affected by the event. These messages are stored in a giant message log. We consider the unsupervised learning problem of identifying the signatures of events that generated these messages; here, the signature of an error event refers to the mixture of messages generated by the event. One of the main contributions of the paper is a novel mapping of our problem which transforms it into a problem of topic discovery in documents. Events in our problem correspond to topics and messages in our problem correspond to words in the topic discovery problem. However, there is no direct analog of documents. Therefore, we use a non-parametric change-point detection algorithm, which has linear computational complexity in the number of messages, to divide the message log into smaller subsets called episodes, which serve as the equivalents of documents. After this mapping has been done, we use a well-known algorithm for topic discovery, called LDA, to solve our problem. We theoretically analyze the change-point detection algorithm, and show that it is consistent and has low sample complexity. We also demonstrate the scalability of our algorithm on a real data set consisting of 97 million messages collected over a period of 15 days, from a distributed data center network which supports the operations of a large wireless service provider.",https://ieeexplore.ieee.org/document/8782613,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 195}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 22}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 161}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 198}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 13}, {'database': 'ACM', 'search_string': ""'clustering' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 91}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 33}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 40}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 41}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 480}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 68}]",1.0,,['logs'],,"['root-cause-analysis', 'failure-detection']",['network'],IEEE/ACM Transactions on Networking,True,['failure-management'],,,['anomaly-detection'],,,10.1109/TNET.2019.2930040,Journal Article,,
148,Long short-term memory based operation log anomaly detection,"['R. Vinayakumar', ' K. P. Soman', ' P. Poornachandran']",2017,"Long short-term memory (LSTM) architecture is an important approach for capturing long-range temporal dependencies in sequences of arbitrary length. Moreover, stacked-LSTM (S-LSTM: formed by adding recurrent LSTM layer to the existing LSTM network in hidden layer) has capability to learn temporal behaviors quickly with sparse representations. To apply this to anomaly detection, we model the operation log samples of normal and anomalous events occurred in 1 minute time interval as time-series with the aim to detect and classify the events as either normal or anomalous. To select an appropriate LSTM network, experiments are conducted for various network parameters and network structures with the dataset provided by Cyber Security Data Mining Competition (CDMC2016). The experiments are run up to 1000 epochs with learning rate in the range [0.01-05]. S-LSTM network architecture has showed its strength by achieving the highest accuracy 0.996 with false positive rate 0.02 on the provided real-world test data by CDMC2016.",https://ieeexplore.ieee.org/document/8125846,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 196}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 92}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 268}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 37}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 209}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 340}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 279}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('anomaly detection' OR 'outlier detection')"", 'index': 208}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 123}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 394}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 328}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 158}]",19.0,['rnn'],['logs'],['novel-use'],['failure-detection'],,"2017 International Conference on Advances in Computing, Communications and Informatics (ICACCI)",True,['failure-management'],,True,['anomaly-detection'],,,,,,
149,Topical behavior prediction from massive logs,['S. Su'],2017,"In this paper, we study the topical behavior in a large scale. We use the network logs where each entry contains the entity ID, the timestamp, and the meta data about the activity. Both the temporal and the spatial relationships of the behavior are explored with the deep learning architectures combing the recurrent neural network (RNN) and the convolutional neural network (CNN). To make the behavioral data appropriate for the spatial learning in the CNN, we propose several reduction steps to form the topical metrics and to place them homogeneously like pixels in the images. The experimental result shows both temporal and spatial gains when compared against a multilayer perceptron (MLP) network. A new learning framework called the spatially connected convolutional networks (SCCN) is introduced to predict the topical metrics more efficiently.",https://ieeexplore.ieee.org/document/8258363,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 210}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1438}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 490}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 58}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 296}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 324}, {'database': 'arxiv', 'search_string': ""('DL' OR 'deep learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 51}, {'database': 'arxiv', 'search_string': ""('DL' OR 'deep learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 268}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 793}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 95}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 201}]",49.0,"['multilayer-perceptron', 'rnn', 'cnn']",['logs'],['novel-use'],['failure-detection'],['user'],2017 IEEE International Conference on Big Data (Big Data),True,['failure-management'],,True,['anomaly-detection'],,,,,,
150,A Hybrid Deep Learning-Based Model for Anomaly Detection in Cloud Datacenter Networks,"['S. Garg', ' K. Kaur', ' N. Kumar', ' G. Kaddoum', ' A. Y. Zomaya', ' R. Ranjan']",2019,"With the emergence of the Internet-of-Things (IoT) and seamless Internet connectivity, the need to process streaming data on real-time basis has become essential. However, the existing data stream management systems are not efficient in analyzing the network log big data for real-time anomaly detection. Further, the existing anomaly detection approaches are not proficient because they cannot be applied to networks, are computationally complex, and suffer from high false positives. Thus, in this paper a hybrid data processing model for network anomaly detection is proposed that leverages grey wolf optimization (GWO) and convolutional neural network (CNN). To enhance the capabilities of the proposed model, GWO and CNN learning approaches were enhanced with: 1) improved exploration, exploitation, and initial population generation abilities and 2) revamped dropout functionality, respectively. These extended variants are referred to as Improved-GWO (ImGWO) and Improved-CNN (ImCNN). The proposed model works in two phases for efficient network anomaly detection. In the first phase, ImGWO is used for feature selection in order to obtain an optimal trade-off between two objectives, i.e., reduced error rate and feature-set minimization. In the second phase, ImCNN is used for network anomaly classification. The efficacy of the proposed model is validated on benchmark (DARPA'98 and KDD'99) and synthetic datasets. The results obtained demonstrate that the proposed cloud-based anomaly detection model is superior in comparison to the other state-of-the-art models (used for network anomaly detection), in terms of accuracy, detection rate, false positive rate, and F-score. In average, the proposed model exhibits an overall improvement of 8.25%, 4.08%, and 3.62% in terms of detection rate, false positives, and accuracy, respectively; relative to standard GWO with CNN.",https://ieeexplore.ieee.org/document/8758843,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 449}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 25}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 201}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 131}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 8}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 178}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1323}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 244}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 672}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1014}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 583}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 112}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 116}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1652}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 46}, {'database': 'IEEE', 'search_string': ""'classification' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 24}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1292}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 8}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 135}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 55}]",1.0,['cnn'],['logs'],"['comparison', 'novel-use']",['failure-detection'],['network'],IEEE Transactions on Network and Service Management,True,['failure-management'],,,,,,,,,
151,"On the Use of Log-Based Model Checking, Clustering and Machine Learning for Process Behavior Prediction","['J. Ezpeleta', ' J. Fabra', ' P. Álvarez']",2018,"The paper proposes the use of Linear Temporal Logic (LTL) formulas for the behavioral description of the traces corresponding to log files. Such descriptions are used to group similar traces into classes applying standard clustering techniques. The classification results are used to feed a machine learning system able to predict, after a few initial events, the cluster to which an in-execution process is probably going to belong. The prediction model could be used to feed an on-line recommendation system so as to drive the process towards a desired cluster or to prevent it from being part of a non-desired one. The paper describes the used methodology and shows its validity by means of the application to a real log.",https://ieeexplore.ieee.org/document/8554490,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 538}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 97}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1294}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1235}, {'database': 'IEEE', 'search_string': ""'classification' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 799}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 105}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 248}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 243}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 308}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 215}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 333}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 819}]",1.0,"['multilayer-perceptron', 'clustering']","['traces', 'logs']",,,['process'],"2018 Fifth International Conference on Social Networks Analysis, Management and Security (SNAMS)",True,['resource-provisioning'],,,,,,,,,
152,Software defect prediction using supervised learning algorithm and unsupervised learning algorithm,"['A. Chug', ' S. Dhall']",2013,"Software defect prediction has recently attracted attention of many software quality researchers. One of the major areas in current project management software is to effectively utilize resources to make meaningful impact on time and cost. A pragmatic assessment of metrics is essential in order to comprehend the quality of software and to ensure corrective measures. Software defect prediction methods are majorly used to study the impact areas in software using different techniques which comprises of neural network (NN) techniques, clustering techniques, statistical method and machine learning methods. These techniques of Data mining are applied in building software defect prediction models which improve the software quality. The aim of this paper is to propose various classification and clustering methods with an objective to predict software defect. To predict software defect we analyzed classification and clustering techniques. The performance of three data mining classifier algorithms named J48, Random Forest, and Naive Bayesian Classifier (NBC) are evaluated based on various criteria like ROC, Precision, MAE, RAE etc. Clustering technique is then applied on the data set using k-means, Hierarchical Clustering and Make Density Based Clustering algorithm. Evaluation of results for clustering is based on criteria like Time Taken, Cluster Instance, Number of Iterations, Incorrectly Clustered Instance and Log Likelihood etc. A thorough exploration of ten real time defect datasets of NASA[1] software project, followed by various applications on them finally results in defect prediction.",https://ieeexplore.ieee.org/document/6832328,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 760}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 566}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 450}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 48}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 745}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1150}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1006}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 35}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1665}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 932}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 564}]",3.0,"['random-forest', 'naive-bayes', 'clustering', 'decision-tree']",['software-metrics'],"['comparison', 'novel-use']",['failure-prevention'],['source-code'],Confluence 2013: The Next Generation Information Technology Summit (4th International Conference),True,['failure-management'],,True,['software-defect-prediction'],,,,,,
153,A Survey on Failure Prediction of Large-Scale Server Clusters,"['Z. Xue', ' X. Dong', ' S. Ma', ' W. Dong']",2007,"As the size and complexity of cluster systems grows, failure rates accelerate dramatically. To reduce the disaster caused by failures, it is desirable to identify the potential failures ahead of their occurrence. In this paper, we survey the state of the art in failure prediction of cluster systems. The characteristic of failures in cluster systems are addressed, and some statistic results are shown. We explore the ways of the collection and preprocessing of data for failure prediction, and suggest a procedure for preprocessing the records in automatically generated log files. Focused on the main idea of five prediction methods, including statistic based threshold, time series analysis, rule-based classification, Bayesian network models and semi-Markov process models, are analyzed respectively. In addition, concerning the accuracy and practicality, we present five metrics for evaluating the failure prediction techniques and compare the five techniques with the five metrics.",https://ieeexplore.ieee.org/document/4287779,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 902}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 924}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1415}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 51}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 82}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 740}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1593}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 155}]",34.0,,,['survey'],['failure-prediction'],['cluster'],"Eighth ACIS International Conference on Software Engineering, Artificial Intelligence, Networking, and Parallel/Distributed Computing (SNPD 2007)",True,['failure-management'],True,,['system-failure-prediction'],,,,,,
154,An alarm correlation algorithm for network management based on root cause analysis,"['D. S. Kim', ' H. Shinbo', ' H. Yokota']",2011,"The alarm correlation is an essential function of network management systems to provide detection, isolation and correlation of unusual operational behaviour of telecommunication network. However, existing alarm correlation approaches still rely on the manual processing, and depend on the knowledge of the network operators. Since, the telecommunication network produces a number of alarms which are so called the alarm floods, it could be very difficult for the network operators to detect the root cause problems in a short period of time. Therefore, we propose the alarm correlation algorithm which is able to isolate and correlate the root causes in a very short time. In addition, we show that the proposed algorithm performs well in terms of efficiency of analyzing alarms and accuracy of identifying root cause.",https://ieeexplore.ieee.org/document/5746028,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 1}]",1.0,"['graph-mining', 'clustering']","['packet-content', 'logs']",['new-method'],['root-cause-analysis'],['network'],13th International Conference on Advanced Communication Technology (ICACT2011),True,['failure-management'],,,['root-cause-diagnosis'],,,,,,
155,Automating Root Cause Analysis via Machine Learning in Agile Software Testing Environments,"['J. Kahles', ' J. Törrönen', ' T. Huuhtanen', ' A. Jung']",2019,"We apply machine learning to automate the root cause analysis in agile software testing environments. In particular, we extract relevant features from raw log data after interviewing testing engineers (human experts). Initial efforts are put into clustering the unlabeled data, and despite obtaining weak correlations between several clusters and failure root causes, the vagueness in the rest of the clusters leads to the consideration of labeling. A new round of interviews with the testing engineers leads to the definition of five ground-truth categories. Using manually labeled data, we train artificial neural networks that either classify the data or pre-process it for clustering. The resulting method achieves an accuracy of 88.9%. The methodology of this paper serves as a prototype or baseline approach for the extraction of expert knowledge and its adaptation to machine learning techniques for root cause analysis in agile environments.",https://ieeexplore.ieee.org/document/8730163,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 2}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 537}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 108}, {'database': 'IEEE', 'search_string': ""'classification' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 5}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 566}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 888}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 2}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 6}]",0.0,"['multilayer-perceptron', 'autoencoder', 'gmm']","['kpis', 'test-cases', 'logs']",,['root-cause-analysis'],,"2019 12th IEEE Conference on Software Testing, Validation and Verification (ICST)",True,['failure-management'],,,['fault-localization'],,,,,,
156,Automatic fault characterization via abnormality-enhanced classification,"['G. Bronevetsky', ' I. Laguna', ' B. R. de Supinski', ' S. Bagchi']",2012,"Enterprise and high-performance computing systems are growing extremely large and complex, employing many processors and diverse software/hardware stacks. As these machines grow in scale, faults become more frequent and system complexity makes it difficult to detect and to diagnose them. The difficulty is particularly large for faults that degrade system performance or cause erratic behavior but do not cause outright crashes. The cost of these errors is high since they significantly reduce system productivity, both initially and by time required to resolve them. Current system management techniques do not work well since they require manual examination of system behavior and do not identify root causes. When a fault is manifested, system administrators need timely notification about the type of fault, the time period in which it occurred and the processor on which it originated. Statistical modeling approaches can accurately characterize normal and abnormal system behavior. However, the complex effects of system faults are less amenable to these techniques. This paper demonstrates that the complexity of system faults makes traditional classification and clustering algorithms inadequate for characterizing them. We design novel techniques that combine classification algorithms with information on the abnormality of application behavior to improve detection and characterization accuracy significantly. Our experiments demonstrate that our techniques can detect and characterize faults with 85% accuracy, compared to just 12% accuracy for direct applications of traditional techniques.",https://ieeexplore.ieee.org/document/6263926,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 23}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 291}, {'database': 'IEEE', 'search_string': ""'classification' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 15}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault detection' OR 'failure detection')"", 'index': 505}]",47.0,"['random-forest', 'decision-tree']",['runs'],['new-method'],['root-cause-analysis'],['hpc'],IEEE/IFIP International Conference on Dependable Systems and Networks (DSN 2012),True,['failure-management'],,False,['root-cause-diagnosis'],,,,,,
157,A GMM and SVM Combined Approach for Automatically Software Fault Localization,"['X. Wu', ' W. Zheng', ' J. Chen', ' H. Bai', ' D. Hu', ' D. Mu']",2018,"To improve the efficiency and accuracy of automatic fault localization. We propose an approach to direct fault localization by applying Gaussian Mixture Model (GMM) and Support Vector Machine (SVM), which are two mathematical models with excellent classification and prediction abilities. We first preprocess the training data using GMM-based clustering algorithm. Then the constant penalty factor of SVM is replaced with two adjustable ones. After that, we find out the mapping relationships between the coverage information and the execution result of each test case by virtue of the robust learning ability of modified SVM. An efficiency comparison between our technique and others on Siemens Suite is carried out afterwards. The experiment result indicates that our localization approach achieves a better accuracy in single and multiple faults localization without increasing testing cost.",https://ieeexplore.ieee.org/document/8706285,True,"[{'database': 'IEEE', 'search_string': ""'classification' AND ('fault localization' OR 'failure localization')"", 'index': 22}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault localization' OR 'failure localization')"", 'index': 18}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault localization' OR 'failure localization')"", 'index': 2}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault localization' OR 'failure localization')"", 'index': 1}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 31}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 23}]",0.0,"['support-vector-machine', 'gmm']",['test-cases'],['new-method'],['root-cause-analysis'],['software'],2018 IEEE International Conference on Progress in Informatics and Computing (PIC),True,['failure-management'],,True,['fault-localization'],,,,,,
158,Fault localization with nearest neighbor queries,"['M. Renieres', ' S. P. Reiss']",2003,"We present a method for performing fault localization using similar program spectra. Our method assumes the existence of a faulty run and a larger number of correct runs. It then selects according to a distance criterion the correct run that most resembles the faulty run, compares the spectra corresponding to these two runs, and produces a report of ""suspicious"" parts of the program. Our method is widely applicable because it does not require any knowledge of the program input and no more information from the user than a classification of the runs as either ""correct"" or ""faulty"". To experimentally validate the viability of the method, we implemented it in a tool, Whither, using basic block profiling spectra. We experimented with two different similarity measures and the Siemens suite of 132 programs with injected bugs. To measure the success of the tool, we developed a generic method for establishing the quality of a report. The method is based on the way an ""ideal user"" would navigate the program using the report to save effort during debugging. The best results obtained were, on average, above 50%, meaning that our ideal user would avoid looking half of the program.",https://ieeexplore.ieee.org/document/1240292,True,"[{'database': 'IEEE', 'search_string': ""'classification' AND ('fault localization' OR 'failure localization')"", 'index': 25}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 16}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault localization' OR 'failure localization')"", 'index': 14}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault localization' OR 'failure localization')"", 'index': 12}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault localization' OR 'failure localization')"", 'index': 59}]",807.0,['similarity-matching'],['runs'],,['root-cause-analysis'],['software'],"18th IEEE International Conference on Automated Software Engineering, 2003. Proceedings.",True,['failure-management'],True,,['fault-localization'],True,72.0,,Conference Paper,,
159,Machine Learning-Based Link Fault Identification and Localization in Complex Networks,"['S. M. Srinivasan', ' T. Truong-Huu', ' M. Gurusamy']",2019,"With the proliferation of network devices and rapid development in information technology, networks such as Internet of Things are increasing in size and becoming more complex with heterogeneous wired and wireless links. In such networks, link faults may result in a link disconnection without immediate replacement or a link reconnection, e.g., a wireless node changes its access point. Identifying whether a link disconnection or a link reconnection has occurred and localizing the failed link become a challenging problem. An active probing approach requires a long time to probe the network by sending signaling messages on different paths, thus incurring significant communication delay and overhead. In this paper, we adopt a passive approach and develop a three-stage machine learning-based technique for link fault identification and localization (ML-LFIL) by analyzing the measurements captured from the normal traffic flows, including aggregate flow rate, end-to-end delay, and packet loss. ML-LFIL learns the traffic behavior in normal working conditions and different link fault scenarios. We train the learning model using support vector machine, multilayer perceptron, and random forest. We implement ML-LFIL and carry out extensive experiments using Mininet platform. Performance studies show that ML-LFIL achieves high accuracy while requiring much lower fault localization time compared to the active probing approach.",https://ieeexplore.ieee.org/document/8676028,True,"[{'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault localization' OR 'failure localization')"", 'index': 30}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault localization' OR 'failure localization')"", 'index': 7}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 44}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 6}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 1}, {'database': 'arxiv', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault localization' OR 'failure localization')"", 'index': 2}, {'database': 'arxiv', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault localization' OR 'failure localization')"", 'index': 4}]",3.0,"['multilayer-perceptron', 'random-forest', 'support-vector-machine']",['network-traffic'],,"['root-cause-analysis', 'failure-detection']","['network', 'link']",IEEE Internet of Things Journal,True,['failure-management'],,,"['anomaly-detection', 'fault-localization']",,,10.1109/JIOT.2019.2908019,,,
160,End-to-end service failure diagnosis using belief networks,"['M. Steinder', ' A. S. Sethi']",2002,"We present fault localization techniques suitable for diagnosing end-to-end service problems in communication systems with complex topologies. We refine a layered system model that represents relationships between services and functions offered between neighboring protocol layers. In a given layer, an end-to-end service between two hosts may be provided using multiple host-to-host services offered in this layer between two hosts on the end-to-end path. Relationships among end-to-end and host-to-host services form a bipartite probabilistic dependency graph whose structure depends on the network topology in the corresponding protocol layer. When an end-to-end service fails or experiences performance problems it is important to efficiently find the responsible host-to-host services. Finding the most probable explanation (MPE) of the observed symptoms is NP-hard. We propose two fault localization techniques based on Pearl's (1988) iterative algorithms for singly connected belief networks. The probabilistic dependency graph is transformed into a belief network, and then the approximations based on Pearl's algorithms and exact bucket tree elimination algorithm are designed and evaluated through extensive simulation study.",https://ieeexplore.ieee.org/document/1015595,True,"[{'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault localization' OR 'failure localization')"", 'index': 32}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 122}]",101.0,['bayesian-network'],['runs'],['new-method'],['root-cause-analysis'],,NOMS 2002. IEEE/IFIP Network Operations and Management Symposium. ' Management Solutions for the New Communications World'(Cat. No.02CH37327),True,['failure-management'],,,['root-cause-diagnosis'],,,,,,
161,Online network performance degradation localization using probabilistic inference and change detection,"['A. Johnsson', ' C. Meirosu', ' C. Flinta']",2014,"Detecting and localizing performance degradations is a difficult problem that increases in importance as telecom network transition to all-packet equipment. Operators require solutions that are accurate in localization and do not impose large additional costs in terms of hardware deployment or manual labor for operations. Existing commercial solutions are generally difficult to operate, while many academic proposals typically require significant computational resources and are difficult to adapt to production networks. This paper describes a novel network fault localization algorithm based on active network measurements, probabilistic inference and change detection. The algorithm is computationally efficient for networks with thousands of nodes and requires few configuration parameters. Results obtained in a simulated environment on tree topologies show that the solution provides fast and accurate localization of performance degradations.",https://ieeexplore.ieee.org/document/6838255,True,"[{'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault localization' OR 'failure localization')"", 'index': 36}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 108}]",7.0,['markov-model'],['kpis'],,['root-cause-analysis'],['network'],2014 IEEE Network Operations and Management Symposium (NOMS),True,['failure-management'],,,,,,,,,
162,Linear and Logistic Regression Based Monitoring for Resource Management in Cloud Networks,"['M. Daraghmeh', ' S. Bani Melhem', ' A. Agarwal', ' N. Goel', ' M. Zaman']",2018,"Understanding and implementing the effective techniques to manage the infrastructure resources of cloud datacenter has recently become important. The energy consumption and ineffective resource utilization can lead to an increase in the operational cost of cloud provider side, which in turn increases the cloud services cost at cloud consumer side. One of the effective techniques to address these issues in cloud datacenters is the server consolidation by allowing multiple virtual machines (VMs) include varying workload to host in a single physical machine. This leads to an increase in the resource utilization, and reduced power consumption by turning off the idle physical machines. However, consolidating the virtual machines due to varying workload in cloud applications can cause a violation of service level agreement. In this paper, we propose a model based on linear and logistic regression to detect overloaded hosts by dynamically generating rules based on historical data of hosts and datacenter in order to update association functions to address and adapt the changes of different types of workloads running on the cloud provider datacenter. The experiments and simulation results based on dynamic workloads show the proposed algorithm significantly outperforms the other competitive host detection algorithms.",https://ieeexplore.ieee.org/document/8458022,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 14}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 102}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('cloud computing')"", 'index': 10}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('cloud')"", 'index': 2}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 1}, {'database': 'IEEE', 'search_string': ""'regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 2}]",2.0,,,,,,2018 IEEE 6th International Conference on Future Internet of Things and Cloud (FiCloud),True,['resource-provisioning'],,True,,,,,,,
163,Using Logistic Regression to Improve Virtual Machines Management in Cloud Computing Systems,"['M. B. Issa', ' M. Daraghmeh', ' Y. Jararweh', ' M. Al-Ayyoub', ' M. Alsmirat', ' E. Benkhelifa']",2017,"Cloud computing (CC) is a computing model that enables its customers to access a shared pool of resources (e.g., storage, network, servers, etc.) through the Internet with a pay-per-use pricing model. Different service models are employed in CC including the Platform-as-a-Service (PaaS) model, in which the costumers request a certain set of resources and the cloud service providers provide these resources in the form of a virtual machine (VM) running on one of the thousands of hosting servers or physical machines (PMs) of a data center. Where to ""place"" VMs, how to ""execute"" them and whether there is a need to ""move/migrate"" them are important decisions that affect the overall resource utilization and power consumption in the hosting data center. VM consolidation is a technique of migrating or consolidating VMs to PMs in order to prevent the PMs from being overloaded or reduce the number of active PMs and increase their utilization. Consolidation techniques measure PM utilization to decide whether to consolidate the VMs running on it or migrate some of them to another PM. This study aims to optimize resource utilization and energy efficiency in cloud data centers by proposing a new Logistic Regression based host overloading prediction technique that can be used by any VM consolidation technique. The new algorithm have been evaluated using a dynamic workload using the CloudSim simulator. The simulation results show that the proposed algorithm outperforms all other known host status prediction techniques.",https://ieeexplore.ieee.org/document/8108812,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 39}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 19}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('cloud computing')"", 'index': 3}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('cloud')"", 'index': 4}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 2}, {'database': 'IEEE', 'search_string': ""'regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 4}]",2.0,,,,,,2017 IEEE 14th International Conference on Mobile Ad Hoc and Sensor Systems (MASS),True,['resource-provisioning'],,True,,,,,,,
164,Experience report: Anomaly detection of cloud application operations using log and cloud metric correlation analysis,"['M. Farshchi', ' J. Schneider', ' I. Weber', ' J. Grundy']",2015,"Failure of application operations is one of the main causes of system-wide outages in cloud environments. This particularly applies to DevOps operations, such as backup, redeployment, upgrade, customized scaling, and migration that are exposed to frequent interference from other concurrent operations, configuration changes, and resources failure. However, current practices fail to provide a reliable assurance of correct execution of these kinds of operations. In this paper, we present an approach to address this problem that adopts a regression-based analysis technique to find the correlation between an operation's activity logs and the operation activity's effect on cloud resources. The correlation model is then used to derive assertion specifications, which can be used for runtime verification of running operations and their impact on resources. We evaluated our proposed approach on Amazon EC2 with 22 rounds of rolling upgrade operations while other types of operations were running and random faults were injected. Our experiment shows that our approach successfully managed to raise alarms for 115 random injected faults, with a precision of 92.3%.",https://ieeexplore.ieee.org/document/7381796,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 50}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 299}, {'database': 'IEEE', 'search_string': ""'regression' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 26}, {'database': 'IEEE', 'search_string': ""'regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 138}, {'database': 'IEEE', 'search_string': ""'regression' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 30}]",34.0,"['linear-regression', 'clustering']","['traces', 'logs']",['novel-use'],['failure-detection'],,2015 IEEE 26th International Symposium on Software Reliability Engineering (ISSRE),True,['failure-management'],,,['anomaly-detection'],,,,,,
165,LiRCUP: Linear Regression Based CPU Usage Prediction Algorithm for Live Migration of Virtual Machines in Data Centers,"['F. Farahnakian', ' P. Liljeberg', ' J. Plosila']",2013,"Virtualization is a vital technology of cloud computing which enables the partition of a physical host into several Virtual Machines (VMs). The number of active hosts can be reduced according to the resources requirements using live migration in order to minimize the power consumption in this technology. However, the Service Level Agreement (SLA) is essential for maintaining reliable quality of service between data centers and their users in the cloud environment. Therefore, reduction of the SLA violation level and power costs are considered as two objectives in this paper. We present a CPU usage prediction method based on the linear regression technique. The proposed approach approximates the short-time future CPU utilization based on the history of usage in each host. It is employed in the live migration process to predict over-loaded and under-loaded hosts. When a host becomes over-loaded, some VMs migrate to other hosts to avoid SLA violation. Moreover, first all VMs migrate from a host while it becomes under-loaded. Then, the host switches to the sleep mode for reducing power consumption. Experimental results on the real workload traces from more than a thousand Planet Lab VMs show that the proposed technique can significantly reduce the energy consumption and SLA violation rates.",https://ieeexplore.ieee.org/document/6619533,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 100}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 37}, {'database': 'IEEE', 'search_string': ""'regression' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 30}]",67.0,,,,['workload-prediction'],,2013 39th Euromicro Conference on Software Engineering and Advanced Applications,True,['resource-provisioning'],,,,,,,,,
166,KSwSVR: A New Load Forecasting Method for Efficient Resources Provisioning in Cloud,"['R. Hu', ' J. Jiang', ' G. Liu', ' L. Wang']",2013,"Cloud provider should ensure QoS while maximizing resources utilization. One optimal strategy is to timely allocate resources in a fine-grained mode according to the actual resources demand of applications. The necessary precondition of this strategy is obtaining future load information in advance. We propose a multi-step-ahead load forecasting method, KSwSVR, based on statistical learning theory which is suitable for the complex and dynamic characteristics of the cloud computing environment. It integrates an improved support vector regression algorithm and Kalman smoother. Public trace data taken from multi-types of resources were used to verify its prediction accuracy, stability and adaptability, comparing with AR, BPNN and standard SVR. CPU allocation experiment indicated that KSwSVR can effectively reduce resources consumption while meeting Service Level Agreements requirement.",https://ieeexplore.ieee.org/document/6649686,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 126}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 1041}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1594}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1170}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 800}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1178}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 125}, {'database': 'IEEE', 'search_string': ""'regression' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 171}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 871}]",9.0,,,,,,2013 IEEE International Conference on Services Computing,True,['resource-provisioning'],,True,,,,,,,
167,Cloud Resource Demand Prediction using Differential Evolution based Learning,"['J. Kumar', ' A. K. Singh']",2019,"Today's digital world generates ample amount of data through interconnected heterogeneous devices that must be stored and processed efficiently for uninterrupted services. The distributed infrastructures have shown the capability of addressing the storage and computing issues of big data. The cloud paradigm is enabled with characteristics including multi-tenancy, on-demand, virtualization, scalability and many more. However, the cloud resources must be used efficiently to reduce power consumption and carbon footprints. This paper presents a workload prediction scheme based on differential evolution that can be used for effective virtual machine allocation. The forecast accuracy of the proposed scheme is evaluated over Google's real world trace and compared with existing state-of-art prediction approaches. We observed a significant reduction in forecast error upto 71% and 88% over back propagation and linear regression based forecasting approaches respectively.",https://ieeexplore.ieee.org/document/8843680,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 132}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 740}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 482}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1333}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 758}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 826}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 214}, {'database': 'IEEE', 'search_string': ""'regression' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 96}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 631}]",1.0,"['multilayer-perceptron', 'genetic-programming']",,"['comparison', 'novel-use']",['workload-prediction'],['vm'],2019 7th International Conference on Smart Computing & Communications (ICSCC),True,['resource-provisioning'],,True,,,,,,,
168,Workload Predicting-Based Automatic Scaling in Service Clouds,"['J. Yang', ' C. Liu', ' Y. Shang', ' Z. Mao', ' J. Chen']",2013,"Service platforms have disadvantages such as they have long construction periods, low resource utilizations and isolated constructions. Migrating service platforms into clouds can solve these problems. The scalability is an important characteristic of service clouds. With the scalability, the service cloud can offer on-demand capacities to different services. In order to achieve the scalability, we need to know when and how to scale virtual resources assigned to different services. In this paper, a linear regression model is used to predict the workload. Based on this predicted workload, an auto-scaling mechanism is proposed to scale virtual resources at different resource levels in service clouds. The automatic scaling mechanism combines the real-time scaling and the pre-scaling. Finally experimental results are provided to demonstrate that our approach can satisfy the user SLA while keeping scaling costs low.",https://ieeexplore.ieee.org/document/6740226,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 140}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 156}]",15.0,,,,,,2013 IEEE Sixth International Conference on Cloud Computing,True,['resource-provisioning'],,,,,,,,,
169,On the Provision of SaaS-Level Quality of Service within Heterogeneous Private Clouds,"['J. P. Orellana', ' M. B. Caminero', ' C. Carrión']",2014,"The efficient utilization of computing resources, consisting of multi-core CPUs, GPUs and FPGAs, has become an interesting research problem for achieving high performance on heterogeneous Cloud computing platforms. In particular, FPGA accelerators can provide significant business value in Cloud environments due to its great computing capacity with predictable latency and low power consumption. In this paper, a Software as a Service (SaaS) model is enhanced with Quality of Service (QoS) support, harnessing such heterogeneous hardware architecture (composed of conventional CPUs plus FPGAs as accelerator). More precisely, the proposal takes into account timing user requirements to manage virtual resources. Hence, novel heterogeneous-aware resource allocation and scheduling algorithms are presented, which can be used both on-demand and in-advance. A lineal regression model that predicts the cost of the requested service is combined with a simple heuristic algorithm in order to allocate different types of Virtual Machines (VMs). Moreover, the framework provides the service efficiently by using an adapted scheduling algorithm that combines CPUs and accelerator resources.",https://ieeexplore.ieee.org/document/7027490,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 144}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 72}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 60}, {'database': 'ACM', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 55}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 57}]",177.0,,,,"['workload-prediction', 'resource-consolidation']",,2014 IEEE/ACM 7th International Conference on Utility and Cloud Computing,True,['resource-provisioning'],,,,,,10.1109/UCC.2014.23,Conference Paper,"['Cloud Computing', 'Heterogeneous Resources', 'SaaS', 'FPGAs', 'Quality of Service']",
170,Spatial Support Vector Regression to Detect Silent Errors in the Exascale Era,"['O. Subasi', ' S. Di', ' L. Bautista-Gomez', ' P. Balaprakash', ' O. Unsal', ' J. Labarta', ' A. Cristal', ' F. Cappello']",2016,"As the exascale era approaches, the increasing capacity of high-performance computing (HPC) systems with targeted power and energy budget goals introduces significant challenges in reliability. Silent data corruptions (SDCs) or silent errors are one of the major sources that corrupt the executionresults of HPC applications without being detected. In this work, we explore a low-memory-overhead SDC detector, by leveraging epsilon-insensitive support vector machine regression, to detect SDCs that occur in HPC applications that can be characterized by an impact error bound. The key contributions are three fold. (1) Our design takes spatialfeatures (i.e., neighbouring data values for each data point in a snapshot) into training data, such that little memory overhead (less than 1%) is introduced. (2) We provide an in-depth study on the detection ability and performance with different parameters, and we optimize the detection range carefully. (3) Experiments with eight real-world HPC applications show thatour detector can achieve the detection sensitivity (i.e., recall) up to 99% yet suffer a less than 1% of false positive rate for most cases. Our detector incurs low performance overhead, 5% on average, for all benchmarks studied in the paper. Compared with other state-of-the-art techniques, our detector exhibits the best tradeoff considering the detection ability and overheads.",https://ieeexplore.ieee.org/document/7515717,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 266}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 420}, {'database': 'ACM', 'search_string': ""'regression' AND ('remediation' OR 'recovery')"", 'index': 92}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('remediation' OR 'recovery')"", 'index': 34}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 77}, {'database': 'ACM', 'search_string': ""'regression' AND ('cloud')"", 'index': 89}]",1.0,['support-vector-machine'],,,['failure-detection'],,"2016 16th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGrid)",True,['failure-management'],,,,,,10.1109/CCGrid.2016.33,Conference Paper,,
171,Predicting web service levels during VM live migrations,"['H. Hlavacs', ' T. Treutner']",2011,"The ability to live migrate virtual machines (VMs) between physical servers without any perceivable service interruption is pivotal for building more energy efficient Cloud Computing infrastructures in the future. Nevertheless, energy efficiency is not worth the effort if quality metrics (e.g., QoS, QoE) are severely decreased by, e.g., dynamic consolidation using live migration. We identify the most significant utilization metrics to predict the service level during live migrations for a web server scenario. We show important correlations, give reasons and draw conclusions for systems using live migration for yielding higher energy efficiency. We also give reasons for extending the current hypervisors' capabilities regarding VM utilization collection and reporting. We present the effects of live migration on service levels for different workload scenarios. In particular, we demonstrate that live migration should be done preventively. This anticipates disproportional high service level degradation due to live migration. We examine the most important utilization metrics for predicting the service level by both stepwise and exhaustive regression. As a result, we can explain 90% of the service level variance during live migration with a single variable, using more variables yields 95%.",https://ieeexplore.ieee.org/document/6096464,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 293}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 421}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 119}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 229}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 72}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 76}]",12.0,,,,,,2011 5th International DMTF Academic Alliance Workshop on Systems and Virtualization Management: Standards and the Cloud (SVM),True,['resource-provisioning'],,,,,,,,,
172,Fisher: An Efficient Container Load Prediction Model with Deep Neural Network in Clouds,"['X. Tang', ' Q. Liu', ' Y. Dong', ' J. Han', ' Z. Zhang']",2018,"Recently, more and more applications have been deployed in container, and prediction of container load is essential in Clouds for improving resource utilization and achieving service-level agreements. However, accurate prediction of container load in Clouds remains an extremely challenge because the container load fluctuates drastically at small timescales. Furthermore, with many metrics of container in Clouds, it is hard to find which metrics are going to be useful. To address these challenges, we design an efficient container load prediction model named Fisher to improve the accuracy and efficiency of prediction. It mainly includes two modules: a metrics selection module and a neural network training module. We first selects relevant metrics by metrics selection module which is a novel algorithm for shape-based time-series clustering. Afterwards, we apply a powerful deep neural network model built with bidirectional long short-term memory to predict actual load one-step-ahead. We evaluate Fisher using a 30-day load trace from a data center with 500 containers. Experiments show that Fisher can reduce training metrics while maintaining prediction accuracy. More importantly, our model significantly improves prediction accuracy over 50\% compared to other state-of-the-art methods based on auto regressive integrated moving-average and long short-term memory.",https://ieeexplore.ieee.org/document/8672282,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 344}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1046}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 283}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1180}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1001}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 578}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 228}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 830}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 648}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 113}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 69}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1711}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 162}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 134}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 145}, {'database': 'IEEE', 'search_string': ""'regression' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 263}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 612}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 193}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 1258}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 236}]",4.0,['rnn'],['traces'],['novel-use'],['workload-prediction'],,"2018 IEEE Intl Conf on Parallel & Distributed Processing with Applications, Ubiquitous Computing & Communications, Big Data & Cloud Computing, Social Computing & Networking, Sustainable Computing & Communications (ISPA/IUCC/BDCloud/SocialCom/SustainCom)",True,['resource-provisioning'],,True,,,,,,,
173,Being Accurate Is Not Enough: New Metrics for Disk Failure Prediction,"['J. Li', ' R. J. Stones', ' G. Wang', ' Z. Li', ' X. Liu', ' K. Xiao']",2016,"Traditionally, disk failure prediction accuracy is used to evaluate disk failure prediction model. However, accuracy may not reflect their practical usage (protecting against failures, rather than only predicting failures) in cloud storage systems. In this paper, we propose two new metrics for disk failure prediction models: migration rate, which measures how much at-risk data is protected as a result of correct failure predictions, and mismigration rate, which measures how much data is migrated needlessly as a result of false failure predictions. To demonstrate their effectiveness, we compare disk failure prediction methods: (a) a classification tree (CT) model vs. a state-of-the-art recurrent neural network (RNN) model, and (b) a proposed residual life prediction model based on gradient boosted regression trees (GBRTs) vs. RNN. While prediction accuracy experiments favor the RNN model, migration rate experiments can favor the CT and GBRT models (depending on transfer rates). We conclude that prediction accuracy can be a misleading metric. Moreover, the proposed GBRT model offers a practical improvement in disk failure prediction in real-world data centers.",https://ieeexplore.ieee.org/document/7794331,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 407}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 1186}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('remediation' OR 'recovery')"", 'index': 130}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 489}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 691}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 501}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 708}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 73}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 1301}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 114}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 303}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 179}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 48}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 29}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 144}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 25}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('remediation' OR 'recovery')"", 'index': 604}, {'database': 'IEEE', 'search_string': ""'classification' AND ('remediation' OR 'recovery')"", 'index': 799}, {'database': 'IEEE', 'search_string': ""'regression' AND ('remediation' OR 'recovery')"", 'index': 281}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('remediation' OR 'recovery')"", 'index': 648}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 954}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1456}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 534}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 108}]",19.0,"['rnn', 'decision-tree', 'regression-tree']",['host-metrics'],['comparison'],['failure-prediction'],['hard-drive'],2016 IEEE 35th Symposium on Reliable Distributed Systems (SRDS),True,['failure-management'],True,True,['hardware-failure-prediction'],True,33.0,,,,
174,Predicting and Mitigating Jobs Failures in Big Data Clusters,"['A. Rosà', ' L. Y. Chen', ' W. Binder']",2015,"In large-scale data enters, software and hardware failures are frequent, resulting in failures of job executions that may cause significant resource waste and performance deterioration. To proactively minimize the resource inefficiency due to job failures, it is important to identify them in advance using key job attributes. However, so far, prevailing research on datacenter workload characterization has overlooked job failures, including their patterns, root causes, and impact. In this paper, we aim to develop prediction models and mitigation policies for unsuccessful jobs, so as to reduce the resource waste in big data enters. In particular, we base our analysis on Google cluster traces, consisting of a large number of big-data jobs with a high task fan-out. We first identify the time-varying patterns of failed jobs and the contributing system features. Based on our characterization study, we develop an on-line predictive model for job failures by applying various statistical learning techniques, namely Linear Discriminate Analysis (LDA), Quadratic Discriminate Analysis (QDA), and Logistic Regression (LR). Furthermore, we propose a delay-based mitigation policy which, after a certain grace period, proactively terminates the execution of jobs that are predicted to fail. The particular objective of postponing job terminations is to strike a good tradeoffs between resource waste and false prediction of successful jobs. Our evaluation results show that the proposed method is able to significantly reduce the resource waste by 41.9% on average, and keep false terminations of jobs low, i.e., only 1%.",https://ieeexplore.ieee.org/document/7152488,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 416}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 19}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 908}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('cloud')"", 'index': 47}, {'database': 'IEEE', 'search_string': ""'regression' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 186}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 638}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 1819}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('cloud')"", 'index': 46}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 48}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('remediation' OR 'recovery')"", 'index': 42}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 77}]",27.0,['logistic-regression'],"['kpis', 'host-metrics']","['discussion', 'novel-use']",['failure-prediction'],['job'],"2015 15th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing",True,['failure-management'],,True,['system-failure-prediction'],,,10.1109/CCGrid.2015.139,Conference Paper,,
175,Host load prediction in a Google compute cloud with a Bayesian model,"['S. Di', ' D. Kondo', ' W. Cirne']",2012,"Prediction of host load in Cloud systems is critical for achieving service-level agreements. However, accurate prediction of host load in Clouds is extremely challenging because it fluctuates drastically at small timescales. We design a prediction method based on Bayes model to predict the mean load over a long-term time interval, as well as the mean load in consecutive future time intervals. We identify novel predictive features of host load that capture the expectation, predictability, trends and patterns of host load. We also determine the most effective combinations of these features for prediction. We evaluate our method using a detailed one-month trace of a Google data center with thousands of machines. Experiments show that the Bayes method achieves high accuracy with a mean squared error of 0.0014. Moreover, the Bayes method improves the load prediction accuracy by 5.6 -- 50% compared to other state-of-the-art methods based on moving averages, auto-regression, and/or noise filters.",https://ieeexplore.ieee.org/document/6468464,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 465}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 467}, {'database': 'IEEE', 'search_string': ""'regression' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 354}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 11}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 75}, {'database': 'ACM', 'search_string': ""'regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 86}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 31}]",53.0,,,,['workload-prediction'],,"SC '12: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis",True,['resource-provisioning'],,,,,,,Conference Paper,,
176,New Metrics for Disk Failure Prediction That Go Beyond Prediction Accuracy,"['J. Li', ' R. J. Stones', ' G. Wang', ' Z. Li', ' X. Liu', ' J. Ding']",2018,"Prediction accuracy (true positives, false positives, and so on) is the usual way for evaluating disk-failure prediction models. Realistically however, we aim not only to correctly predict failures, but also to protect data against failure, i.e., we need to take appropriate action after a failure prediction. In the context of storage systems, protecting data requires that we migrate at-risk data, but this consumes network and disk bandwidth, which is particularly problematic for large-scale and cloud systems. This paper consolidates and builds on Li et al. (2016), where we propose using two new metrics, migration rate (MR) and mismigration rate (MMR), to measure the quality of disk failure prediction: MR measures how much at-risk data is migrated (and therefore protected) as a result of correct failure predictions, while MMR measures how much data is migrated needlessly as a result of incorrect failure predictions. In this paper, we additionally propose measuring quality in terms of migration time and mismigration time, which measure the time spent migrating at-risk disks, and the time spent mismigrating healthy disks caused by false alarms, respectively. To demonstrate these metrics' usefulness, we use them to compare disk-failure prediction methods: we compare: 1) a classification tree (CT) model against a state-of-the-art recurrent neural network (RNN) model and 2) a gradient-boosted regression tree (GBRT) model (which predicts residual life) against RNN. We observe that while RNN performs best in the prediction accuracy experiments, the CT and GBRT models sometimes outperform RNN in the resource-dependent migrationrate experiments. We conclude that prediction accuracy is sometimes misleading: correct predictions do not necessarily imply protected data. We additionally present an improved GBRT model (GBRT+), which offers a practical improvement in disk residual-life prediction accordingly to the newly proposed metrics.",https://ieeexplore.ieee.org/document/8552338,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('cloud')"", 'index': 482}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 1344}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 730}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 438}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 301}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 824}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 5}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 1357}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 137}, {'database': 'IEEE', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 397}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 205}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 21}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 22}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 121}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 9}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 1154}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 511}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 355}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 66}]",1.0,"['rnn', 'decision-tree']",['host-metrics'],,['failure-prediction'],['hard-drive'],,True,['failure-management'],,,,,,,,,
177,A Comparison of Predictive Algorithms for Failure Prevention in Smart Environment Applications,"['E. U. Warriach', ' T. Ozcelebi', ' J. J. Lukkien']",2015,"The functional correctness and the performance of smart environment applications can be hampered by faults. Fault tolerance solutions aim to achieve graceful performance degradation in the presence of faults, ideally without leading to application failures. This is a reactive approach and, by itself, gives little flexibility and time for preventing potential failures. We argue that the key step in achieving high dependability is to predict faults before they occur. We propose a proactive fault prevention framework, which predicts potential low-level hardware, software and network faults and tries to prevent them via dynamic adaptation. Many statistical fault prediction algorithms have been proposed in the literature. In this paper, we evaluate and compare the performances of two fault prediction models, namely, multiple linear regression, and artificial neural networks by using them to predict the remaining useful life of a battery-powered wireless sensor network node. The results show that the proposed framework will provide better control over performance degradation of smart environment applications, and will increase reliability and availability, and reduce manual user interventions.",https://ieeexplore.ieee.org/document/7194268,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prevention' OR 'failure prevention')"", 'index': 1}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault prevention' OR 'failure prevention')"", 'index': 1}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 60}, {'database': 'IEEE', 'search_string': ""'regression' AND ('remediation' OR 'recovery')"", 'index': 206}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('remediation' OR 'recovery')"", 'index': 568}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 135}]",1.0,"['multilayer-perceptron', 'linear-regression']",['host-metrics'],['comparison'],"['failure-prediction', 'failure-prevention']",['battery'],2015 International Conference on Intelligent Environments,True,['failure-management'],,True,['hardware-failure-prediction'],,,,,,
178,LogTree: A Framework for Generating System Events from Raw Textual Logs,"['L. Tang', ' T. Li']",2010,"Modern computing systems are instrumented to generate huge amounts of system logs and these data can be utilized for understanding and complex system behaviors. One main fundamental challenge in automated log analysis is the generation of system events from raw textual logs. Recent works apply clustering techniques to translate the raw log messages into system events using only the word/term information. In this paper, we first illustrate the drawbacks of existing techniques for event generation from system logs. We then propose Log Tree, a novel and algorithm-independent framework for events generation from raw system log messages. Log Tree utilizes the format and structural information of the raw logs in the clustering process to generate system events with better accuracy. In addition, an indexing data structure, Message Segment Table, is proposed in Log Tree to significantly improve the efficiency of events creation. Extensive experiments on real system logs demonstrate the effectiveness and efficiency of Log Tree.",https://ieeexplore.ieee.org/document/5694003,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 15}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 23}]",27.0,['clustering'],['logs'],['new-method'],['failure-detection'],['log'],2010 IEEE International Conference on Data Mining,True,['failure-management'],,,['log-enhancement'],,,,,,
179,Log Clustering Based Problem Identification for Online Service Systems,"['Q. Lin', ' H. Zhang', ' J. Lou', ' Y. Zhang', ' X. Chen']",2016,"Logs play an important role in the maintenance of large-scale online service systems. When an online service fails, engineers need to examine recorded logs to gain insights into the failure and identify the potential problems. Traditionally, engineers perform simple keyword search (such as “error” and “exception”) of logs that may be associated with the failures. Such an approach is often time consuming and error prone. Through our collaboration with Microsoft service product teams, we propose LogCluster, an approach that clusters the logs to ease log-based problem identification. LogCluster also utilizes a knowledge base to check if the log sequences occurred before. Engineers only need to examine a small number of previously unseen, representative log sequences extracted from the clusters to identify a problem, thus significantly reducing the number of logs that should be examined, meanwhile improving the identification accuracy. Through experiments on two Hadoop-based applications and two large-scale Microsoft online service systems, we show that our approach is effective and outperforms the state-of-the-art work proposed by Shang et al. in ICSE 2013. We have successfully applied LogCluster to the maintenance of many actual Microsoft online service systems. In this paper, we also share our success stories and lessons learned.",https://ieeexplore.ieee.org/document/7883294,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 26}, {'database': 'ACM', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 61}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault localization' OR 'failure localization')"", 'index': 72}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault localization' OR 'failure localization')"", 'index': 58}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 65}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 36}, {'database': 'ACM', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 100}]",81.0,['clustering'],['logs'],,['root-cause-analysis'],,2016 IEEE/ACM 38th International Conference on Software Engineering Companion (ICSE-C),True,['failure-management'],True,True,['rca-others'],True,86.0,10.1145/2889160.2889232,Conference Paper,"['problem identification', 'online service system', 'diagnosis', 'logs', 'log clustering']",
180,A data clustering algorithm for mining patterns from event logs,['R. Vaarandi'],2003,"Today, event logs contain vast amounts of data that can easily overwhelm a human. Therefore, mining patterns from event logs is an important system management task. The paper presents a novel clustering algorithm for log file data sets which helps one to detect frequent patterns from log files, to build log file profiles, and to identify anomalous log file lines.",https://ieeexplore.ieee.org/document/1251233,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 51}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 156}]",336.0,['clustering'],['logs'],['novel-use'],['failure-detection'],['application'],Proceedings of the 3rd IEEE Workshop on IP Operations & Management (IPOM 2003) (IEEE Cat. No.03EX764),True,['failure-management'],,,['anomaly-detection'],,,,,,
181,Digging deeper into cluster system logs for failure prediction and root cause diagnosis,"['X. Fu', ' R. Ren', ' S. A. McKee', ' J. Zhan', ' N. Sun']",2014,"As the sizes of supercomputers and data centers grow towards exascale, failures become normal. System logs play a critical role in the increasingly complex tasks of automatic failure prediction and diagnosis. Many methods for failure prediction are based on analyzing event logs for large scale systems, but there is still neither a widely used one to predict failures based on both non-fatal and fatal events, nor a precise one that uses fine-grained information (such as failure type, node location, related application, and time of occurrence). A deeper and more precise log analysis technique is needed. We propose a three-step approach to draw out event dependencies and to identify failure-event generating processes. First, we cluster frequent event sequences into event groups based on common events. Then we infer causal dependencies between events in each event group. Finally, we extract failure rules based on the observation that events of the same event types, on the same nodes or from the same applications have similar operational behaviors. We use this rich information to improve failure prediction. Our approach semi-automates diagnosing the root causes of failure events, making it a valuable tool for system administrators.",https://ieeexplore.ieee.org/document/6968768,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 166}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prediction' OR 'failure prediction')"", 'index': 18}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 39}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 36}]",23.0,"['correlation', 'graph-mining', 'clustering']",['logs'],['new-method'],"['failure-prediction', 'root-cause-analysis']",['cluster'],2014 IEEE International Conference on Cluster Computing (CLUSTER),True,['failure-management'],,,"['system-failure-prediction', 'root-cause-diagnosis']",,,,,,
182,LOGAIDER: A Tool for Mining Potential Correlations of HPC Log Events,"['S. Di', ' R. Gupta', ' M. Snir', ' E. Pershey', ' F. Cappello']",2017,"Today's large-scale supercomputers are producing a huge amount of log data. Exploring various potential correlations of fatal events is crucial for understanding their causality and improving the working efficiency for system administrators. To this end, we developed a toolkit, named LogAider, that can reveal three types of potential correlations: across-field, spatial, and temporal. Across-field correlation refers to the statistical correlation across fields within a log or across multiple logs based on probabilistic analysis. For analyzing the spatial correlation of events, we developed a generic, easy-to-use visualizer that can view any events queried by userson a system machine graph. LogAider can also mine spatial correlations by an optimized K-meaning clustering algorithm over a Torus network topology. It is also able to disclose the temporal correlations (or error propagations) over a certain period inside a log or across multiple logs, based on an effective similarity analysis strategy. We assessed LogAider using theone-year reliability-availability-serviceability (RAS) log of Mira system (one of the world's most powerful supercomputers), as well as its job log. We find that LogAider very helpful for revealing the potential correlations of fatal system events and job events, with an accurate mining of across-field correlation with both precision and recall of 99.9-100%, as well as precisedetection of temporal-correlation with a high similarity (up to 95%) to the ground-truth.",https://ieeexplore.ieee.org/document/7973730,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 256}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 1667}]",14.0,['clustering'],['logs'],,['root-cause-analysis'],,"2017 17th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID)",True,['failure-management'],,,['rca-others'],,,,,,
183,Spatio-temporal factorization of log data for understanding network events,"['T. Kimura', ' K. Ishibashi', ' T. Mori', ' H. Sawada', ' T. Toyono', ' K. Nishimatsu', ' A. Watanabe', ' A. Shimoda', ' K. Shiomoto']",2014,"Understanding the impacts and patterns of network events such as link flaps or hardware errors is crucial for diagnosing network anomalies. In large production networks, analyzing the log messages that record network events has become a challenging task due to the following two reasons. First, the log messages are composed of unstructured text messages generated by vendor-specific rules. Second, network equipment such as routers, switches, and RADIUS severs generate various log messages induced by network events that span across several geographical locations, network layers, protocols, and services. In this paper, we have tackled these obstacles by building two novel techniques: statistical template extraction (STE) and log tensor factorization (LTF). STE leverages a statistical clustering technique to automatically extract primary templates from unstructured log messages. LTF aims to build a statistical model that captures spatial-temporal patterns of log messages. Such spatial-temporal patterns provide useful insights into understanding the impacts and root cause of hidden network events. This paper first formulates our problem in a mathematical way. We then validate our techniques using massive amount of network log messages collected from a large operating network. We also demonstrate several case studies that validate the usefulness of our technique.",https://ieeexplore.ieee.org/document/6847986,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 386}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 355}]",53.0,"['dimensionality-reduction', 'clustering']",['logs'],['new-method'],['root-cause-analysis'],"['switch', 'network', 'router']",IEEE INFOCOM 2014 - IEEE Conference on Computer Communications,True,['failure-management'],,,"['root-cause-diagnosis', 'fault-localization']",,,,,,
184,Diagnosing the root-causes of failures from cluster log files,"['E. Chuah', ' S. Kuo', ' P. Hiew', ' W. Tjhi', ' G. Lee', ' J. Hammond', ' M. T. Michalewicz', ' T. Hung', ' J. C. Browne']",2010,"System event logs are often the primary source of information for diagnosing (and predicting) the causes of failures for cluster systems. Due to interactions among the system hardware and software components, the system event logs for large cluster systems are comprised of streams of interleaved events, and only a small fraction of the events over a small time span are relevant to the diagnosis of a given failure. Furthermore, the process of troubleshooting the causes of failures is largely manual and ad-hoc. In this paper, we present a systematic methodology for reconstructing event order and establishing correlations among events which indicate the root-causes of a given failure from very large syslogs. We developed a diagnostics tool, FDiag, to extract the log entries as structured message templates and uses statistical correlation analysis to establish probable cause and effect relationships for the fault being analyzed. We applied FDiag to analyze failures due to breakdowns in interactions between the Lustre file system and its clients on the Ranger supercomputer at the Texas Advanced Computing Center (TACC). The results are positive. FDiag is able to identify the dates and the time periods that contain the significant events which eventually led to the occurrence of compute node soft lockups.",https://ieeexplore.ieee.org/document/5713159,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 417}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('remediation' OR 'recovery')"", 'index': 520}]",34.0,['correlation'],"['kpis', 'host-metrics', 'logs']",,['root-cause-analysis'],['cluster'],2010 International Conference on High Performance Computing,True,['failure-management'],,,['root-cause-diagnosis'],,,,,,
185,LogMaster: Mining Event Correlations in Logs of Large-Scale Cluster Systems,"['X. Fu', ' R. Ren', ' J. Zhan', ' W. Zhou', ' Z. Jia', ' G. Lu']",2012,"This paper presents a set of innovative algorithms and a system, named Log Master, for mining correlations of events that have multiple attributions, i.e., node ID, application ID, event type, and event severity, in logs of large-scale cloud and HPC systems. Different from traditional transactional data, e.g., supermarket purchases, system logs have their unique characteristics, and hence we propose several innovative approaches to mining their correlations. We parse logs into an n-ary sequence where each event is identified by an informative nine-tuple. We propose a set of enhanced apriori-like algorithms for improving sequence mining efficiency, we propose an innovative abstraction-event correlation graphs (ECGs) to represent event correlations, and present an ECGs-based algorithm for fast predicting events. The experimental results on three logs of production cloud and HPC systems, varying from 433490 entries to 4747963 entries, show that our method can predict failures with a high precision and an acceptable recall rates.",https://ieeexplore.ieee.org/document/6424841,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 436}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 1717}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 67}]",63.0,"['rule-mining', 'graph-mining']","['events', 'logs']",['new-method'],['failure-prediction'],"['hpc', 'cloud']",2012 IEEE 31st Symposium on Reliable Distributed Systems,True,['failure-management'],True,,['system-failure-prediction'],,,,,,
186,Real-Time Anomaly Detection in Streams of Execution Traces,"['W. Zhang', ' F. Bastani', ' I. Yen', ' K. Hulin', ' F. Bastani', ' L. Khan']",2012,"For deployed systems, software fault detection can be challenging. Generally, faulty behaviors are detected based on execution logs, which may contain a large volume of execution traces, making analysis extremely difficult. This paper investigates and compares the effectiveness and efficiency of various data mining techniques for software fault detection based on execution logs, including clustering based, density based, and probabilistic automata based methods. However, some existing algorithms suffer from high complexity and do not scale well to large datasets. To address this problem, we present a suite of prefix tree based anomaly detection techniques. The prefix tree model serves as a compact loss less data representation of execution traces. Also, the prefix tree distance metric provides an effective heuristic to guide the search for execution traces having close proximity to each other. In the density based algorithm, the prefix tree distance is used to confine the K-nearest neighbor search to a small subset of the nodes, which greatly reduces the computing time without sacrificing accuracy. Experimental studies show a significant speedup in our prefix tree based and prefix tree distance guided approaches, from days to minutes in the best cases, in automated identification of software failures.",https://ieeexplore.ieee.org/document/6375634,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 504}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 605}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 328}, {'database': 'IEEE', 'search_string': ""'classification' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 288}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 258}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 102}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 402}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault detection' OR 'failure detection')"", 'index': 1493}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault detection' OR 'failure detection')"", 'index': 135}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 934}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1101}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 35}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 148}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 149}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 287}]",8.0,"['automaton', 'similarity-matching', 'clustering']",['traces'],"['comparison', 'novel-use']",['failure-detection'],,2012 IEEE 14th International Symposium on High-Assurance Systems Engineering,True,['failure-management'],,True,['anomaly-detection'],,,,,,
187,Kahuna: Problem diagnosis for Mapreduce-based cloud computing environments,"['Jiaqi Tan', ' Xinghao Pan', ' E. Marinelli', ' S. Kavulya', ' R. Gandhi', ' P. Narasimhan']",2010,"We present Kahuna, an approach that aims to diagnose performance problems in MapReduce systems. Central to Kahuna's approach is our insight on peer-similarity, that nodes behave alike in the absence of performance problems, and that a node that behaves differently is the likely culprit of a performance problem. We present applications of Kahuna's insight in techniques and their algorithms to statistically compare black-box (OS-level performance metrics) and white-box (Hadoop-log statistics) data across the different nodes of a MapReduce cluster, in order to identify the faulty node(s). We also present empirical evidence of our peer-similarity observations from the 4000-processor Yahoo! M45 Hadoop cluster. In addition, we demonstrate Kahuna's effectiveness through experimental evaluation of two algorithms for a number of reported performance problems, on four different workloads in a 100-node Hadoop cluster running on Amazon's EC2 infrastructure.",https://ieeexplore.ieee.org/document/5488446,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 978}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1346}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 913}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 1772}]",77.0,['clustering'],"['host-metrics', 'logs']",,['failure-detection'],['mapreduce'],2010 IEEE Network Operations and Management Symposium - NOMS 2010,True,['failure-management'],,,['anomaly-detection'],,,,,,
188,Bad Words: Finding Faults in Spirit's Syslogs,"['J. Stearley', ' A. J. Oliner']",2008,"Accurate fault detection is a key element of resilient computing. Syslogs provide key information regarding faults, and are found on nearly all computing systems. Discovering new fault types requires expert human effort, however, as no previous algorithm has been shown to localize faults in time and space with an operationally acceptable false positive rate. We present experiments on three weeks of syslogs from Sandia's 512-node ""Spirit"" Linux cluster, showing one algorithm that localizes 50% of faults with 75% precision, corresponding to an excellent false positive rate of 0.05%. The salient characteristics of this algorithm are (1) calculation of nodewise information entropy, and (2) encoding of word position. The key observation is that similar computers correctly executing similar work should produce similar logs.",https://ieeexplore.ieee.org/document/4534301,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1151}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 355}]",39.0,['entropy-selection'],['logs'],['new-method'],['failure-detection'],,2008 Eighth IEEE International Symposium on Cluster Computing and the Grid (CCGRID),True,['failure-management'],,True,['anomaly-detection'],,,,,,
189,"Detecting traffic anomaly in wireless networks, an analytics methodology","['Z. Li', ' Y. Ouyang', ' L. Su', ' W. Jiang', ' Y. Hu', ' Z. Lin']",2018,"Since anomalies in wireless networks have different behaviors, not all of them can be detected and recognized. With the capacity of detecting and describing hidden structure from unlabeled data, unsupervised algorithms are able to automatically characterize the nature of traffic behavior and detect anomalies from normal behaviors in the wireless network. In this paper, co-occurrence data is studied since it combines traffic data with generating entities. Gaussian probabilistic latent semantic analysis (GPLSA) model is leveraged to compare the Gaussian Mixture Model (GMM) with temporal network data. A novel “Donut” algorithm of anomaly detection is proposed with model log-likelihood. Experimental results validated that the proposed GPLSA model could hold better promise in the early detection with low false alarm rate and low implementation complexity.",https://ieeexplore.ieee.org/document/8363936,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1154}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 864}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 836}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 237}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 381}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1483}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('anomaly detection' OR 'outlier detection')"", 'index': 168}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 259}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 431}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 504}]",1.0,,,,['failure-detection'],['network'],2018 Wireless Telecommunications Symposium (WTS),True,['failure-management'],,True,"['anomaly-detection', 'traffic-classification']",,,,,,
190,An Empirical Evaluation of Deep Learning for Network Anomaly Detection,"['R. K. Malaiya', ' D. Kwon', ' S. C. Suh', ' H. Kim', ' I. Kim', ' J. Kim']",2019,"Deep learning has been widely studied in many technical domains such as image analysis and speech recognition, with its benefits that effectively deal with complex and high-dimensional data. Our preliminary experiments show a high degree of non-linearity from the network connection data, which explains why it is hard to improve the performance of identifying network anomalies by using conventional learning methods (e.g., Adaboosting, SVM, and Random Forest). In this study, we design and examine deep learning models constructed based on Fully Connected Networks (FCNs), Variational AutoEncoder (VAE), and Sequence-to-Sequence (Seq2Seq) structures. For the extensive evaluation, we employ a broad range of the public datasets with unique characteristics. Our experimental results confirm the feasibility of deep learning-based network anomaly detection, with the improved performance compared to the conventional learning techniques. In particular, the detection model based on Seq2Seq with LSTM is highly promising, consistently yielding over 99% of accuracy to identify network anomalies from the entire datasets employed in the evaluation.",https://ieeexplore.ieee.org/document/8846674,True,"[{'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 10}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 47}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 437}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('anomaly detection' OR 'outlier detection')"", 'index': 238}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('anomaly detection' OR 'outlier detection')"", 'index': 283}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 178}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 216}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 431}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 304}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 485}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 95}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 266}]",10.0,"['multilayer-perceptron', 'rnn', 'autoencoder']",,"['comparison', 'novel-use']",['failure-detection'],['network'],"2018 International Conference on Computing, Networking and Communications (ICNC)",True,['failure-management'],False,True,['anomaly-detection'],,,,,,
191,Anomaly Detection of System Logs Based on Natural Language Processing and Deep Learning,"['M. Wang', ' L. Xu', ' L. Guo']",2018,"System logs record the execution trajectory of the system and exist in all components of the system. Nowadays, the systems are deployed in a distributed environment and they generate logs which contain complex format and rich semantic information. Simple statistical analysis methods cannot fully capture log information for effective abnormal detection of software systems. In this paper, we propose to analyze the logs by combining feature extraction methods from natural language processing and anomaly detection methods from deep learning. Two feature extraction algorithms, Word2vec and Term Frequency-Inverse Document Frequency (TF-IDF), are respectively adopted and compared here to obtain the log information, and then one deep learning method named Long Short-Term Memory (LSTM) is applied for the anomaly detection. To validate the effectiveness of the proposed method, we compare LSTM with other machine learning algorithms, including Gradient Boosting Decision Tree (GBDT) and Naïve Bayes, the results show that LSTM can perform the best for anomaly detection of system logs with both of the two feature extraction methods, indicating that LSTM can capture contextual semantic information effectively in log anomaly detection and will be a promising tool for log analysis.",https://ieeexplore.ieee.org/document/8552075,True,"[{'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 20}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 20}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 180}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 3}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 22}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 28}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 44}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 58}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 64}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 175}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 154}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 44}]",0.0,['rnn'],['logs'],,['failure-detection'],,2018 4th International Conference on Frontiers of Signal Processing (ICFSP),True,['failure-management'],,True,['anomaly-detection'],,,,,,
192,Root Cause Detection using Dynamic Dependency Graphs from Time Series Data,"['S. Y. Shah', ' X. Dang', ' P. Zerfos']",2018,"Change detection in system behavior and its root cause detection is essential for many large-scale systems such as, manufacturing plants, in order to keep systems running uninterrupted and avoid costly machine breakdown via predictive maintenance. In this paper, we present a novel graph based technique that uses time variant interdependencies and lagged dependencies among different components of a system to detect changes in the system behavior. We further find the root causes for these detected changes by pointing out the component and its historical values that are responsible for initiating and changing the system to the new state. The proposed mechanism extracts these time variant dependencies using a deep learning system and converts them into weighted directed graphs and applies graph based techniques for change detection. For each detected change, our system uses graph theoretic techniques to uncover the root causes for the change. Such a mechanism provides us with valuable insights about the inner workings of a system from a different perspective as opposed to traditional techniques for root cause analysis that directly apply statistical models to the time series data for analysis. Experimental results on real manufacturing data show, that we can detect changes in system behavior and accurately identify the root causes in almost 71% of the cases for which we have the ground truth. For synthetic data, our system can correctly identify root causes in 87% of the cases.",https://ieeexplore.ieee.org/document/8622059,True,"[{'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 181}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 807}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 420}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 2}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 21}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 971}]",1.0,"['rnn', 'graph-mining']",['events'],,"['root-cause-analysis', 'failure-detection']",,2018 IEEE International Conference on Big Data (Big Data),True,['failure-management'],,,"['anomaly-detection', 'root-cause-diagnosis', 'fault-localization']",,,,,,
193,Fault detection for cloud computing systems with correlation analysis,"['T. Wang', ' W. Zhang', ' J. Wei', ' H. Zhong']",2015,"The large-scale dynamic cloud computing environment has raised great challenges for fault diagnosis in Web applications. First, fluctuating workloads cause traditional application models to change over time. Moreover, modeling the behaviors of complex applications always requires domain knowledge which is difficult to obtain. Finally, managing large-scale applications manually is impractical for operators. This paper addresses these issues and proposes an automatic fault diagnosis method for Web applications in cloud computing. We propose an online incremental clustering method to recognize access behavior patterns, and uses CCA to model the correlation between workloads and the metrics of application performance/resource utilization in a specific access behavior pattern. Our method detects anomalies by discovering the abrupt change of correlation coefficients with a EWMA control chart, and then locates suspicious metrics using a feature selection method combining ReliefF and SVM-RFE. We validate our method by injecting typical faults in TPC-W an industry-standard benchmark, and the experimental results demonstrate that it can effectively detect typical faults.",https://ieeexplore.ieee.org/document/7140351,True,"[{'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault detection' OR 'failure detection')"", 'index': 333}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1133}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 404}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 1109}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 180}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 319}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('anomaly detection' OR 'outlier detection')"", 'index': 440}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 127}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 488}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 1020}]",9.0,,,,['failure-detection'],,2015 IFIP/IEEE International Symposium on Integrated Network Management (IM),True,['failure-management'],,True,,,,,,,
194,Proactive drive failure prediction for large scale storage systems,"['B. Zhu', ' G. Wang', ' X. Liu', ' D. Hu', ' S. Lin', ' J. Ma']",2013,"Most of the modern hard disk drives support Self-Monitoring, Analysis and Reporting Technology (SMART), which can monitor internal attributes of individual drives and predict impending drive failures by a thresholding method. As the prediction performance of the thresholding algorithm is disappointing, some researchers explored various statistical and machine learning methods for predicting drive failures based on SMART attributes. However, the failure detection rates of these methods are only up to 50% ~ 60% with low false alarm rates (FARs). We explore the ability of Backpropagation (BP) neural network model to predict drive failures based on SMART attributes. We also develop an improved Support Vector Machine (SVM) model. A real-world dataset concerning 23,395 drives is used to verify these models. Experimental results show that the prediction accuracy of both models is far higher than previous works. Although the SVM model achieves the lowest FAR (0.03%), the BP neural network model is considerably better in failure detection rate which is up to 95% while keeping a reasonable low FAR.",https://ieeexplore.ieee.org/document/6558427,True,"[{'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault detection' OR 'failure detection')"", 'index': 340}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 510}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('remediation' OR 'recovery')"", 'index': 94}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 1358}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 118}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 43}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('remediation' OR 'recovery')"", 'index': 493}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 433}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 106}]",77.0,['multilayer-perceptron'],['host-metrics'],['novel-use'],['failure-prediction'],['hard-drive'],2013 IEEE 29th Symposium on Mass Storage Systems and Technologies (MSST),True,['failure-management'],True,,['hardware-failure-prediction'],True,30.0,,,,
195,Hidden Markov Model for hard-drive failure detection,"['T. Teoh', ' S. Cho', ' Y. Nguwi']",2012,"This paper illustrates the use of Hidden Markov Model (HMM) to model hard disk failure. The reason we use HMM is because HMM is a formal foundation for making probabilistic models of linear sequence `labeling' problem. We use the database provided by University of California, San Diego for detection of hard-drive failure. We have selected 24 attributes and obtain accuracy of about 90%. We compare machine-learning methods applied to a difficult real-world problem: predicting computer hard-drive failure using attributes monitored internally by individual drives. The problem is one of detecting rare events in a time series of noisy and non-parametrically distributed data. We develop a new algorithm HMM which is specifically designed for the low false-alarm case, and is shown to have promising performance. Other methods compared are support vector machines (SVMs), unsupervised clustering, and non-parametric statistical tests (rank-sum and reverse arrangements). The failure-prediction performance of the SVM, rank-sum and mi-NB algorithm is considerably better than the threshold method currently implemented in drives, while maintaining low false alarm rates [13]. Our results suggest that non-parametric statistical tests should be considered for learning problems involving detecting rare events.",https://ieeexplore.ieee.org/document/6295014,True,"[{'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault detection' OR 'failure detection')"", 'index': 352}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 525}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 271}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 265}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 134}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 11}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 205}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 77}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 80}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 266}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 50}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 25}]",1.0,['markov-model'],['host-metrics'],,['failure-prediction'],['hard-drive'],2012 7th International Conference on Computer Science & Education (ICCSE),True,['failure-management'],,,,,,,,,
196,A Two-Step Parametric Method for Failure Prediction in Hard Disk Drives,"['Y. Wang', ' E. W. M. Ma', ' T. W. S. Chow', ' K. Tsui']",2014,"Predicting the impending failure of hard disk drives (HDDs) is crucial for preventing essential data from losing. In this paper, a two-step parametric method was developed to predict the impending failure of HDDs using the aggregate of statistical models. This method deals with the problem of failure prediction in two steps: anomaly detection and failure prediction. First, Mahalanobis distance was used for aggregating all the monitored variables into one index, which was then transformed into Gaussian variables by Box-Cox transformation. By defining an appropriate threshold, anomalies in HDDs were detected as a result. Second, a sliding-window-based generalized likelihood ratio test was proposed to track the anomaly progression in an HDD. When the occurrence of anomalies in a time interval is found to be statistically significant, indicating the HDD is approaching failure. In this work, we also derived a new cost function to adjust the prediction rate. This is important in a way to balance the failure detection rate and false alarm rate as well as to provide an advanced warning of HDD failures to the users, whereby the users can back up their data in time. Then the developed method was applied on a synthetic data set showing its effectiveness on predicting failures. To demonstrate the practical usefulness, this method was also applied on a real-life HDD data set. The result shows that our method could achieve 68% failure detection rate with 0% false alarm rate. This is much better than the results achieved by the state-of-the-art methods, such as support vector machine and hidden Markov models.",https://ieeexplore.ieee.org/document/6517515,True,"[{'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault detection' OR 'failure detection')"", 'index': 452}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('anomaly detection' OR 'outlier detection')"", 'index': 407}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 49}]",60.0,"['similarity-matching', 'correlation', 'optimization', 'genetic-programming']",['host-metrics'],"['new-method', 'comparison']","['failure-detection', 'failure-prediction']",['hard-drive'],IEEE Transactions on Industrial Informatics,True,['failure-management'],False,,"['anomaly-detection', 'hardware-failure-prediction']",,,,,,
197,Data mining approaches to software fault diagnosis,"['R. P. J. C. Bose', ' S. H. Srinivasan']",2005,Automatic identification of software faults has enormous practical significance. This requires characterizing program execution behavior and the use of appropriate data mining techniques on the chosen representation. In this paper we use the sequence of system calls to characterize program execution. The data mining tasks addressed are learning to map system call streams to fault labels and automatic identification of fault causes. Spectrum kernels and SVM are used for the former while latent semantic analysis is used for the latter The techniques are demonstrated for the intrusion dataset containing system call traces. The results show that kernel techniques are as accurate as the best available results but are faster by orders of magnitude. We also show that latent semantic indexing is capable of revealing fault-specific features.,https://ieeexplore.ieee.org/document/1498230,True,"[{'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault detection' OR 'failure detection')"", 'index': 459}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 177}]",7.0,,,,['failure-detection'],,15th International Workshop on Research Issues in Data Engineering: Stream Data Mining and Applications (RIDE-SDMA'05),True,['failure-management'],,,,,,,,,
198,Automatic and Generic Periodicity Adaptation for KPI Anomaly Detection,"['N. Zhao', ' J. Zhu', ' Y. Wang', ' M. Ma', ' W. Zhang', ' D. Liu', ' M. Zhang', ' D. Pei']",2019,"Key performance indicator (KPI) anomaly detection (AD) is critical to ensure service quality and reliability. Due to the effects of work days, off days, festivals, and business activities on user behavior, KPIs may exhibit different patterns within different days, which we call periodicity profiles of KPIs. However, existing KPI AD approaches have difficulties in adapting to diverse periodicity profiles due to the lack of generality. In this paper, we propose an automatic and generic framework called Period, which can accurately detect the periodicity profiles through daily subsequences clustering, and improve the performance of AD methods by robustly and automatically adapting to different periodicity profiles. In our evaluation using several real-world KPIs with different periodicity profiles from large Internet-based services, the clustering algorithm used to detect periodicity can achieve about 0.95 accuracy on average. More importantly, further evaluation on 56 KPIs shows that Period can significantly improve the best F-score of several widely used AD approaches by up to 0.66.",https://ieeexplore.ieee.org/document/8723601,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 172}, {'database': 'IEEE', 'search_string': ""'AIOps'"", 'index': 4}, {'database': 'IEEE', 'search_string': ""'clustering' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 29}]",0.0,['clustering'],['kpis'],['new-method'],['failure-detection'],,IEEE Transactions on Network and Service Management,True,['failure-management'],,,,,,,,,
199,Robust and Rapid Clustering of KPIs for Large-Scale Anomaly Detection,"['Z. Li', ' Y. Zhao', ' R. Liu', ' D. Pei']",2018,"For large Internet companies, it is very important to monitor a large number of KPIs (Key Performance Indicators) and detect anomalies to ensure the service quality and reliability. However, large-scale anomaly detection on millions of KPIs is very challenging due to the large overhead of model selection, parameter tuning, model training, or labeling. In this paper we argue that KPI clustering can help: we can cluster millions of KPIs into a small number of clusters and then select and train model on a per-cluster basis. However, KPI clustering faces new challenges that are not present in classic time series clustering: KPIs are typically much longer than other time series, and noises, anomalies, phase shifts and amplitude differences often change the shape of KPIs and mislead the clustering algorithm. To tackle the above challenges, in this paper we propose a robust and rapid KPI clustering algorithm, ROCKA. It consists of four steps: preprocessing, baseline extraction, clustering and assignment. These techniques help group KPIs according to their underlying shapes with high accuracy and efficiency. Our evaluation using real-world KPIs shows that ROCKA gets F-score higher than 0.85, and reduces model training time of a state-of-the-art anomaly detection algorithm by 90%, with only 15% performance loss.",https://ieeexplore.ieee.org/document/8624168,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 16}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 140}]",2.0,,,,['failure-detection'],,2018 IEEE/ACM 26th International Symposium on Quality of Service (IWQoS),True,['failure-management'],,,,,,,,,
200,Rapid Deployment of Anomaly Detection Models for Large Number of Emerging KPI Streams,"['J. Bu', ' Y. Liu', ' S. Zhang', ' W. Meng', ' Q. Liu', ' X. Zhu', ' D. Pei']",2018,"Internet-based services monitor and detect anomalies on KPIs (Key Performance Indicators, say CPU utilization, number of queries per second, response latency) of their applications and systems in order to keep their services reliable. This paper identifies a common, important, yet little-studied problem of KPI anomaly detection: rapid deployment of anomaly detection models for large number of emerging KPI streams, without manual algorithm selection, parameter tuning, or new anomaly labeling for any newly emerging KPI streams. We propose the first framework ADS (Anomaly Detection through Self-training) that tackles the above problem, via clustering and semi-supervised learning. Our extensive experiments using real-world data show that, with the labels of only the 5 cluster centroids of 70 historical KPI streams, ADS achieves an averaged best F-score of 0.92 on 81 new KPI streams, almost the same as a state-of-art supervised approach, and greatly outperforming a state-of-art unsupervised approach by 61.40% on average.",https://ieeexplore.ieee.org/document/8711315,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 34}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 28}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 433}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 33}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 257}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 243}]",4.0,,,,['failure-detection'],,2018 IEEE 37th International Performance Computing and Communications Conference (IPCCC),True,['failure-management'],,,,,,,,,
201,A model for early prediction of faults in software systems,"['P. S. Sandhu', ' R. Goel', ' A. S. Brar', ' J. Kaur', ' S. Anand']",2010,"Quality of a software component can be measured in terms of fault proneness of data. Quality estimations are made using fault proneness data available from previously developed similar type of projects and the training data consisting of software measurements. To predict faulty modules in software data different techniques have been proposed which includes statistical method, machine learning methods, neural network techniques and clustering techniques. Predicting faults early in the software life cycle can be used to improve software process control and achieve high software reliability. The aim of proposed approach is to investigate that whether metrics available in the early lifecycle (i.e. requirement metrics), metrics available in the late lifecycle (i.e. code metrics) and metrics available in the early lifecycle (i.e. requirement metrics) combined with metrics available in the late lifecycle (i.e. code metrics) can be used to identify fault prone modules using decision tree based Model in combination of K-means clustering as preprocessing technique. This approach has been tested with CM1 real time defect datasets of NASA software projects. The high accuracy of testing results show that the proposed Model can be used for the prediction of the fault proneness of software modules early in the software life cycle.",https://ieeexplore.ieee.org/document/5451695,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 65}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 181}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 170}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 170}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 238}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 72}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 164}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 40}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 301}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 156}]",10.0,,,,['failure-prediction'],,2010 The 2nd International Conference on Computer and Automation Engineering (ICCAE),True,['failure-management'],,True,,,,,,,
202,PerfInsight: A Robust Clustering-Based Abnormal Behavior Detection System for Large-Scale Cloud,"['X. Zhang', ' F. Meng', ' J. Xu']",2018,"Anomalous behaviors of cloud services usually lead to performance degradation or even unplanned outages, which dramatically harms their Quality of Services. Performance monitoring and anomaly detection systems have been widely applied to mitigate these risks. However, huge volume of collected data, prevalence of trends and noises in data distribution, lack of labelled anomalies and unpredictability of various types of anomalies bring great challenges to existing anomaly detection systems in real world. Recently, unsupervised clustering-based anomaly detection approaches become promising solutions due to less dependency on labelled data and adaption to various types of anomalies. To achieve better quality with clustering-based anomaly detection approaches, huge amount of data normalization work is required. In this paper, we present a practical robust anomaly detection system for large-scale cloud called PerfInsight. First, it detects potential trends from these collected data and automatically transforms them to reduce their negative impact to clustering results. Then, an entropy-based feature selection of transformed metrics is designed to improve the detection efficiency. Finally, more robust clustering models can be trained and used based on these well transformed and selected features. Our evaluation results prove that PerfInsight could significantly reduce the cardinality of models.",https://ieeexplore.ieee.org/document/8457898,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 131}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 14}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 74}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 1}]",2.0,"['entropy-selection', 'clustering']",['host-metrics'],,['failure-detection'],,2018 IEEE 11th International Conference on Cloud Computing (CLOUD),True,['failure-management'],,,['anomaly-detection'],,,,,,
203,Class level fault prediction using software clustering,"['G. Scanniello', ' C. Gravino', ' A. Marcus', ' T. Menzies']",2013,"Defect prediction approaches use software metrics and fault data to learn which software properties associate with faults in classes. Existing techniques predict fault-prone classes in the same release (intra) or in a subsequent releases (inter) of a subject software system. We propose an intra-release fault prediction technique, which learns from clusters of related classes, rather than from the entire system. Classes are clustered using structural information and fault prediction models are built using the properties of the classes in each cluster. We present an empirical investigation on data from 29 releases of eight open source software systems from the PROMISE repository, with predictors built using multivariate linear regression. The results indicate that the prediction models built on clusters outperform those built on all the classes of the system.",https://ieeexplore.ieee.org/document/6693126,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 154}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 323}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 17}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 11}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 13}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 7}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 2}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prediction' OR 'failure prediction')"", 'index': 4}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 2}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 8}]",6.0,,,,['failure-prediction'],,2013 28th IEEE/ACM International Conference on Automated Software Engineering (ASE),True,['failure-management'],,,,,,10.1109/ASE.2013.6693126,Conference Paper,"['software clustering', 'empirical study', 'fault prediction']",
204,A Density-Based Anomaly Detection Method for MapReduce,"['K. Wang', ' Y. Wang', ' B. Yin']",2012,"Cloud computing has been more and more popular and widely used as a new model of information technology. In order to achieve a reliable and efficient operation of the cloud environment, it is important for cloud providers to detect and deal with system anomalies in time. In this paper, we present a method for anomaly detection in MapReduce environment. This method is based on peer-similarity and uses density based clustering on OS-level metrics to perform real time analysis. The peer-similarity as well as our anomaly detection method is evaluated through experiments. Compared with other methods, the method proposed in this paper reflects the characteristics of simple, sensitive and efficient. And it can be deployed in both online and offline environment.",https://ieeexplore.ieee.org/document/6299088,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 221}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 210}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 70}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 518}]",2.0,,,,['failure-detection'],,2012 IEEE 11th International Symposium on Network Computing and Applications,True,['failure-management'],,False,,,,,,,
205,Internet Traffic Classification Using Constrained Clustering,"['Y. Wang', ' Y. Xiang', ' J. Zhang', ' W. Zhou', ' G. Wei', ' L. T. Yang']",2014,"Statistics-based Internet traffic classification using machine learning techniques has attracted extensive research interest lately, because of the increasing ineffectiveness of traditional port-based and payload-based approaches. In particular, unsupervised learning, that is, traffic clustering, is very important in real-life applications, where labeled training data are difficult to obtain and new patterns keep emerging. Although previous studies have applied some classic clustering algorithms such as K-Means and EM for the task, the quality of resultant traffic clusters was far from satisfactory. In order to improve the accuracy of traffic clustering, we propose a constrained clustering scheme that makes decisions with consideration of some background information in addition to the observed traffic statistics. Specifically, we make use of equivalence set constraints indicating that particular sets of flows are using the same application layer protocols, which can be efficiently inferred from packet headers according to the background knowledge of TCP/IP networking. We model the observed data and constraints using Gaussian mixture density and adapt an approximate algorithm for the maximum likelihood estimation of model parameters. Moreover, we study the effects of unsupervised feature discretization on traffic clustering by using a fundamental binning method. A number of real-world Internet traffic traces have been used in our evaluation, and the results show that the proposed approach not only improves the quality of traffic clusters in terms of overall accuracy and per-class metrics, but also speeds up the convergence.",https://ieeexplore.ieee.org/document/6684161,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 233}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 132}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 179}, {'database': 'IEEE', 'search_string': ""'classification' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 522}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1200}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 167}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 139}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 552}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 206}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1257}]",48.0,['clustering'],,,['failure-detection'],['network'],IEEE Transactions on Parallel and Distributed Systems,True,['failure-management'],,True,['traffic-classification'],,,,,,
206,Workload-Aware Online Anomaly Detection in Enterprise Applications with Local Outlier Factor,"['T. Wang', ' W. Zhang', ' J. Wei', ' H. Zhong']",2012,"Detecting anomalies are essential for improving the reliability of enterprise applications. Current approaches set thresholds for metrics or model correlations between metrics, and anomalies are detected when the thresholds are violated or the correlations are broken. However, we have found that the dynamic workload fluctuating over multiple time scales causes system metrics and their correlations to change. Moreover, it is difficult to model various metric correlations in complex applications. This paper addresses these problems and proposes an online anomaly detection approach for enterprise applications. A method is presented for recognizing workload patterns with an incremental clustering algorithm. The Local Outlier Factor (LOF) based on the specific workload pattern is adopted for detecting anomalies. Our approach is evaluated on a testbed running the TPC-W benchmark. The experimental results show that our approach can capture workload fluctuations accurately and detect the typical faults effectively.",https://ieeexplore.ieee.org/document/6340251,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 446}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 478}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 305}]",4.0,,,,['failure-detection'],,2012 IEEE 36th Annual Computer Software and Applications Conference,True,['failure-management'],,,,,,,,,
207,Identifying Recurrent and Unknown Performance Issues,"['M. Lim', ' J. Lou', ' H. Zhang', ' Q. Fu', ' A. B. J. Teoh', ' Q. Lin', ' R. Ding', ' D. Zhang']",2014,"For a large-scale software system, especially an online service system, when a performance issue occurs, it is desirable to check whether this issue has occurred before. If there are past similar issues, a known remedy could be applied. Otherwise, a new troubleshooting process may have to be initiated. The symptom of a performance issue can be characterized by a set of metrics. Due to the sophisticated nature of software systems, manual diagnosis of performance issues based on metric data is typically expensive and laborious. In this paper, we propose a Hidden Markov Random Field (HMRF) based approach to automatic identification of recurrent and unknown performance issues. We formulate the problem of issue identification as a HMRF-based clustering problem. Our approach incorporates the learning of metric discretization thresholds and the optimization of issue clustering. Based on the learned thresholds and cluster centroids, we can achieve accurate identification of recurrent issues and unknown issues. Experimental evaluations on an open benchmark and a large-scale industrial production system show that our approach is effective and outperforms the related state-of-the-art approaches.",https://ieeexplore.ieee.org/document/7023349,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 528}]",9.0,,,,['failure-detection'],,2014 IEEE International Conference on Data Mining,True,['failure-management'],,True,,,,,,,
208,A measurement-based model for estimation of resource exhaustion in operational software systems,"['K. Vaidyanathan', ' K. S. Trivedi']",1999,"Software systems are known to suffer from outages due to transient errors. Recently, the phenomenon of ""software aging"", in which the state of the software system degrades with time, has been reported (S. Garg et al., 1998). The primary causes of this degradation are the exhaustion of operating system resources, data corruption and numerical error accumulation. This may eventually lead to performance degradation of the software or crash/hang failure, or both. Earlier work in this area to detect aging and to estimate its effect on system resources did not take into account the system workload. In this paper, we propose a measurement-based model to estimate the rate of exhaustion of operating system resources both as a function of time and the system workload state. A semi-Markov reward model is constructed based on workload and resource usage data collected from the UNIX operating system. We first identify different workload states using statistical cluster analysis and build a state-space model. Corresponding to each resource, a reward function is then defined for the model based on the rate of resource exhaustion in the different states. The model is then solved to obtain trends and the estimated exhaustion rates and the time-to-exhaustion for the resources. With the help of this measure, proactive fault management techniques such as ""software rejuvenation"" (Y. Huang et al., 1995) may be employed to prevent unexpected outages.",https://ieeexplore.ieee.org/document/=809313,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 784}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('remediation' OR 'recovery')"", 'index': 345}]",233.0,"['clustering', 'markov-model']",['host-metrics'],['new-method'],['failure-prevention'],['software'],Proceedings 10th International Symposium on Software Reliability Engineering (Cat. No.PR00443),True,['failure-management'],True,,['rejuvenation'],True,18.0,,,,
209,Astro: A predictive model for anomaly detection and feedback-based scheduling on Hadoop,"['C. Gupta', ' M. Bansal', ' T. Chuang', ' R. Sinha', ' S. Ben-romdhane']",2014,"The sheer growth in data volume and Hadoop cluster size make it a significant challenge to diagnose and locate problems in a production-level cluster environment efficiently and within a short period of time. Often times, the distributed monitoring systems are not capable of detecting a problem well in advance when a large-scale Hadoop cluster starts to deteriorate i n performance or becomes unavailable. Thus, inc o m i n g workloads, scheduled between the time when cluster starts to deteriorate and the time when the problem is identified, suffer from longer execution times. As a result, both reliability and throughput of the cluster reduce significantly. In this paper, we address this problem by proposing a system called Astro, which consists of a predictive model and an extension to the Hadoop scheduler. The predictive model in Astro takes into account a rich set of cluster behavioral information that are collected by monitoring processes and model them using machine learning algorithms to predict future behavior of the cluster. The Astro predictive model detects anomalies in the cluster and also identifies a ranked set of metrics that have contributed the most towards the problem. The Astro scheduler uses the prediction outcome and the list of metrics to decide whether it needs to move and reduce workloads from the problematic cluster nodes or to prevent additional workload allocations to them, in order to improve both throughput and reliability of the cluster. The results demonstrate that the Astro scheduler improves usage of cluster compute resources significantly by 64.23% compared to traditional Hadoop. Furthermore, the runtime of the benchmark application reduced by 26.68% during the time of anomaly, thus improving the cluster throughput.",https://ieeexplore.ieee.org/document/7004315,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1466}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('remediation' OR 'recovery')"", 'index': 111}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 416}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 922}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('remediation' OR 'recovery')"", 'index': 493}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1469}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 165}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('anomaly detection' OR 'outlier detection')"", 'index': 244}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 721}]",1.0,"['dimensionality-reduction', 'reasoning', 'case-based-reasoning']",['metrics'],,['failure-detection'],['hadoop'],2014 IEEE International Conference on Big Data (Big Data),True,"['resource-provisioning', 'failure-management']",,,,,,,,,
210,Online Anomaly Prediction for Robust Cluster Systems,"['X. Gu', ' H. Wang']",2009,"In this paper, we present a stream-based mining algorithm for online anomaly prediction. Many real-world applications such as data stream analysis requires continuous cluster operation. Unfortunately, today's large-scale cluster systems are still vulnerable to various software and hardware problems. System administrators are often overwhelmed by the tasks of correcting various system anomalies such as processing bottlenecks (i.e., full stream buffers), resource hot spots, and service level objective (SLO) violations. Our anomaly prediction scheme raises early alerts for impending system anomalies and suggests possible anomaly causes. Specifically, we employ Bayesian classification methods to capture different anomaly symptoms and infer anomaly causes. Markov models are introduced to capture the changing patterns of different measurement metrics. More importantly, our scheme combines Markov models and Bayesian classification methods to predict when a system anomaly will appear in the foreseeable future and what are the possible anomaly causes. To the best of our knowledge, our work provides the first stream-based mining algorithm for predicting system anomalies. We have implemented our approach within the IBM System S distributed stream processing cluster, and conducted case study experiments using fully implemented distributed data analysis applications processing real application workloads. Our experiments show that our approach efficiently predicts and diagnoses several bottleneck anomalies with high accuracy while imposing low overhead to the cluster system.",https://ieeexplore.ieee.org/document/4812472,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1643}]",36.0,,,,['failure-detection'],,2009 IEEE 25th International Conference on Data Engineering,True,['failure-management'],,,,,,,,,
211,Integrating Clustering and Learning for Improved Workload Prediction in the Cloud,"['Y. Yu', ' V. Jindal', ' I. Yen', ' F. Bastani']",2016,"Good resource management is very important in the cloud and workload prediction is a crucial step towards achieving good resource management. While it is possible to predict the workloads of long-running tasks based on the seasonality in their historical workloads, it is difficult to do so for tasks which do not have such recurring workload patterns. In this paper, we consider a different solution for task workload prediction. Instead of using the historical workload of a task to predict the future workload of the same task, we use the knowledge about the workloads of a pool of tasks to help predict the workloads of new tasks. In this paper, we develop a clustering and learning based approach to realize this concept. First, the workloads of existing tasks are grouped into multiple clusters. Then, neural network is used to learn the characteristics of the workloads of each cluster. For each new task, we collect its initial workload, determine its cluster, and use the trained neural network of its cluster to predict its future workload. Our approach is experimentally evaluated using Google dataset. The results confirm the effectiveness of our integrated scheme.",https://ieeexplore.ieee.org/document/7820364,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 4}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 9}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 130}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 18}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 13}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 2}]",11.0,"['multilayer-perceptron', 'clustering']",,['new-method'],['workload-prediction'],,2016 IEEE 9th International Conference on Cloud Computing (CLOUD),True,['resource-provisioning'],,,,,,,,,
212,Evolutionary Neural Network Based Energy Consumption Forecast for Cloud Computing,"['Y. W. Foo', ' C. Goh', ' H. C. Lim', ' Z. Zhan', ' Y. Li']",2015,"The success of Hadoop, an open-source framework for massively parallel and distributed computing, is expected to drive energy consumption of cloud data centers to new highs as service providers continue to add new infrastructure, services and capabilities to meet the market demands. While current research on data center airflow management, HVAC (Heating, Ventilation and Air Conditioning) system design, workload distribution and optimization, and energy efficient computing hardware and software are all contributing to improved energy efficiency, energy forecast in cloud computing remains a challenge. This paper reports an evolutionary computation based modeling and forecasting approach to this problem. In particular, an evolutionary neural network is developed and structurally optimized to forecast the energy load of a cloud data center. The results, both in terms of forecasting speed and accuracy, suggest that the evolutionary neural network approach to energy consumption forecasting for cloud computing is highly promising.",https://ieeexplore.ieee.org/document/7421894,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 74}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 24}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 39}]",9.0,,,,,,2015 International Conference on Cloud Computing Research and Innovation (ICCCRI),True,['resource-provisioning'],,,,,,,,,
213,Resource scheduling in cloud computing using back propagation algorithm,"['B. C. Ashwini', ' Y. S. Nijagunarya']",2017,"Cloud computing is also called as internet -based computing. Cloud computing is one of the distributed computing system. Its main principle is, to use the resources optimally for achieving higher efficiency. Cloud computing is mainly used to solve vast scale issues. IT assets gave by cloud system, then again, are devoted to providing back-end preparing abilities and client based access to these capabilities. In cloud computing technology, clients lease the resources and only pay for what they use. Cloud computing users should be particular while choosing cloud service providers. Main advantage of cloud computing is “agility”. Scheduling of resources is one of the tough problem in cloud computing. The main task of Scheduling is, to assign the resources for the user request, so that the task can be finished in minimum time. In this paper Multiple Artificial Neural Network (MANN) is used for allocating the resources for the user requests. Thus with the help of Multiple Artificial Neural Network, scheduling of resources in cloud computing can be improved.",https://ieeexplore.ieee.org/document/8453280,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 86}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 15}]",0.0,,,,,,2017 2nd International Conference On Emerging Computation and Information Technologies (ICECIT),True,['resource-provisioning'],,True,,,,,,,
214,Data center power management for regulation service using neural network-based power prediction,"['N. Liu', ' X. Lin', ' Y. Wang']",2017,"The underlying infrastructure of cloud computing relies on data centers monitored and maintained by the cloud service providers. Data centers usually incur enormous power consumption and are expected to have a significant impact on the local power grid due to dramatically increasing power consumption and fluctuation. In order to mitigate such fluctuation and balance the power demand and supply in the power grid in real time, the regulation service (RS) opportunity has been provided, which offers the electricity consumers to dynamically adjust their power consumption and reduce their electricity cost. Data centers can be active RS participants due to their flexibility and controllability in load dispatching and scheduling temporally (within a server) and spatially (among multiple servers). In order for the data centers to provide better RS, prediction on the data center power consumption becomes essential. In this work, we first adopt artificial neural network (ANN)-based method and long short term memory (LSTM) neural network-based method for the prediction of future data center power consumption. Based on the prediction results, we formulate a novel optimal power management problem of data center to minimize the total cost. Experimental results demonstrate that the total cost of the data center can be reduced by up to 20.6% compared with the baseline systems.",https://ieeexplore.ieee.org/document/7918343,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 124}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 67}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 1}]",3.0,,,,,,2017 18th International Symposium on Quality Electronic Design (ISQED),True,['resource-provisioning'],,,,,,,,,
215,An SLA-aware load balancing scheme for cloud datacenters,"['Chung-Cheng Li', ' Kuochen Wang']",2014,"One of the most important issues about cloud computing is how to achieve load balancing among thousands of virtual machines (VMs) in a large datacenter. In this paper, we propose a novel decentralized load balancing architecture, called tldlb (two-level decentralized load balancer). This distributed load balancer takes advantage of the decentralized architecture for providing scalability and high availability capabilities to service more cloud users. We also propose a neural network-based dynamic load balancing algorithm, called nn-dwrr (neural network-based dynamic weighted round-robin), to dispatch a large number of requests to different VMs, which are actually providing services. In nn-dwrr, we combine VM load metrics (CPU, memory, network bandwidth, and disk I/O utilizations) monitoring and neural network-based load prediction to adjust the weight of each VM. Experimental results support that our proposed load balancing algorithm, nn-dwrr, can be applied to a large cloud datacenter, and it is 1.86 times faster than the wrr, 1.49 times faster than the Capacity-based, and 1.21 times faster than the ANN-based load balancing algorithms in terms of average response time. In addition, tldlb can reduce the SLA (service-level agreement) violation rate via in-time activating VMs from a spare VM pool.",https://ieeexplore.ieee.org/document/6799665,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 134}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 323}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 162}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 36}]",20.0,,,,['resource-consolidation'],,The International Conference on Information Networking 2014 (ICOIN2014),True,['resource-provisioning'],,,,,,,,,
216,Prediction of cloud data center networks loads using stochastic and neural models,"['J. J. Prevost', ' K. Nagothu', ' B. Kelley', ' M. Jamshidi']",2011,"The increasing demand for cloud computing resources has led to a commensurate increase in the operating power consumption of the systems that comprise the cloud. In this paper, we introduce a novel framework combining load demand prediction and stochastic state transition models. We claim that our model will lead to optimal cloud resource allocation by minimizing energy consumed while maintaining required performance levels. We characterize the ability of neural network and auto-regressive linear prediction algorithms to forecast loads in cloud data center applications. In this paper, the performance of our models against two sets of data at multiple look-ahead times is also presented.",https://ieeexplore.ieee.org/document/5966610,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 180}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 275}]",65.0,['autoregression'],,,['workload-prediction'],['network'],2011 6th International Conference on System of Systems Engineering,True,['resource-provisioning'],,,,,,,,,
217,Dynamic resource scaling in cloud using neural network and black hole algorithm,"['J. Kumar', ' A. K. Singh']",2016,"Cloud computing has gained much attention in recent years. In spite of several advantages, cloud computing involves a number of issues such as dynamic resource scaling and power consumption. These factors lead a cloud system to be inefficient and costly. Workload prediction is one of the factors by which the efficiency of a cloud can be improved and operational cost would be reduced. In this paper, we present a workload prediction model using neural network and black hole algorithm. The experiments were performed on the benchmark data sets of HTTP traces from NASA, Calgary and Saskatchewan web servers. We achieved an improvement on mean squared error upto 134 times over back propagation.",https://ieeexplore.ieee.org/document/7893243,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 287}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 207}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 111}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 161}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 237}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 259}]",9.0,,,,,,2016 Fifth International Conference on Eco-friendly Computing and Communication Systems (ICECCS),True,['resource-provisioning'],,,,,,,,,
218,Dynamic Load Balancing Strategy based on Resource Classification Technique in IaaS Cloud,"['S. Paul', ' M. Adhikari']",2018,"Cloud computing is a utility-based model in the distributed environment which consists of various numbers of resources with heterogeneous servers. The diversity and increasing demands of the user applications lead to increasing resource demands, which makes the whole cloud data center as load imbalanced. The existing algorithms deal with the load distribution in a static and dynamic environment without dealing with a current load of the servers which may balance the load of the servers at certain time interval but not in the long run. So one of the biggest challenges in a cloud environment is to maximize the resource utilization of the servers and balance the load of the whole cloud data center for the long-term process. In order to meet the above-mentioned challenge, in this paper, we have devised a dynamic load balancing strategy based on a baseline neural network technique which will dynamically classify the servers based on the remaining load capacity of the server and deploy the task to the best-fit virtual machine instances on the optimal loaded server. This may minimize the total execution time of the tasks and maximize the resource utilization of the servers while balancing the load of the cloud data center for the long-term process. Finally, we compare the proposed approach over the existing strategies using various performance metrics.",https://ieeexplore.ieee.org/document/8554440,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 406}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 101}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 18}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1357}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 227}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 561}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 404}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 178}, {'database': 'IEEE', 'search_string': ""'classification' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 43}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 509}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 53}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1884}]",0.0,,,,,,"2018 International Conference on Advances in Computing, Communications and Informatics (ICACCI)",True,['resource-provisioning'],,True,,,,,,,
219,Distributed Resource Allocation to Virtual Machines via Artificial Neural Networks,"['D. Minarolli', ' B. Freisleben']",2014,"The goal of the provider with respect to dynamic resource allocation in cloud computing is to maintain application performance according to service level agreements while reducing electrical power costs. To achieve this goal, we present a resource manager that optimizes a utility function expressing the trade-off between the conflicting objectives of maintaining application performance and reducing power costs. It is based on an artificial neural network (ANN) to find the best resource allocation to virtual machines that optimizes the utility function. To provide support for a potentially large number of virtual machines, we present a distributed version of the resource manager consisting of several ANNs in which each ANN is responsible for modeling application performance and power consumption of a single VM while exchanging information with other ANNs to coordinate resource allocation. Simulated and real experiments show the effectiveness of the distributed ANN resource manager over static allocation, a centralized version and a distributed non-coordinated version.",https://ieeexplore.ieee.org/document/6787320,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 462}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 194}]",8.0,,,,,,"2014 22nd Euromicro International Conference on Parallel, Distributed, and Network-Based Processing",True,['resource-provisioning'],,,,,,,,,
220,An Efficient Predictive technique to Autoscale the Resources for Web applications in Private cloud,"['E. G. Radhika', ' G. Sudha Sadasivam', ' J. Fenila Naomi']",2018,"Cloud computing refers to the delivery of computing resources over the network based on user demand. Some web applications may experience different workload at different times, automatic provisioning needs to work efficiently and viably at any point of time. Autoscaling is a feature of cloud computing that has the capability to scale the resources according to demand. Autoscaling provides better fault tolerance, availability and cost management. Although Autoscaling is beneficial, it is not easy to implement. Effective Autoscaling requires techniques to foresee future workload as well as the resources needed to handle the workload. Reactive Autoscaling strategy adds or reduces resources based on threshold set. The predictive strategy is used to address the issues like rapid spike in demand, outages and variable traffic patterns from web applications by providing necessary scaling actions beforehand. In the proposed work, Auto Regressive Integrated Moving Average (ARIMA) and Recurrent Neural Network-Long Short Term Memory (RNN-LSTM) techniques are used for predicting the future workload of based on CPU and RAM usage rate collected from three tier architecture of web application integrated in private cloud. On comparing the performance metrics of both the techniques, the RNN-LSTM deep learning technique gives the minimum error rate and can be applied on large datasets for predicting the future workload of web applications in a private cloud.",https://ieeexplore.ieee.org/document/8480899,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 489}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1907}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 216}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 967}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 213}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 762}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 413}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 863}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 359}]",1.0,['rnn'],['metrics'],,,,"2018 Fourth International Conference on Advances in Electrical, Electronics, Information, Communication and Bio-Informatics (AEEICB)",True,['resource-provisioning'],,,,,,,,,
221,Anomaly Detection Using LSTM in IP Networks,"['G. Qin', ' Y. Chen', ' Y. Lin']",2018,A Long Short Term Memory network technique is being applied in the anomaly detection at Huawei in IP Network Performance Monitor. LSTM networks are trained and used to predict instant network traffic at port level. The output is sent to an anomaly identifier to detect abnormal traffic events. Good precision and recall rate are achieved by using the traffic data from routers.,https://ieeexplore.ieee.org/document/8530862,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 540}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 25}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 36}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 19}]",2.0,,,,['failure-detection'],['network'],2018 Sixth International Conference on Advanced Cloud and Big Data (CBD),True,"['resource-provisioning', 'failure-management']",,,,,,,,,
222,Application-Aware Resource Allocation for SDN-based Cloud Datacenters,"['W. Hong', ' K. Wang', ' Y. Hsu']",2013,"In cloud datacenters, since resource requirements change frequently, how to assign and manage resources efficiently while meeting service level agreements (SLAs) of different types of applications is an important research issue. In this paper, we propose an Application-aware Resource Allocation (App-RA) scheme to predict resource requirements and allocate an appropriate number of virtual machines (VMs) for each application in SDN-based cloud datacenters. To the best of our knowledge, the proposed App-RA is the first application-aware resource allocation scheme that adapts to all types of applications. The App-RA can meet SLAs, allocate resources efficiently, and reduce power consumption for each application in cloud datacenters. The proposed App-RA adopts the neural network based predictor to forecast the requirements of resources (CPU, Memory, GPU, Disk I/O and bandwidth) for an application. In the proposed App-RA, we have designed two algorithms which allocate appropriate numbers of virtual machines and use the VM allocation threshold to avoid SLA violations for five different types of applications. In addition, we adopt an SDN-based OpenFlow network with CICQ switches to appropriately schedule packets for different types of application in the network layer. Finally, simulation results show that the power consumption of the proposed App-RA is only 9.21% higher than that of the best case (oracle) and the power consumption of EAACVA, which is a representative resource allocation method for non-graphic applications, is 104.58% worse than that of App-RA. Furthermore, the SLA violation rate of the proposed App-RA is less than 4% for all applications.",https://ieeexplore.ieee.org/document/6820980,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 641}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 679}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 56}]",12.0,,,,,,2013 International Conference on Cloud Computing and Big Data,True,['resource-provisioning'],,,,,,,,,
223,Reducing Operational Costs through Consolidation with Resource Prediction in the Cloud,"['J. Li', ' K. Shuang', ' S. Su', ' Q. Huang', ' P. Xu', ' X. Cheng', ' J. Wang']",2012,"How to achieve energy efficiency to run a cloud data center is a major challenge in the era of rising electricity cost and environmental protection. Various techniques have been devised to help reduce energy consumption for cloud data centers that consist of a large number of identical servers, including dynamic allocation of active servers, consolidating diverse applications to run on them, and adjusting the CPU speed of an active server. Leveraging these techniques, we use an Online Coloring Bin Packing problem to model the consolidation problem and devise an effective application-aware approximation algorithm to find a near-optimal solution. We show a 1.7 asymptotic approximation ratio. We then apply a Predictive Bayesian Network model to identify daily workload patterns and adjust resource provisioning accordingly. We evaluate our approaches using traces collected from a real data center and demonstrate that (1) our prediction algorithm is effective in estimating future demands, (2) our coordinated approaches can provide significant savings of energy and operational costs close to the near-optimal offline solution, and (3) our approaches incur little reliability costs in term of wear-and-tear of server components.",https://ieeexplore.ieee.org/document/6217513,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 1003}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 943}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 838}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 55}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 17}]",6.0,,,,,,"2012 12th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (ccgrid 2012)",True,['resource-provisioning'],,True,,,,10.1109/CCGrid.2012.50,Conference Paper,"['Consolidation', 'Forecast-based Resource Provisioning', 'Energy Efficiency', 'Data Center']",
224,Genetic Algorithm Optimized BP-network Model and its Application in Fault Detection of Complicate Equipments,"['M. Xianyao', ' W. Haojun']",2010,"The BP nerve network has been used widely in fault detection field, but the usage of solution seeking algorithm by along gradient descent often results in low convergence speed and frequently getting into the part minimum. On the contrary, the genetic algorithm has the advantage of fast seeking speed in full-scale. Therefore, to optimize the BP nerve network, this essay adopts the auto-fit genetic algorithm. Later, the example of shipping main shafting fault detection proves that the optimized BP nerve network combined with the genetic algorithm is more adapted to complicate equipments' fault detection.",https://ieeexplore.ieee.org/document/5523050,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 37}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 162}]",0.0,"['genetic-programming', 'multilayer-perceptron']",,,['failure-detection'],,2010 International Conference on Intelligent Computation Technology and Automation,True,['failure-management'],,,['anomaly-detection'],,,,,,
225,AFD: Adaptive failure detection system for cloud computing infrastructures,"['H. S. Pannu', ' J. Liu', ' Q. Guan', ' S. Fu']",2012,"Cloud computing has become increasingly popular by obviating the need for users to own and maintain complex computing infrastructure. However, due to their inherent complexity and large scale, production cloud computing systems are prone to various runtime problems caused by hardware and software failures. Autonomic failure detection is a crucial technique for understanding emergent, cloud-wide phenomena and self-managing cloud resources for system-level dependability assurance. To detect failures, we need to monitor the cloud execution and collect runtime performance data. These data are usually unlabeled, and thus a prior failure history is not always available in production clouds, especially for newly managed or deployed systems. In this paper, we present an Adaptive Failure Detection (AFD) framework for cloud dependability assurance. AFD employs data description using hypersphere for adaptive failure detection. Based on the cloud performance data, AFD detects possible failures, which are verified by the cloud operators. They are confirmed as either true failures with failure types or normal states. AFD adapts itself by recursively learning from these newly verified detection results to refine future detections. Meanwhile, AFD exploits the observed but undetected failure records reported by the cloud operators to identify new types of failures. We have implemented a prototype of the AFD system and conducted experiments in an on-campus cloud computing environment. Our experimental results show that AFD can achieve more efficient and accurate failure detection than other existing schemes.",https://ieeexplore.ieee.org/document/6407740,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 107}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 219}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 411}]",4.0,,,,['failure-detection'],,2012 IEEE 31st International Performance Computing and Communications Conference (IPCCC),True,['failure-management'],,,,,,,,,
226,Latent fault detection in large scale services,"['M. Gabel', ' A. Schuster', ' R. Bachrach', ' N. Bjørner']",2012,"Unexpected machine failures, with their resulting service outages and data loss, pose challenges to datacenter management. Existing failure detection techniques rely on domain knowledge, precious (often unavailable) training data, textual console logs, or intrusive service modifications. We hypothesize that many machine failures are not a result of abrupt changes but rather a result of a long period of degraded performance. This is confirmed in our experiments, in which over 20% of machine failures were preceded by such latent faults. We propose a proactive approach for failure prevention. We present a novel framework for statistical latent fault detection using only ordinary machine counters collected as standard practice. We demonstrate three detection methods within this framework. Derived tests are domain-independent and unsupervised, require neither background information nor tuning, and scale to very large services. We prove strong guarantees on the false positive rates of our tests.",https://ieeexplore.ieee.org/document/6263932,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 151}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 41}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 81}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 338}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1469}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 135}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault prevention' OR 'failure prevention')"", 'index': 3}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prevention' OR 'failure prevention')"", 'index': 8}]",8.0,,,,['failure-detection'],,IEEE/IFIP International Conference on Dependable Systems and Networks (DSN 2012),True,['failure-management'],,True,,,,,,,
227,AAD: Adaptive Anomaly Detection System for Cloud Computing Infrastructures,"['H. S. Pannu', ' J. Liu', ' S. Fu']",2012,"Cloud computing has become increasingly popular by obviating the need for users to own and maintain complex computing infrastructure. However, due to their inherent complexity and large scale, production cloud computing systems are prone to various runtime problems caused by hardware and software failures. Autonomic failure detection is a crucial technique for understanding emergent, cloudwide phenomena and self-managing cloud resources for system-level dependability assurance. To detect failures, we need to monitor the cloud execution and collect runtime performance data. These data are usually unlabeled, and thus a prior failure history is not always available in production clouds, especially for newly managed or deployed systems. In this paper, we present an Adaptive Anomaly Detection (AAD) framework for cloud dependability assurance. It employs data description using hypersphere for adaptive failure detection. Based on the cloud performance data, AAD detects possible failures, which are verified by the cloud operators. They are confirmed as either true failures with failure types or normal states. The algorithm adapts itself by recursively learning from these newly verified detection results to refine future detections. Meanwhile, it exploits the observed but undetected failure records reported by the cloud operators to identify new types of failures. We have implemented a prototype of the algorithm and conducted experiments in an on-campus cloud computing environment. Our experimental results show that AAD can achieve more efficient and accurate failure detection than other existing scheme.",https://ieeexplore.ieee.org/document/6424881,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 228}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 77}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 149}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 404}]",17.0,,,,['failure-detection'],,2012 IEEE 31st Symposium on Reliable Distributed Systems,True,['failure-management'],,True,,,,,,,
228,Anomaly-based Fault Detection System in Distributed System,"['B. u. Kim', ' S. Hariri']",2007,"One of the important design criteria for distributed systems and their applications is their reliability and robustness to hardware and software failures. The increase in complexity, inter connectedness, dependency and the asynchronous interactions between the components that include hardware resources (computers, servers, network devices), and software (application services, middleware, web services, etc.) makes the fault detection and tolerance a challenging research problem. In this paper, we present an innovative approach based on statistical and data mining techniques to detect faults (hardware or software) and also identify the source of the fault. In our approach, we monitor and analyze in realtime all the interactions between all the components of a distributed system. We used data mining and supervised learning techniques to obtain the rules that can accurately model the normal interactions among these components. Our anomaly analysis engine will immediately produce an alert whenever one or more of the interaction rules that capture normal operations is violated due to a software or hardware failure. We evaluate the effectiveness of our approach and its performance to detect software faults that we inject asynchronously, and compare the results for different noise level.",https://ieeexplore.ieee.org/document/4297016,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 448}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 124}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 723}]",5.0,,,,['failure-detection'],,"5th ACIS International Conference on Software Engineering Research, Management & Applications (SERA 2007)",True,['failure-management'],,False,,,,,,,
229,A neural net based approach for fault diagnosis in distribution networks,"['K. L. Butler', ' J. A. Momoh']",1999,"This paper discusses the application of field data to a new supervised clustering-based arcing distribution fault diagnosis method. The fault diagnosis method can perform three functions that provide preliminary fault location information for grounded and ungrounded power distribution systems: fault detection, faulted type classification, and faulted phase identification. It contains two main modules: a preprocessor and a pattern classifier which was implemented as a supervised clustering-based neural net. The inputs to the fault diagnosis method are the three phase and neutral currents for a feeder. The preprocessor computes a vector of statistical features from the phase currents and passes them to the neural net pattern classifier. The neural net determines the features pattern as normal or faulted. If detected as faulted, the neural net also identifies the fault type and classifies the faulted phase. Field studies were conducted in which the fault diagnosis method was trained and tested with normal and faulted phase currents generated from data recorded by events staged in the field for two, four-wire systems. The fault diagnosis method was highly successful during test to validate the fault detection and identification functions. Also the fault diagnosis method was able to recognize the difference between faulted test patterns and fault-like test patterns representing line switching and load tap changer operations. Further the clustering-based fault diagnosis approach was evaluated using simulated data generated for a 3-feeder ungrounded system.",https://ieeexplore.ieee.org/document/=747478,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 491}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 492}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 218}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 219}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 137}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 138}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault detection' OR 'failure detection')"", 'index': 529}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault detection' OR 'failure detection')"", 'index': 862}]",46.0,"['multilayer-perceptron', 'clustering']",,['new-method'],['root-cause-analysis'],['network'],IEEE Power Engineering Society. 1999 Winter Meeting (Cat. No.99CH36233),True,['failure-management'],,,,,,,,,
230,Proactive failure detection learning generation patterns of large-scale network logs,"['T. Kimura', ' A. Watanabe', ' T. Toyono', ' K. Ishibashi']",2015,"With the growth of services in IP networks, network operators are required to perform proactive operation that quickly detects the signs of critical failures and prevents future problems. Network log data, including router syslog, are rich sources for such operations. However, it has become impossible to find genuinely important logs that lead to serious problems due to the large volume and complexity of log data. We propose a log analysis system for proactive detection of failures. Our key observation is that the abnormality of logs depends on not just the keywords in the messages (e.g. ERROR, FAIL), but generation patterns such as burstiness. Our system consists of three functions: (i) extracting log templates automatically and quickly from a massive amount of unstructured log data; (ii) constructing log feature vectors to characterize the generation patterns of logs; and (iii) using a supervised machine learning approach to associate failures with the log data that appeared before them. We validated our system using real log data collected from a large network and determined its effectiveness.",https://ieeexplore.ieee.org/document/7367332,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 526}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 99}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 56}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 33}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 47}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 267}]",51.0,['support-vector-machine'],"['tickets', 'logs']",['new-method'],['failure-prediction'],['network'],2015 11th International Conference on Network and Service Management (CNSM),True,['failure-management'],True,,['system-failure-prediction'],,,,,,
231,Detecting application-level failures in component-based Internet services,"['E. Kiciman', ' A. Fox']",2005,"Most Internet services (e-commerce, search engines, etc.) suffer faults. Quickly detecting these faults can be the largest bottleneck in improving availability of the system. We present Pinpoint, a methodology for automating fault detection in Internet services by: 1) observing low-level internal structural behaviors of the service; 2) modeling the majority behavior of the system as correct; and 3) detecting anomalies in these behaviors as possible symptoms of failures. Without requiring any a priori application-specific information, Pinpoint correctly detected 89%-96% of major failures in our experiments, as compared with 20%-70% detected by current application-generic techniques.",https://ieeexplore.ieee.org/document/1510707,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 744}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1065}]",77.0,,,,['root-cause-analysis'],,IEEE Transactions on Neural Networks,True,['failure-management'],,,,,,,,,
232,Statistical analysis of network traffic for adaptive faults detection,['H. Hajji'],2005,"This paper addresses the problem of normal operation baselining for automatic detection of network anomalies. A model of network traffic is presented in which studied variables are viewed as sampled from a finite mixture model. Based on the stochastic approximation of the maximum likelihood function, we propose baselining network normal operation, using the asymptotic distribution of the difference between successive estimates of model parameters. The baseline random variable is shown to be stationary, with mean zero under normal operation. Anomalous events are shown to induce an abrupt jump in the mean. Detection is formulated as an online change point problem, where the task is to process the baseline random variable realizations, sequentially, and raise alarms as soon as anomalies occur. An analytical expression of false alarm rate allows us to choose the design threshold, automatically. Extensive experimental results on a real network showed that our monitoring agent is able to detect unusual changes in the characteristics of network traffic, adapt to diurnal traffic patterns, while maintaining a low alarm rate. Despite large fluctuations in network traffic, this work proves that tailoring traffic modeling to specific goals can be efficiently achieved.",https://ieeexplore.ieee.org/document/1510709,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 819}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 1348}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 824}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 371}]",50.0,,,,['failure-detection'],,IEEE Transactions on Neural Networks,True,['failure-management'],,,,,,,,,
233,Adaptive Anomaly Identification by Exploring Metric Subspace in Cloud Computing Infrastructures,"['Q. Guan', ' S. Fu']",2013,"Cloud computing has become increasingly popular by obviating the need for users to own and maintain complex computing infrastructures. However, due to their inherent complexity and large scale, production cloud computing systems are prone to various runtime problems caused by hardware and software faults and environmental factors. Autonomic anomaly detection is a crucial technique for understanding emergent, cloud-wide phenomena and self-managing cloud resources for system-level dependability assurance. To detect anomalous cloud behaviors, we need to monitor the cloud execution and collect runtime cloud performance data. These data consist of values of performance metrics for different types of failures, which display different correlations with the performance metrics. In this paper, we present an adaptive anomaly identification mechanism that explores the most relevant principal components of different failure types in cloud computing infrastructures. It integrates the cloud performance metric analysis with filtering techniques to achieve automated, efficient, and accurate anomaly identification. The proposed mechanism adapts itself by recursively learning from the newly verified detection results to refine future detections. We have implemented a prototype of the anomaly identification system and conducted experiments in an on-campus cloud computing environment and by using the Google data center traces. Our experimental results show that our mechanism can achieve more efficient and accurate anomaly detection than other existing schemes.",https://ieeexplore.ieee.org/document/6656276,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 1044}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 889}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 180}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 371}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 141}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1610}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 980}]",49.0,,,,['failure-detection'],,2013 IEEE 32nd International Symposium on Reliable Distributed Systems,True,['failure-management'],,,['anomaly-detection'],,,,,,
234,Recent Advances in Fault Localization in Computer Networks,"['A. Dusia', ' A. S. Sethi']",2016,"Fault localization, a core element in network fault management, is the process of inferring the exact failure in a network from the set of observed symptoms. Since faults in network systems can be unavoidable, their quick and accurate detection and diagnosis is important for the stability, consistency, and performance of a communication system. In this paper, we discuss the challenges of fault localization in complex communication systems and present an overview of recent techniques proposed in the literature along with their advantages and limitations. We start by briefly surveying passive monitoring techniques which were previously reviewed in a survey by Steinder. We then describe more recent fault localization research in five categories: 1) active monitoring techniques; 2) techniques for overlay and virtual networks; 3) decentralized probabilistic management techniques; 4) temporal correlation techniques; and 5) learning techniques.",https://ieeexplore.ieee.org/document/7471418,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 1110}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 19}]",13.0,,,,['root-cause-analysis'],,IEEE Communications Surveys & Tutorials,True,['failure-management'],,,,,,,,,
235,Alarm correlation and network fault resolution using the Kohonen self-organising map,"['R. D. Gardner', ' D. A. Harle']",1997,"A vital role for a large telecommunications operator is the ability to effectively manage its network. With increasing system size, complexity and bandwidth threatening to overload current management systems, it is no surprise that network management is a high priority for leading telecommunications providers. In this paper, a novel approach to fault management is investigated in which the advantages of artificial neural network technologies are used to complement and extend the capability of traditional systems. In particular, it is shown how a system based on Kohonen (1990) self-organising map theory can be successfully implemented as a means of identifying fault scenarios arising in an SDH-based environment.",https://ieeexplore.ieee.org/document/=644365,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 1320}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 824}]",78.0,,,,['root-cause-analysis'],,GLOBECOM 97. IEEE Global Telecommunications Conference. Conference Record,True,['failure-management'],,,,,,,,,
236,Failure diagnosis using decision trees,"['M. Chen', ' A. X. Zheng', ' J. Lloyd', ' M. I. Jordan', ' E. Brewer']",2004,"We present a decision tree learning approach to diagnosing failures in large Internet sites. We record runtime properties of each request and apply automated machine learning and data mining techniques to identify the causes of failures. We train decision trees on the request traces from time periods in which user-visible failures are present. Paths through the tree are ranked according to their degree of correlation with failure, and nodes are merged according to the observed partial order of system components. We evaluate this approach using actual failures from eBay, and find that, among hundreds of potential causes, the algorithm successfully identifies 13 out of 14 true causes of failure, along with 2 false positives. We discuss some results in applying simplified decision trees on eBay's production site for several months. In addition, we give a cost-benefit analysis of manual vs. automated diagnosis systems. Our contributions include the statistical learning approach, the adaptation of decision trees to the context of failure diagnosis, and the deployment and evaluation of our tools on a high-volume production service.",https://ieeexplore.ieee.org/document/1301345,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 1326}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1399}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 258}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 277}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault detection' OR 'failure detection')"", 'index': 149}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('remediation' OR 'recovery')"", 'index': 186}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 419}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 574}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 692}]",89.0,['decision-tree'],,,['root-cause-analysis'],,"International Conference on Autonomic Computing, 2004. Proceedings.",True,['failure-management'],,,,,,,,,
237,Discovering Rules from Disk Events for Predicting Hard Drive Failures,"['V. Agarwal', ' C. Bhattacharyya', ' T. Niranjan', ' S. Susarla']",2009,"Detecting impending failure of hard disks is an important prediction task which might help computer systems to prevent loss of data and performance degradation. Currently most of the hard drive vendors support self-monitoring, analysis and reporting technology (SMART) which are often considered unreliable for such tasks. The problem of finding alternatives to SMART for predicting disk failure is an area of active research. In this paper, we consider events recorded from live disks and show that it is possible to construct decision support systems which can detect such failures. It is desired that any such prediction methodology should have high accuracy and ease of interpretability. Black box models can deliver highly accurate solutions but do not provide an understanding of events which explains the decision given by it. To this end we explore rule based classifiers for predicting hard disk failures from various disk events. We show that it is possible to learn easy to understand rules, from disk events, which have extremely low false alarm rates on real world data.",https://ieeexplore.ieee.org/document/5382104,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 1330}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 165}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 264}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 972}]",7.0,,,,['failure-prediction'],,2009 International Conference on Machine Learning and Applications,True,['failure-management'],,,,,,,,,
238,Modeling and Tracking of Transaction Flow Dynamics for Fault Detection in Complex Systems,"['G. Jiang', ' H. Chen', ' K. Yoshihira']",2006,"With the prevalence of Internet services and the increase of their complexity, there is a growing need to improve their operational reliability and availability. While a large amount of monitoring data can be collected from systems for fault analysis, it is hard to correlate this data effectively across distributed systems and observation time. In this paper, we analyze the mass characteristics of user requests and propose a novel approach to model and track transaction flow dynamics for fault detection in complex information systems. We measure the flow intensity at multiple checkpoints inside the system and apply system identification methods to model transaction flow dynamics between these measurements. With the learned analytical models, a model-based fault detection and isolation method is applied to track the flow dynamics in real time for fault detection. We also propose an algorithm to automatically search and validate the dynamic relationship between randomly selected monitoring points. Our algorithm enables systems to have self-cognition capability for system management. Our approach is tested in a real system with a list of injected faults. Experimental results demonstrate the effectiveness of our approach and algorithms",https://ieeexplore.ieee.org/document/4012644,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('fault detection' OR 'failure detection')"", 'index': 131}]",83.0,,,,['failure-detection'],,IEEE Transactions on Dependable and Secure Computing,True,['failure-management'],,,,,,,,,
239,Predicting the location and number of faults in large software systems,"['T. J. Ostrand', ' E. J. Weyuker', ' R. M. Bell']",2005,"Advance knowledge of which files in the next release of a large software system are most likely to contain the largest numbers of faults can be a very valuable asset. To accomplish this, a negative binomial regression model has been developed and used to predict the expected number of faults in each file of the next release of a system. The predictions are based on the code of the file in the current release, and fault and modification history of the file from previous releases. The model has been applied to two large industrial systems, one with a history of 17 consecutive quarterly releases over 4 years, and the other with nine releases over 2 years. The predictions were quite accurate: for each release of the two systems, the 20 percent of the files with the highest predicted number of faults contained between 71 percent and 92 percent of the faults that were actually detected, with the overall average being 83 percent. The same model was also used to predict which files of the first system were likely to have the highest fault densities (faults per KLOC). In this case, the 20 percent of the files with the highest predicted fault densities contained an average of 62 percent of the system's detected faults. However, the identified files contained a much smaller percentage of the code mass than the files selected to maximize the numbers of faults. The model was also used to make predictions from a much smaller input set that only contained fault data from integration testing and later. The prediction was again very accurate, identifying files that contained from 71 percent to 93 percent of the faults, with the average being 84 percent. Finally, a highly simplified version of the predictor selected files containing, on average, 73 percent and 74 percent of the faults for the two systems.",https://ieeexplore.ieee.org/document/1435354,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('fault detection' OR 'failure detection')"", 'index': 306}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 93}]",680.0,['linear-regression'],"['code-metrics', 'code-history']",,['failure-prevention'],['source-code'],IEEE Transactions on Software Engineering,True,['failure-management'],True,,['software-defect-prediction'],True,7.0,,,,
240,Robust prediction of fault-proneness by random forests,"['L. Guo', ' Y. Ma', ' B. Cukic', ' Harshinder Singh']",2004,"Accurate prediction of fault prone modules (a module is equivalent to a C function or a C+ + method) in software development process enables effective detection and identification of defects. Such prediction models are especially beneficial for large-scale systems, where verification experts need to focus their attention and resources to problem areas in the system under development. This paper presents a novel methodology for predicting fault prone modules, based on random forests. Random forests are an extension of decision tree learning. Instead of generating one decision tree, this methodology generates hundreds or even thousands of trees using subsets of the training data. Classification decision is obtained by voting. We applied random forests in five case studies based on NASA data sets. The prediction accuracy of the proposed methodology is generally higher than that achieved by logistic regression, discriminant analysis and the algorithms in two machine learning software packages, WEKA [I. H. Witten et al. (1999)] and See5. The difference in the performance of the proposed methodology over other methods is statistically significant. Further, the classification accuracy of random forests is more significant over other methods in larger data sets.",https://ieeexplore.ieee.org/document/1383136,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('fault detection' OR 'failure detection')"", 'index': 470}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('fault detection' OR 'failure detection')"", 'index': 19}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 896}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1072}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault detection' OR 'failure detection')"", 'index': 1800}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault detection' OR 'failure detection')"", 'index': 322}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 90}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 517}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 590}]",30.0,,,,['failure-prediction'],,15th International Symposium on Software Reliability Engineering,True,['failure-management'],,,,,,,,,
241,Spectrum-Based Multiple Fault Localization,"['R. Abreu', ' P. Zoeteweij', ' A. J. C. v. Gemund']",2009,"Fault diagnosis approaches can generally be categorized into spectrum-based fault localization (SFL, correlating failures with abstractions of program traces), and model-based diagnosis (MBD, logic reasoning over a behavioral model). Although MBD approaches are inherently more accurate than SFL, their high computational complexity prohibits application to large programs. We present a framework to combine the best of both worlds, coined BARINEL. The program is modeled using abstractions of program traces (as in SFL) while Bayesian reasoning is used to deduce multiple-fault candidates and their probabilities (as in MBD). A particular feature of BARINEL is the usage of a probabilistic component model that accounts for the fact that faulty components may fail intermittently. Experimental results on both synthetic and real software programs show that BARINEL typically outperforms current SFL approaches at a cost complexity that is only marginally higher. In the context of single faults this superiority is established by formal proof.",https://ieeexplore.ieee.org/document/5431781,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 355}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 29}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 10}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault localization' OR 'failure localization')"", 'index': 9}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault localization' OR 'failure localization')"", 'index': 8}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 20}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault localization' OR 'failure localization')"", 'index': 42}]",290.0,"['collaborative-filtering', 'reasoning', 'naive-bayes']","['runs', 'test-cases']",['new-method'],['root-cause-analysis'],['software'],2009 IEEE/ACM International Conference on Automated Software Engineering,True,['failure-management'],True,True,['fault-localization'],True,76.0,10.1109/ASE.2009.25,Conference Paper,"['program spectra', 'Software fault diagnosis', 'statistical and reasoning approaches']",
242,Intelligent Automated Diagnosis of Client Device Bottlenecks in Private Clouds,"['C. Widanapathirana', ' J. Li', ' Y. A. Sekercioglu', ' M. Ivanovich', ' P. Fitzpatrick']",2011,"We present an automated solution for rapid diagnosis of client device problems in private cloud environments: the Intelligent Automated Client Diagnostic (IACD) system. Clients are diagnosed with the aid of Transmission Control Protocol (TCP) packet traces, by (i) observation of anomalous artifacts occurring as a result of each fault and (ii) subsequent use of the inference capabilities of soft-margin Support Vector Machine (SVM) classifiers. The IACD system features a modular design and is extendible to new faults, with detection capability unaffected by the TCP variant used at the client. Experimental evaluation of the IACD system in a controlled environment demonstrated an overall diagnostic accuracy of 98%.",https://ieeexplore.ieee.org/document/6123506,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 774}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 223}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 908}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 82}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 634}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 148}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 29}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 85}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 133}, {'database': 'arxiv', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 32}, {'database': 'arxiv', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 272}, {'database': 'arxiv', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 979}, {'database': 'arxiv', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 7}, {'database': 'arxiv', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 9}, {'database': 'arxiv', 'search_string': ""('support vector machine' OR 'SVM') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 6}]",1.0,['support-vector-machine'],['traces'],,['root-cause-analysis'],,2011 Fourth IEEE International Conference on Utility and Cloud Computing,True,['failure-management'],,True,,,,10.1109/UCC.2011.42,,,
243,Reasoning Based Workload Performance Prediction in Cloud Data Centers,"['A. Aslam', ' H. Chen', ' J. Xiao', ' H. Jin']",2019,"Cloud computing provides utility-based and scalable services to end-users. In the past decade, the demands for resource management in cloud computing have increased substantially which lead to certain challenges such as optimal resource utilization, power consumption, and service level agreement violations. Workload performance prediction serves as an assistance to address these issues. In this paper, we propose a prediction model based on clustered Case-Based Reasoning (CBR). The proposed model determines the performance metrics for workload prior to the co-operation of autonomic computing characteristics. Thus, CBR provides optimal scheduling of resources and workload monitoring for cloud data centers. In order to validate the proposed CBR-based prediction model, we perform a series of experiments and evaluate the effectiveness in terms of precision, recall, f-measure, and mean square error rate. We generate the cases for CBR using traces from the Google cluster data center. Moreover, we also validate our proposed prediction model against Support Vector Machine (SVM) prediction scheme. Experimental results show that the proposed CBR outperforms the SVM-based approach and yields 10% improvement in terms of precision.",https://ieeexplore.ieee.org/document/8968830,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 819}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 13}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 106}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 895}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 427}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 77}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 70}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 354}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 87}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 196}]",0.0,['case-based-reasoning'],['traces'],"['new-method', 'comparison']",['workload-prediction'],,2019 IEEE International Conference on Cloud Computing Technology and Science (CloudCom),True,['resource-provisioning'],,,,,,,,,
244,Robust and Unsupervised KPI Anomaly Detection Based on Conditional Variational Autoencoder,"['Z. Li', ' W. Chen', ' D. Pei']",2018,"To ensure undisrupted web-based services, operators need to closely monitor various KPIs (Key Performance Indicator, such as CPU usages, network throughput, page views, number of online users, and etc), detect anomalies in them, and trigger timely troubleshooting or mitigation. There can be hundreds of thousands to even millions of KPIs to be monitored, thus operators need automatic anomaly detection approaches. However, neither traditional statistical approaches nor supervised ensemble approaches satisfy this requirement in practice when facing large number of KPIs. A state-of-art unsupervised approach Donut offering promising results, but it is not a sequential model thus cannot deal with the time information related anomalies. Thus, in this paper we propose Bagel, a robust and unsupervised anomaly detection algorithm for KPI that can handle time information related anomalies, using CVAE to incorporate time information and dropout layer to avoid overfitting. Our experiments using real data from Internet companies show that, compared to Donut, Bagel improves the anomaly detection best F1-score by 0.08 to 0.43.",https://ieeexplore.ieee.org/document/8710885,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 17}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 9}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 38}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 110}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 245}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 338}]",2.0,,,,['failure-detection'],,2018 IEEE 37th International Performance Computing and Communications Conference (IPCCC),True,['failure-management'],,,,,,,,,
245,"Applying Machine Learning to Predict Software Fault Proneness Using Change Metrics, Static Code Metrics, and a Combination of Them","['Y. A. Alshehri', ' K. Goseva-Popstojanova', ' D. G. Dzielski', ' T. Devine']",2018,"Predicting software fault proneness is very important as the process of fixing these faults after the release is very costly and time-consuming. In order to predict software fault proneness, many machine learning algorithms (e.g., Logistic regression, Naive Bayes, and J48) were used on several datasets, using different metrics as features. The question is what algorithm is the best under which circumstance and what metrics should be applied. Related works suggested that using change metrics leads to the highest accuracy in prediction. In addition, some algorithms perform better than others in certain circumstances. In this work, we use three machine learning algorithms (i.e., logistic regression, Naive Bayes, and J48) on three Eclipse releases (i.e., 2.0, 2.1, 3.0). The results showed that accuracy is slightly better and false positive rates are lower, when we use the reduced set of metrics compared to all change metrics set. However, the recall and the G score are better when we use the complete set of change metrics. Furthermore, J48 outperformed the other classifiers with respect to the G score for the reduced set of change metrics, as well as in almost all cases when the complete set of change metrics, static code metrics, and the combination of both were used.",https://ieeexplore.ieee.org/document/8478911,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 21}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 13}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 2}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 2}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 10}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 64}]",1.0,,,,['failure-prediction'],,SoutheastCon 2018,True,['failure-management'],,,,,,,,,
246,A comparative analysis of the efficiency of change metrics and static code attributes for defect prediction,"['R. Moser', ' W. Pedrycz', ' G. Succi']",2008,"In this paper we present a comparative analysis of the predictive power of two different sets of metrics for defect prediction. We choose one set of product related and one set of process related software metrics and use them for classifying Java files of the Eclipse project as defective respective defect-free. Classification models are built using three common machine learners: logistic regression, naive Bayes, and decision trees. To allow different costs for prediction errors we perform cost-sensitive classification, which proves to be very successful: >75% percentage of correctly classified files, a recall of >80%, and a false positive rate <30%. Results indicate that for the Eclipse data, process metrics are more efficient defect predictors than code metrics.",https://ieeexplore.ieee.org/document/4814129,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 40}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 41}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 49}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 8}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 5}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 54}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 65}]",303.0,"['naive-bayes', 'logistic-regression', 'decision-tree']",['code-metrics'],['comparison'],['failure-prevention'],['source-code'],2008 ACM/IEEE 30th International Conference on Software Engineering,True,['failure-management'],True,True,['software-defect-prediction'],True,8.0,10.1145/1368088.1368114,Conference Paper,"['defect prediction', 'software metrics', 'cost-sensitive classification']",
247,Deep Learning Approach for Software Maintainability Metrics Prediction,"['S. Jha', ' R. Kumar', ' L. Hoang Son', ' M. Abdel-Basset', ' I. Priyadarshini', ' R. Sharma', ' H. Viet Long']",2019,"Software maintainability predicts changes or failures that may occur in software after it has been deployed. Since it deals with the degree to which an application may be understood, repaired, or enhanced, it also takes into account the overall cost of the project. In the past, several measures have been taken into account for predicting metrics that influence software maintainability. However, deep learning is yet to be explored for the same. In this paper, we perform deep learning for software maintainability metrics' prediction on a large number of datasets. Unlike the previous research works, we have relied on large datasets from 299 software and subsequently applied various metrics and functions to the same; 29 object-oriented metrics have been considered along with their impact on software maintainability of open source software. Several metrics have been analyzed and descriptive statistics of these metrics have been pointed out. The proposed long short term memory has been evaluated using measures, such as mean absolute error, root mean square error and accuracy. Five machine learning algorithms, namely, ridge regression with variable selection, decision tree, quantile regression forest, support vector machine, and principal component analysis have been applied to the original datasets, as well as, to the refined datasets. It was found that this paper provides results in the form of metrics that may be used in the prediction of software maintenance and the proposed deep learning model outperforms all of the other methods that were considered. Furthermore, the results of experiment affirm the efficiency of the proposed deep learning model for software maintainability prediction.",https://ieeexplore.ieee.org/document/8698760,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 72}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 51}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 36}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 2}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 54}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 92}]",8.0,"['random-forest', 'linear-regression', 'support-vector-machine', 'decision-tree', 'dimensionality-reduction', 'rnn']",['code-metrics'],"['comparison', 'novel-use']",['failure-prevention'],['source-code'],,True,['failure-management'],False,True,['software-defect-prediction'],,,,,,
248,Deep Semantic Feature Learning with Embedded Static Metrics for Software Defect Prediction,"['G. Fan', ' X. Diao', ' H. Yu', ' K. Yang', ' L. Chen']",2019,"Software defect prediction, which locates defective code snippets, can assist developers in finding potential bugs and assigning their testing efforts. Traditional defect prediction features are static code metrics, which only contain statistic information of programs and fail to capture semantics in programs, leading to the degradation of defect prediction performance. To take full advantage of the semantics and static metrics of programs, we propose a framework called Defect Prediction via Attention Mechanism (DP-AM) in this paper. Specifically, DPAM first extracts vectors which are then encoded as digital vectors by mapping and word embedding from abstract syntax trees (ASTs) of programs. Then it feeds these numerical vectors into Recurrent Neural Network to automatically learn semantic features of programs. After that, it applies self-attention mechanism to further build relationship among these features. Furthermore, it employs global attention mechanism to generate significant features among them. Finally, we combine these semantic features with traditional static metrics for accurate software defect prediction. We evaluate our method in terms of F1-measure on seven open-source Java projects in Apache. Our experimental results show that DP-AM improves F1-measure by 11% in average, compared with the state-of-the-art methods.",https://ieeexplore.ieee.org/document/8946058,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 87}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 193}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 154}]",0.0,['rnn'],['source-code'],['novel-use'],['failure-prevention'],['apache'],2019 26th Asia-Pacific Software Engineering Conference (APSEC),True,['failure-management'],,,['software-defect-prediction'],,,,,,
249,An Investigation of the Effect of Discretization on Defect Prediction Using Static Measures,"['P. Singh', ' S. Verma']",2009,"Software repositories with defect logs are main resource for defect prediction. In recent years, researchers have used the vast amount of data that is contained by software repositories to predict the location of defect in the code that caused problems. In this paper we evaluate the effectiveness of software fault prediction with Naive-Bayes classifiers and J48 classifier by integrating with supervised discretization algorithm developed by Fayyad and Irani. Public datasets from the promise repository have been explored for this purpose. The repository contains software metric data and error data at the function/method level. Our experiment shows that integration of discretization method improves the software fault prediction accuracy when integrated with Naive-Bayes and J48 classifiers.",https://ieeexplore.ieee.org/document/5375760,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 99}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 78}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 141}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 122}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 425}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 48}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 204}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1271}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 93}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1336}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 163}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 26}]",8.0,,,,['failure-prediction'],,"2009 International Conference on Advances in Computing, Control, and Telecommunication Technologies",True,['failure-management'],,True,,,,,,,
250,Predicting Defect-Prone Software Modules at Different Logical Levels,"['P. Huang', ' J. Zhu']",2009,"Effective software defect estimation can bring cost reduction and efficient resources allocation in software development and testing. Usually, estimation of defect-prone modules is based on the supervised learning of the modules at the same logical level. Various practical issues may limit the availability or quality of the attribute-value vectors extracting from the high-level modules by software metrics. In this paper, the problem of estimating the defect in high-level software modules is investigated with a multi-instance learning (MIL) perspective. In detail, each high-level module is regarded as a bag of its low-level components, and the learning task is to estimate the defect-proneness of the bags. Several typical supervised learning and MIL algorithms are evaluated on a mission critical project from NASA. Compared to the selected supervised schemas, the MIL methods improve the performance of the software defect estimation models.",https://ieeexplore.ieee.org/document/5401286,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 109}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 28}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 55}]",5.0,,['code-metrics'],,['failure-prevention'],,2009 International Conference on Research Challenges in Computer Science,True,['failure-management'],,,['software-defect-prediction'],,,,,,
251,Robust and Rapid Adaption for Concept Drift in Software System Anomaly Detection,"['M. Ma', ' S. Zhang', ' D. Pei', ' X. Huang', ' H. Dai']",2018,"Anomaly detection is critical for web-based software systems. Anecdotal evidence suggests that in these systems, the accuracy of a static anomaly detection method that was previously ensured is bound to degrade over time. It is due to the significant change of data distribution, namely concept drift, which is caused by software change or personal preferences evolving. Even though dozens of anomaly detectors have been proposed over the years in the context of software system, they have not tackled the problem of concept drift. In this paper, we present a framework, StepWise, which can detect concept drift without tuning detection threshold or per-KPI (Key Performance Indicator) model parameters in a large scale KPI streams, take external factors into account to distinguish the concept drift which under operators' expectations, and help any kind of anomaly detection algorithm to handle it rapidly. For the prototype deployed in Sogou, our empirical evaluation shows StepWise improve the average F-score by 206% for many widely-used anomaly detectors over a baseline without any concept drift detection.",https://ieeexplore.ieee.org/document/8539065,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 111}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 122}]",6.0,,,,['failure-detection'],,2018 IEEE 29th International Symposium on Software Reliability Engineering (ISSRE),True,['failure-management'],,,,,,,,,
252,Software metrics for fault prediction using machine learning approaches: A literature review with PROMISE repository dataset,"['Meiliana', ' S. Karim', ' H. L. H. S. Warnars', ' F. L. Gaol', ' E. Abdurachman', ' B. Soewito']",2017,"Software testing is an important and critical phase of software development life cycle to find software faults or defects and then correct those faults. However, testing process is a time-consuming activity that requires good planning and a lot of resources. Therefore, technique and methodology for predicting the testing effort is important process prior the testing process to significantly increase efficiency of time, effort and cost usage. Correspond to software metric usage for measuring software quality, software metric can be used to identify the faulty modules in software. Furthermore, implementing machine learning technique will allow computer to “learn” and able to predict the fault prone modules. Research in this field has become a hot issue for more than ten years ago. However, considering the high importance of software quality with support of machine learning methods development, this research area is still being highlighted until this year. In this paper, a survey of various software metric used for predicting software fault by using machine learning algorithm is examined. According to our review, this is the first study of software fault prediction that focuses to PROMISE repository dataset usage. Some conducted experiments from PROMISE repository dataset are compared to contribute a consensus on what constitute effective software metrics and machine learning method in software fault prediction.",https://ieeexplore.ieee.org/document/8311708,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 126}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 95}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 80}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 14}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 9}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 256}]",5.0,,['software-metrics'],['survey'],"['failure-prediction', 'root-cause-analysis']",['source-code'],2017 IEEE International Conference on Cybernetics and Computational Intelligence (CyberneticsCom),True,['failure-management'],True,,"['software-defect-prediction', 'rca-others']",,,,,,
253,Empirical Investigation of Code and Process Metrics for Defect Prediction,"['W. Han', ' C. Lung', ' S. A. Ajila']",2016,"Data science is becoming more important for software engineering problems. Software defect prediction is a critical area which can help the development team allocate test resource efficiently and better understand the root cause of defects. Furthermore, it can help find the reason why a component or even a project is failure-prone. This paper deals with binary classification in predicting if a software component has a bug by using three widely used machine learning algorithms: Random Forest (RF), Neural Networks (NN), and Support Vector Machine (SVM). The paper investigates the applications of these algorithms to the challenging issue of predicting defects in software components. This paper combines code metrics and process metrics as indicators for the Eclipse environment using the aforementioned three algorithms for a sample of weekly Eclipse features. Feature reduction is also adopted using General Linear Model (GLM) to save computational time. The results confirm the predictive capabilities of using two features -- NBD_max and Pre-defects -- and are comparable to the results of using all 61 features. Additionally, this paper evaluates the performance of the three algorithms. NN and RF turn out to have the best fit.",https://ieeexplore.ieee.org/document/7545064,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 152}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 91}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 67}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 345}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 248}]",2.0,,,,['failure-prediction'],,2016 IEEE Second International Conference on Multimedia Big Data (BigMM),True,['failure-management'],,True,,,,,,,
254,Experimental Study on Software Fault Prediction Using Machine Learning Model,"['T. M. Phuong Ha', ' D. Hung Tran', ' L. T. My Hanh', ' N. Thanh Binh']",2019,"Faults are the leading cause of time consuming and cost wasting during software life cycle. Predicting faults in early stage improves the quality and reliability of the system and also reduces cost for software development. Many researches proved that software metrics are effective elements for software fault prediction. In addition, many machine learning techniques have been developed for software fault prediction. It is important to determine which set of metrics are effective for predicting fault by using machine learning techniques. In this paper, we conduct an experimental study to evaluate the performance of seven popular techniques including Logistic Regression, K-nearest Neighbors, Decision Tree, Random Forest, Naïve Bayes, Support Vector Machine and Multilayer Perceptron using software metrics from Promise repository dataset usage. Our experiment is performed on both method-level and class-level datasets. The experimental results show that Support Vector Machine archives a higher performance in class-level datasets and Multilayer Perception produces a better accuracy in method-level datasets among seven techniques above.",https://ieeexplore.ieee.org/document/8919429,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 166}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 168}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 7}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 55}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 29}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 35}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 17}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 30}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 8}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 31}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 85}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 196}]",0.0,"['random-forest', 'naive-bayes', 'logistic-regression', 'similarity-matching', 'multilayer-perceptron', 'support-vector-machine', 'decision-tree', 'clustering']",['code-metrics'],"['comparison', 'novel-use']",['failure-prevention'],['source-code'],2019 11th International Conference on Knowledge and Systems Engineering (KSE),True,['failure-management'],,True,['software-defect-prediction'],,,,,,
255,Device-Agnostic Log Anomaly Classification with Partial Labels,"['W. Meng', ' Y. Liu', ' S. Zhang', ' D. Pei', ' H. Dong', ' L. Song', ' X. Luo']",2018,"Anomaly classification, i.e., detecting whether a network device is anomalous and determining its anomaly category if yes, plays a crucial role in troubleshooting. Compared to KPI curves, device logs contain too much more valuable information for anomaly classification. However, the regular expression based anomaly classification techniques cannot tackle the challenges lying in log anomaly classification. We propose LogClass, a data-driven framework to detect and classify anomalies based on device logs. LogClass combines a word representation method and the PU learning model to construct device-agnostic vocabulary with partial labels. We evaluate LogClass on tens of millions of switch logs collected from several real-world datacenters owned by a top global search engine. Our results show that LogClass achieves 99.515% F1 score in anomalous log detection, 95.32% Macro-F1 and 99.74% Micro-F1 in anomalous log classification in a computationally efficient manner.",https://ieeexplore.ieee.org/document/8624141,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 211}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 62}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 19}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 53}]",6.0,,,,['failure-detection'],,2018 IEEE/ACM 26th International Symposium on Quality of Service (IWQoS),True,['failure-management'],,,,,,,,,
256,Application of genetic algorithm as feature selection technique in development of effective fault prediction model,"['L. Kumar', ' S. K. Rath']",2016,"Prediction of faults in a proposed software is helpful in deciding the amount of effort to be given for software development. We observed that, a good number of authors hypothesized that the performance of fault prediction model depends on the source code metrics which are used as input of the model. Feature selection technique is a process of selecting suitable set of source code metrics which may improve the performance of fault prediction model. In this work, genetic algorithm (GA) has been applied as feature selection technique to select the suitable set of source code metrics. This selected set of source code metrics are used as requisite input data to develop a classifier using five different classification techniques such as logistic regression, extreme learning machine, support vector machine (SVM) with three different kernel functions (linear, polynomial, and radial basis kernel functions) in order to predict the faulty and non-faulty classes. In this study, we propose a cost evaluation framework to perform cost based analysis for evaluating the effectiveness of fault prediction model. We perform experiments on thirty number of Java Open Source projects. From the obtained results, it is observed that the model developed using selected set of source code metrics obtained better result as compared to all metrics. From costs analysis framework, it is observed that the developed fault prediction model is best suitable for software with % of faulty classes less than the threshold value depending on fault identification efficiency (low-46.44%, median-45.37%, and high-36.63%).",https://ieeexplore.ieee.org/document/7894693,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 276}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 375}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 12}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 92}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 50}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 73}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 36}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 34}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 105}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 331}]",4.0,"['genetic-programming', 'support-vector-machine', 'logistic-regression']",,"['new-method', 'novel-use']",['failure-prevention'],,"2016 IEEE Uttar Pradesh Section International Conference on Electrical, Computer and Electronics Engineering (UPCON)",True,['failure-management'],,True,['software-defect-prediction'],,,,,,
257,Software defect prediction using semi-supervised learning with dimension reduction,"['H. Lu', ' B. Cukic', ' M. Culp']",2012,"Accurate detection of fault prone modules offers the path to high quality software products while minimizing non essential assurance expenditures. This type of quality modeling requires the availability of software modules with known fault content developed in similar environment. Establishing whether a module contains a fault or not can be expensive. The basic idea behind semi-supervised learning is to learn from a small number of software modules with known fault content and supplement model training with modules for which the fault information is not available. In this study, we investigate the performance of semi-supervised learning for software fault prediction. A preprocessing strategy, multidimensional scaling, is embedded in the approach to reduce the dimensional complexity of software metrics. Our results show that the semi-supervised learning algorithm with dimension-reduction preforms significantly better than one of the best performing supervised learning algorithms, random forest, in situations when few modules with known fault content are available for training.",https://ieeexplore.ieee.org/document/6494944,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 408}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 27}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 134}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 71}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 657}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 9}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 77}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 90}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 51}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 47}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 15}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prediction' OR 'failure prediction')"", 'index': 25}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 21}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 23}]",27.0,,,,['failure-prediction'],,2012 Proceedings of the 27th IEEE/ACM International Conference on Automated Software Engineering,True,['failure-management'],,True,,,,10.1145/2351676.2351734,Conference Paper,"['dimension reduction', 'software metrics', 'Software fault prediction', 'semi-supervised learning']",
258,Towards Interpretable Defect-Prone Component Analysis Using Genetic Fuzzy Systems,"['T. Diamantopoulos', ' A. Symeonidis']",2015,"The problem of Software Reliability Prediction is attracting the attention of several researchers during the last few years. Various classification techniques are proposed in current literature which involve the use of metrics drawn from version control systems in order to classify software components as defect-prone or defect-free. In this paper, we create a novel genetic fuzzy rule-based system to efficiently model the defect-proneness of each component. The system uses a Mamdani-Assilian inference engine and models the problem as a one-class classification task. System rules are constructed using a genetic algorithm, where each chromosome represents a rule base (Pittsburgh approach). The parameters of our fuzzy system and the operators of the genetic algorithm are designed with regard to producing interpretable output. Thus, the output offers not only effective classification, but also a comprehensive set of rules that can be easily visualized to extract useful conclusions about the metrics of the software.",https://ieeexplore.ieee.org/document/7168329,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 475}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 464}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 168}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 185}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prediction' OR 'failure prediction')"", 'index': 60}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1396}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 65}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 42}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 62}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 14}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prediction' OR 'failure prediction')"", 'index': 35}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 66}]",0.0,,,,['failure-prediction'],,2015 IEEE/ACM 4th International Workshop on Realizing Artificial Intelligence Synergies in Software Engineering,True,['failure-management'],,,,,,,Conference Paper,"['software reliability prediction', 'genetic fuzzy systems', 'software fault prediction', 'defect-prone components']",
259,Comparative Study on Defect Prediction Algorithms of Supervised Learning Software Based on Imbalanced Classification Data Sets,"['J. Ge', ' J. Liu', ' W. Liu']",2018,"With the development of high complexity and high integration of software systems, the quality of software has gradually received widespread attention in scientific research and engineering. Software defect prediction technology plays an important role in improving software quality, reducing software development time, and reducing testing expenses. It has also become one of the hot issues in the field of software engineering research in recent years. As an important learning method in machine learning, supervised learning is widely used in the classification and regression prediction with annotation data because of its high accuracy, mature theory, and simple calculation. However, the imbalanced classification problem of data sets is common in practical applications and seriously affects the performance of learning algorithm. This paper analyzes the characteristics of software forecasting technology from the perspective of supervised learning, and performs balance like processing on imbalanced classification NASA data sets (JM1, KC3, MC1). NASA data sets after application processing are used to conduct experiments on LWL, C4.5, Random forest, Bagging, Bayesian Belief Network, Multilayer Feed forward Neural Network, SVM and NB-K algorithms, and experimental data is analyzed and evaluated.",https://ieeexplore.ieee.org/document/8441143,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 553}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 99}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 694}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 621}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1069}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 409}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 173}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1127}]",1.0,,,,['failure-prediction'],,"2018 19th IEEE/ACIS International Conference on Software Engineering, Artificial Intelligence, Networking and Parallel/Distributed Computing (SNPD)",True,['failure-management'],,,,,,,,,
260,Fault prediction model for software using soft computing techniques,"['I. U. Nisa', ' S. N. Ahsan']",2015,"Faulty modules of any software can be problematic in terms of accuracy, hence may encounter more costly redevelopment efforts in later phases. These problems could be addressed by incorporating the ability of accurate prediction of fault prone modules in the development process. Such ability of the software enables developers to reduce the faults in the whole life cycle of software development, at the same time it benefits automation process, and reduces the overall cost and efforts of the software maintenance. In this paper, we propose to design fault prediction model by using a set of code and design metrics; applying various machine learning (ML) classifiers; also used transformation techniques for feature reduction and dealing class imbalance data to improve fault prediction model. The data sets were obtained from publicly available PROMISE repositories. The results of the study revealed that there was no significant impact on the ability to accurately predict the fault-proneness of modules by applying PCA in reducing the dimensions; the results were improved after balancing data by SMOTE, Resample techniques, and by applying PCA with Resample in combination. It has also been seen that Random Forest, Random Tree, Logistic Regression, and Kstar machine learning classifiers have relatively better consistency in prediction accuracy as compared to other techniques.",https://ieeexplore.ieee.org/document/7396406,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 571}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 410}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 8}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 56}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 343}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 37}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 18}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 60}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 30}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 445}]",1.0,"['dimensionality-reduction', 'random-forest', 'decision-tree']",['software-metrics'],,['failure-prediction'],['software'],2015 International Conference on Open Source Systems & Technologies (ICOSST),True,['failure-management'],,,,,,,,,
261,A Learning-to-Rank Approach to Software Defect Prediction,"['X. Yang', ' K. Tang', ' X. Yao']",2015,"Software defect prediction can help to allocate testing resources efficiently through ranking software modules according to their defects. Existing software defect prediction models that are optimized to predict explicitly the number of defects in a software module might fail to give an accurate order because it is very difficult to predict the exact number of defects in a software module due to noisy data. This paper introduces a learning-to-rank approach to construct software defect prediction models by directly optimizing the ranking performance. In this paper, we build on our previous work, and further study whether the idea of directly optimizing the model performance measure can benefit software defect prediction model construction. The work includes two aspects: one is a novel application of the learning-to-rank approach to real-world data sets for software defect prediction, and the other is a comprehensive evaluation and comparison of the learning-to-rank method against other algorithms that have been used for predicting the order of software modules according to the predicted number of defects. Our empirical studies demonstrate the effectiveness of directly optimizing the model performance measure for the learning-to-rank approach to construct defect prediction models for the ranking task.",https://ieeexplore.ieee.org/document/6996020,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 681}]",83.0,"['entropy-selection', 'linear-regression']","['code-metrics', 'code-history']","['new-method', 'comparison']",['failure-prevention'],['source-code'],IEEE Transactions on Reliability,True,['failure-management'],True,,['software-defect-prediction'],True,11.0,,,,
262,Software Defect Prediction Based on Fourier Learning,"['K. Yang', ' H. Yu', ' G. Fan', ' X. Yang', ' S. Zheng', ' C. Leng']",2018,"Modern software systems have grown significantly in their size and complexity, therefore software systems have more and more potential defects. Software defect prediction uses a defect data set to build a predictive model, where the data set is composed of software defect metrics. Then, this predictive model is used to predict potential defect program modules in the project. This paper uses the Fourier expression of Boolean function to build a software defect prediction model. We provide the algorithms to calculate the Fourier coefficients and get the predicted function which can predict software defect. And, we compare the Fourier learning algorithm with the traditional machine learning algorithms, such as the random forest algorithm. Finally, the experiment results show that the Fourier learning algorithm is not only better than other algorithms, but also more stable.",https://ieeexplore.ieee.org/document/8706304,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 738}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 470}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1133}]",0.0,,,,['failure-prediction'],,2018 IEEE International Conference on Progress in Informatics and Computing (PIC),True,['failure-management'],,,,,,,,,
263,Network Fault Diagnosis Using Hierarchical SVMs Based on Kernel Method,"['L. Zhang', ' X. Meng', ' H. Zhou']",2009,"A new method based on kernel which can measure class separability in feature space is proposed in this paper for existing error accumulation when the hierarchical SVMs is used to diagnose multiclass network fault. This method has defined metrics of sample distribution in feature space, which are used as the rule of constructing hierarchical SVMs. Experiment results show that this method can restrain error accumulation and has higher multiclass classification accuracy, and offer an effective way for network fault diagnosis.",https://ieeexplore.ieee.org/document/4772045,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 824}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1423}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 110}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1621}]",1.0,['support-vector-machine'],['metrics'],,['root-cause-analysis'],['network'],2009 Second International Workshop on Knowledge Discovery and Data Mining,True,['failure-management'],,,,,,,,,
264,Software aging analysis and prediction in a web server based on multiple linear regression algorithm,"['S. Jia', ' C. Hou', ' J. Wang']",2017,"In the last few years, software aging has been reported. The phenomenon of software aging is that the performance of the software system is degradation in a long-running state, which is the result of exhaustion of system resources, the accumulation of internal error conditions and so on. In order to counteract software aging, a technique, which called software rejuvenation, has been proposed, this procedure involves occasionally stopping a system process, cleaning its running environment and restarting it. Due to the direct and indirect costs incurred by software rejuvenation, when to carry out this action is very important. Traditionally, most scholars focused on time series or analytic methods to model software aging process, however, machine learning algorithm has been neglected. In this paper, we make a detailed analysis and predict about the web server parameters by multiple linear regression algorithm. Firstly, the aging phenomenon of the system is simulated by the pressure testing tool and then collecting data and preprocessed. Secondly, we fit time series models to the data collected and determine the trend of resource consumption. Thirdly, using the feature selection algorithm to select a subset set as the input parameters of the algorithm. Fourthly, using the multiple linear regression algorithm to analysis and predict the aging process. Finally, we evaluate the feasibility of the algorithm by evaluation metrics. The result shows that we can use this algorithm to predict the aging process in the allowable error range.",https://ieeexplore.ieee.org/document/8230349,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 845}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 100}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 805}]",1.0,,,,['failure-prediction'],,2017 IEEE 9th International Conference on Communication Software and Networks (ICCSN),True,['failure-management'],,,,,,,,,
265,Predicting Software Anomalies Using Machine Learning Techniques,"['J. Alonso', ' L. Belanche', ' D. R. Avresky']",2011,"In this paper, we present a detailed evaluation of a set of well-known Machine Learning classifiers in front of dynamic and non-deterministic software anomalies. The system state prediction is based on monitoring system metrics. This allows software proactive rejuvenation to be triggered automatically. Random Forest approach achieves validation errors less than 1% in comparison to the well-known ML algorithms under a valuation. In order to reduce automatically the number of monitored parameters, needed to predict software anomalies, we analyze Lasso Regularization technique jointly with the Machine Learning classifiers to evaluate how the prediction accuracy could be guaranteed within an acceptable threshold. This allows to reduce drastically (around 60% in the best case) the number of monitoring parameters. The framework, based on ML and Lasso regularization techniques, has been validated using an ecommerce environment with Apache Tomcat server, and MySql database server.",https://ieeexplore.ieee.org/document/6038598,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1116}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 147}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1744}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 669}]",24.0,,,,['failure-prediction'],,2011 IEEE 10th International Symposium on Network Computing and Applications,True,['failure-management'],,,,,,,,,
266,HotSpot: Anomaly Localization for Additive KPIs With Multi-Dimensional Attributes,"['Y. Sun', ' Y. Zhao', ' Y. Su', ' D. Liu', ' X. Nie', ' Y. Meng', ' S. Cheng', ' D. Pei', ' S. Zhang', ' X. Qu', ' X. Guo']",2018,"Additive key performance indicators (KPIs) (such as page view (PV), revenue, and error count) with multi-dimensional attributes (such as ISP, Province, and DataCenter) are common and important in monitoring metrics in Internet companies. When an anomaly happens to an overall KPI, it is critical but challenging to localize the root cause, which is one (or more) combination of attribute values in multiple dimensions. For example, is the total PV decrease caused by the PV decrease from “Beijing”or “China Mobile in Beijing”, or “Beijing and Shanghai”? However, this task is very challenging for two major reasons. First, the PVs of different combinations are interdependent; thus, the PV anomalies at the root cause can cause the changes of many other PVs at different aggregation levels. Second, there could be tens of thousands of combinations to investigate in multi-dimensional attribute space. It is a difficulty to find the root cause from a huge search space. To address the first challenge, our approach HotSpot uses a novel potential score based on the ripple effect for anomaly propagation that we reveal. To address the second challenge, HotSpot adopts the Monte Carlo Tree Search algorithm and a hierarchical pruning strategy. Using the real-world data from a top global search engine, we show that HotSpot achieves a great improvement on effectiveness and robustness, i.e., 95% of all types of root cause cases using HotSpot (compared with only 15% using existing approaches) achieves an F-score over 90%. Operational experiences show that HotSpot can reduce the localization time from more than 1 h in manual efforts to less than 20 s.",https://ieeexplore.ieee.org/document/8288614,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1174}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 481}]",9.0,,,,['failure-detection'],,,True,['failure-management'],,,,,,,,,
267,A Novel Approach for Software Defect Prediction Using Fuzzy Decision Trees,"['Z. Marian', ' I. Mircea', ' I. Czibula', ' G. Czibula']",2016,"Detecting defective entities from existing software systems is a problem of great importance for increasing both the software quality and the efficiency of software testing related activities. We introduce in this paper a novel approach for predicting software defects using fuzzy decision trees. Through the fuzzy approach we aim to better cope with noise and imprecise information. A fuzzy decision tree will be trained to identify if a software module is or not a defective one. Two open source software systems are used for experimentally evaluating our approach. The obtained results highlight that the fuzzy decision tree approach outperforms the non-fuzzy one on almost all case studies used for evaluation. Compared to the approaches used in the literature, the fuzzy decision tree classifier is shown to be more efficient than most of the other machine learning-based classifiers.",https://ieeexplore.ieee.org/document/7829618,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1313}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 869}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 741}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1877}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 202}]",3.0,"['fuzzy-logic', 'decision-tree']",,,['failure-prevention'],,2016 18th International Symposium on Symbolic and Numeric Algorithms for Scientific Computing (SYNASC),True,['failure-management'],,True,['software-defect-prediction'],,,,,,
268,Reducing Web Latency Through Dynamically Setting TCP Initial Window with Reinforcement Learning,"['X. Nie', ' Y. Zhao', ' D. Pei', ' G. Chen', ' K. Sui', ' J. Zhang']",2018,"Latency, which directly affects the user experience and revenue of web services, is far from ideal in reality, due to the well-known TCP flow startup problem. Specifically, since TCP starts from a conservative and static initial window (IW, 2~4 or 10), most of the web flows are too short to have enough time to find its best congestion window before the session ends. As a result, TCP cannot fully utilize the available bandwidth in the modern Internet. In this paper, we propose to use group-based reinforcement learning (RL) to enable a web server, through trial-and-error, to dynamically set a suitable IW for a web flow before its transmission starts. Our proposed system, SmartIW, collects TCP flow performance metrics (e.g., transmission time, loss rate, RTT) in real-time without any client assistance. Then these metrics are aggregated into groups with similar features (subnet, ISP, province, etc.) to satisfy RL's requirement. SmartIW has been deployed in one of the top global search engines for more than a year. Our online and testbed experiments show that, compared to the common practice of IW=10, SmartIW can reduce the average transmission time by 23% to 29%.",https://ieeexplore.ieee.org/document/8624175,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1452}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 152}]",2.0,,,,,,2018 IEEE/ACM 26th International Symposium on Quality of Service (IWQoS),True,['resource-provisioning'],,,,,,,,,
269,Software Bug Prediction Using Supervised Machine Learning Algorithms,"['S. Delphine Immaculate', ' M. Farida Begam', ' M. Floramary']",2019,"Machine Learning algorithms sprawl their application in various fields relentlessly. Software Engineering is not exempted from that. Software bug prediction at the initial stages of software development improves the important aspects such as software quality, reliability, and efficiency and minimizes the development cost. In majority of software projects which are becoming increasingly large and complex programs, bugs are serious challenge for system consistency and efficiency. In our approach, three supervised machine learning algorithms are considered to build the model and predict the occurrence of the software bugs based on historical data by deploying the classifiers Logistic regression, Naive Bayes, and Decision Tree. Historical data has been used to predict the future software faults by deploying the classifier algorithms and make the models a better choice for predictions using random forest ensemble classifiers and validating the models with K-Fold cross validation technique which results in the model effectively working for all the scenarios.",https://ieeexplore.ieee.org/document/8816965,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1507}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 137}, {'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 722}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 159}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 98}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 924}]",1.0,"['naive-bayes', 'logistic-regression', 'decision-tree']",['code-history'],,['failure-prevention'],['source-code'],2019 International Conference on Data Science and Communication (IconDSC),True,['failure-management'],,True,['software-defect-prediction'],,,,,,
270,Wavelet-based multi-scale anomaly identification in cloud computing systems,"['Q. Guan', ' S. Fu']",2013,"Modern cloud computing systems contain thousands of computing and storage servers. Such a scale combined with ever-growing system complexity of their components and interactions, introduces a key challenge to failure and resource management for highly dependable cloud computing. Automated anomaly detection is a crucial technique for understanding emergent, cloud-wide phenomena and self-managing cloud resources for system level dependability assurance. In this paper, we present a wavelet-based multi-scale anomaly identification mechanism, that can analyze profiled cloud performance metrics in both time and frequency domains and identify anomalous cloud behaviors. Learning technologies are exploited to adapt the selection of mother wavelets and a sliding detection window is employed handle cloud dynamicity and improve anomaly detection accuracy. We test a prototype implementation of our cloud anomaly detection mechanism on an institute-wide cloud system. Experimental results show our approach can identify cloud failures accurately.",https://ieeexplore.ieee.org/document/6831266,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1928}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 324}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 564}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 899}]",3.0,,,,['failure-detection'],,2013 IEEE Global Communications Conference (GLOBECOM),True,['failure-management'],,,,,,,,,
271,Automatic Generation of Workload Profiles Using Unsupervised Learning Pipelines,"['D. B. Prats', ' J. L. Berral', ' D. Carrera']",2018,"The complexity of resource usage and power consumption on cloud-based applications makes the understanding of application behavior through expert examination difficult. The difficulty increases when applications are seen as “black boxes,” where only external monitoring can be retrieved. Furthermore, given the different amount of scenarios and applications, automation is required. Here, we examine and model application behavior by finding behavior phases. We use conditional restricted Boltzmann machines (CRBMs) to model time-series containing resources traces measurements like CPU, memory, and IO. CRBMs can be used to map a given historic window of trace behavior into a single vector. This low dimensional and time-aware vector can be passed through clustering methods, from simplistic ones like k-means to more complex ones like those based on hidden Markov models. We use these methods to find phases of similar behavior in the workloads. Our experimental evaluation shows that the proposed method is able to identify different phases of resource consumption across different workloads. We show that the distinct phases contain specific resource patterns that distinguish them.",https://ieeexplore.ieee.org/document/8240924,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 61}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 1044}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 116}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 298}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 165}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 481}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 120}]",0.0,['boltzmann-machine'],['host-metrics'],,['anomaly-detection'],,IEEE Transactions on Network and Service Management,True,['resource-provisioning'],,True,,,,,,,
272,DReAM: Deep Recursive Attentive Model for Anomaly Detection in Kernel Events,"['O. M. Ezeme', ' Q. H. Mahmoud', ' A. Azim']",2019,"System logs and traces contain information that reflects the state of the system and serves as a rich source of knowledge for system monitoring from the application to the kernel layer. Moreover, logging of traces as a tool for monitoring the operation of a cyber-physical system is recommended by most safety standard organizations. However, because the data can be overwhelmingly huge within a short space of time, the use of models that do not rely only on known signatures for online anomaly detection becomes difficult to use due to the challenge of processing such an enormous amount of data at runtime. Hence, most practitioners resort to the use of signature-based tools. In this paper, we introduce an anomaly detection model that uses intra-trace and inter-trace context vectors with long short-term memory networks to overcome the challenge of online anomaly detection in cyber-physical systems. We test the performance of the model with publicly available datasets that reflect the internal and external control flow of an embedded application and our model demonstrates both the effectiveness and robustness in detecting an anomalous sequence in a system call stream.",https://ieeexplore.ieee.org/document/8633836,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 278}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 623}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 778}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 291}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 330}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1457}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 182}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 320}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 259}]",3.0,['rnn'],"['events', 'logs']",['new-method'],['failure-detection'],['kernel'],,True,['failure-management'],,True,,,,,,,
273,Failure Root Cause Analysis Automation on Functional Simulation Regressions,['C. J. Yen'],2019,"Performing a functional simulation on the developed test bench and design codes is the common practice to secure the design behavior and quality. It is non-trivial to launch a large functional simulation regression and deal with many failures. Triage, the process of deciding which failures to debug first, can itself be time consuming. Practical considerations such as disk space, run time, and computation resources also come in the picture. Given the relatively low throughput of debugging a failure, it is always not the best solution to automatically re-run all failed tests, have waveform dumping, and debug the error traces one by one. In this paper, we introduce an automatic approach to analyze the root cause of regression failures. The approach enabled several machine learning methods to speed up the triage process. In addition, by combining the concept of binning and the self-propelling root cause tracing process, it can suggest probable sources of the problem. Engineers can thus benefit from checking a smaller set of failures, narrowing down the debugging scope, and justifying the fixes with the same failure bin. This results in a significant productivity boost for the verification efforts.",https://ieeexplore.ieee.org/document/8741656,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 456}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 3}, {'database': 'IEEE', 'search_string': ""'regression' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 53}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 204}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 5}, {'database': 'IEEE', 'search_string': ""'regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 1}]",0.0,['clustering'],['logs'],,['root-cause-analysis'],,"2019 International Symposium on VLSI Design, Automation and Test (VLSI-DAT)",True,['failure-management'],,True,['rca-others'],,,,,,
274,SmartSLA: Cost-Sensitive Management of Virtualized Resources for CPU-Bound Database Services,"['P. Xiong', ' Y. Chi', ' S. Zhu', ' H. J. Moon', ' C. Pu', ' H. Hacgümüş']",2015,"Virtualization-based multi-tenant database consolidation is an important technique for database-as-a-service (DBaaS) providers to minimize their total cost which is composed of SLA penalty cost, infrastructure cost and action cost. Due to the bursty and diverse tenant workloads, over-provisioning for the peak or under-provisioning for the off-peak often results in either infrastructure cost or service level agreement (SLA) penalty cost. Moreover, although the process of scaling out database systems will help DBaaS providers satisfy tenants' service level agreement, its indiscriminate use has performance implications or incurs action cost. In this paper, we propose SmartSLA, a cost-sensitive virtualized resource management system for CPU-bound database services which is composed of two modules. The system modeling module uses machine learning techniques to learn a model for predicting the SLA penalty cost for each tenant under different resource allocations. Based on the learned model, the resource allocating module dynamically adjusts the resource allocation by weighing the potential reduction of SLA penalty cost against increase of infrastructure cost and action cost. SmartSLA is evaluated by using the TPC-W and modified YCSB benchmarks with dynamic workload trace and multiple database tenants. The experimental results show that SmartSLA is able to minimize the total cost under time-varying workloads compared to the other cost-insensitive approaches.",https://ieeexplore.ieee.org/document/6803071,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 981}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 620}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1693}]",10.0,,,,,,IEEE Transactions on Parallel and Distributed Systems,True,['resource-provisioning'],,,,,,,,,
275,Dynamic Meta-Learning for Failure Prediction in Large-Scale Systems: A Case Study,"['J. Gu', ' Z. Zheng', ' Z. Lan', ' J. White', ' E. Hocks', ' B. Park']",2008,"Despite great efforts on the design of ultra-reliable components, the increase of system size and complexity has outpaced the improvement of component reliability. As a result, fault management becomes crucial in high performance computing. The advance of fault management relies on effective failure prediction. Despite years of research on failure prediction, it remains an open problem, especially in large-scale systems. In this paper, we address the problem by presenting a dynamic meta-learning prediction engine. It extends our previous work by exploring dynamic training, testing and prediction. Here, the ""dynamic"" part is from two perspectives: one is to continuously increase the training set during the system operation; and the other is to dynamically modify the rules of failure patterns by tracing prediction accuracy at runtime. Our case study indicates that the proposed predictor is promising by being capable of capturing more than 70% of failures, with the false alarm rate less than 10%.",https://ieeexplore.ieee.org/document/4625845,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1337}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 60}]",33.0,,,,['failure-prediction'],,2008 37th International Conference on Parallel Processing,True,['failure-management'],,,,,,,,,
276,SENATUS: An Approach to Joint Traffic Anomaly Detection and Root Cause Analysis,"['A. Abdelkefi', ' Y. Jiang', ' S. Sharma']",2018,"In this paper, we propose a novel approach, called SENATUS, for joint anomaly detection and root-cause analysis. Inspired from the concept of a senate, the key idea of the proposed approach is divided into three stages: election, voting and decision. At the election stage, a small number of traffic flow sets (termed as senator flows) are chosen based on the K-sparse approximation technique, which can be used to represent approximately the total (usually huge) set of traffic flows. In the voting stage, Principal Component Pursuit (PCP) analysis is used for anomaly detection on the senator flows. In addition, the detected anomalies are correlated across traffic features to identify the most possible anomalous time bins. Finally, in the decision stage, a machine learning (ML) technique is applied to the senator flows of anomalous time bins to find the root cause of the anomalies. The performance of SENATUS is evaluated using real traffic traces collected from a Pan European network, GEANT, and compared against another approach which detects anomalies using lossless compression of traffic histograms. The evaluation shows that SENATUS has higher effectiveness in diagnosing traffic anomalies.",https://ieeexplore.ieee.org/document/8602689,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1374}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 10}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 766}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 103}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 11}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 274}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 702}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 16}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 275}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 283}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 6}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 370}]",0.0,,,,['failure-detection'],,2018 2nd Cyber Security in Networking Conference (CSNet),True,['failure-management'],,True,,,,,,,
277,Online System Problem Detection by Mining Patterns of Console Logs,"['W. Xu', ' L. Huang', ' A. Fox', ' D. Patterson', ' M. Jordan']",2009,"We describe a novel application of using data mining and statistical learning methods to automatically monitor and detect abnormal execution traces from console logs in an online setting. Different from existing solutions, we use a two stage detection system. The first stage uses frequent pattern mining and distribution estimation techniques to capture the dominant patterns (both frequent sequences and time duration). The second stage use principal component analysis based anomaly detection technique to identify actual problems. Using real system data from a 203-node Hadoop cluster, we show that we can not only achieve highly accurate and fast problem detection, but also help operators better understand execution patterns in their system.",https://ieeexplore.ieee.org/document/5360285,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1473}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 88}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 86}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1122}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 778}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1509}]",58.0,,,,['failure-detection'],,2009 Ninth IEEE International Conference on Data Mining,True,['failure-management'],,,,,,,,,
278,Job failure prediction in grid environment based on workload characteristics,"['H. Fadishei', ' H. Saadatfar', ' H. Deldari']",2009,"The power of grid technology in aggregating autonomous resources owned by several organizations into a single virtual system has made it popular in compute-intensive and data-intensive applications. Complex and dynamic nature of grid makes failure of users' jobs fairly probable. Furthermore, traditional methods for job failure recovery have proven costly and thus a need to shift toward proactive and predictive management strategies is necessary in such systems. In this paper, an innovative effort is made to predict the futurity of jobs submitted to a production grid environment (AuverGrid). By analyzing grid workload traces and extracting patterns describing common failure characteristics, the success or failure status of jobs during 6 months of AuverGrid activity was predicted with around 96% accuracy. The quality of services on grid can be improved by integrating the result of this work into management services like scheduling and monitoring.",https://ieeexplore.ieee.org/document/5349381,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1494}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 835}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 120}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 117}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 781}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 1824}]",15.0,,,,['failure-prediction'],,2009 14th International CSI Computer Conference,True,['failure-management'],,,,,,,,,
279,A novel weighted combination technique for traffic classification,"['J. Yan', ' X. Yun', ' Z. Wu', ' H. Luo', ' S. Zhang']",2012,"Accurate classification of traffic flows is highly beneficial for network management and security monitoring. Nowadays, many researchers have proposed machine learning techniques (i.e., decision tree, SVM, BayesNet and Naïve Bayes) for traffic classification. However, none of these classification techniques can achieve the highest accuracy for all traffic classification tasks. Recently, more and more researchers tried to combine multiple classifiers to obtain better performance. In this paper, we propose a weighted combination technique for traffic classification. The weighted combination approach first takes advantage of the confidence values inferred by each individual classifier; then assigns weight for each classifier according to its prediction accuracy on a validation traffic dataset. Experimental results on two different traffic traces demonstrate that our new weighted multi-classification framework is able to obtain satisfactory results.",https://ieeexplore.ieee.org/document/6664277,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1506}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 1182}, {'database': 'IEEE', 'search_string': ""'classification' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 168}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 603}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 224}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 200}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1060}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 610}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 336}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 834}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1472}]",1.0,"['similarity-matching', 'bayesian-network', 'decision-tree']",['traces'],,['failure-detection'],['network'],2012 IEEE 2nd International Conference on Cloud Computing and Intelligence Systems,True,['failure-management'],,,,,,,,,
280,PRESS: PRedictive Elastic ReSource Scaling for cloud systems,"['Zhenhuan Gong', ' Xiaohui Gu', ' J. Wilkes']",2010,"Cloud systems require elastic resource allocation to minimize resource provisioning costs while meeting service level objectives (SLOs). In this paper, we present a novel PRedictive Elastic reSource Scaling (PRESS) scheme for cloud systems. PRESS unobtrusively extracts fine-grained dynamic patterns in application resource demands and adjust their resource allocations automatically. Our approach leverages light-weight signal processing and statistical learning algorithms to achieve online predictions of dynamic application resource requirements. We have implemented the PRESS system on Xen and tested it using RUBiS and an application load trace from Google. Our experiments show that we can achieve good resource prediction accuracy with less than 5% over-estimation error and near zero under-estimation error, and elastic resource scaling can both significantly reduce resource waste and SLO violations.",https://ieeexplore.ieee.org/document/5691343,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1608}]",146.0,"['autoregression', 'markov-model']",['host-metrics'],,"['resource-consolidation', 'workload-prediction']",,2010 International Conference on Network and Service Management,True,['resource-provisioning'],,,,,,,,,
281,Dynamic Analysis for Diagnosing Integration Faults,"['L. Mariani', ' F. Pastore', ' M. Pezze']",2011,"Many software components are provided with incomplete specifications and little access to the source code. Reusing such gray-box components can result in integration faults that can be difficult to diagnose and locate. In this paper, we present Behavior Capture and Test (BCT), a technique that uses dynamic analysis to automatically identify the causes of failures and locate the related faults. BCT augments dynamic analysis techniques with model-based monitoring. In this way, BCT identifies a structured set of interactions and data values that are likely related to failures (failure causes), and indicates the components and the operations that are likely responsible for failures (fault locations). BCT advances scientific knowledge in several ways. It combines classic dynamic analysis with incremental finite state generation techniques to produce dynamic models that capture complementary aspects of component interactions. It uses an effective technique to filter false positives to reduce the effort of the analysis of the produced data. It defines a strategy to extract information about likely causes of failures by automatically ranking and relating the detected anomalies so that developers can focus their attention on the faults. The effectiveness of BCT depends on the quality of the dynamic models extracted from the program. BCT is particularly effective when the test cases sample the execution space well. In this paper, we present a set of case studies that illustrate the adequacy of BCT to analyze both regression testing failures and rare field failures. The results show that BCT automatically filters out most of the false alarms and provides useful information to understand the causes of failures in 69 percent of the case studies.",https://ieeexplore.ieee.org/document/5611554,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('fault localization' OR 'failure localization')"", 'index': 28}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 154}]",44.0,,,,['root-cause-analysis'],,IEEE Transactions on Software Engineering,True,['failure-management'],,,,,,,,,
282,A Function Clustering Algorithm for Resource Utilization in Service Function Chaining,"['H. Kanemitsu', ' K. Kanai', ' J. Katto', ' H. Nakazato']",2019,"Virtualized service and network functions are deployed on virtual machines (VMs) to realize essential processing to realize service function chaining (SFC). Issues on SFC is SF allocation to a VM and to minimize the response time and number of function instances. In this paper, we propose an SF clustering-based scheduling algorithm, called ""SF-clustering for utilizing virtual CPUs"" (SFCUV), to solve the SF allocation and SF selection problems simultaneously. Experimental results show that SF-CUV can utilize vCPUs to minimize the response time.",https://ieeexplore.ieee.org/document/8814541,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 127}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 39}]",0.0,['clustering'],,['new-method'],['scheduling'],['vm'],2019 IEEE 12th International Conference on Cloud Computing (CLOUD),True,['resource-provisioning'],,,,,,,,,
283,Analysis and Clustering of Workload in Google Cluster Trace Based on Resource Usage,"['M. Alam', ' K. A. Shakil', ' S. Sethi']",2016,"Cloud computing has gained interest amongst commercial organizations, research communities, developers and other individuals during the past few years. In order to move ahead with research in field of data management and to enable processing of such data, we need benchmark datasets and freely available data which are publicly accessible. Google in May 2011 released a trace of a cluster of 11k machines referred as ""Google Cluster Trace"". This trace contains cell information of about 29 days. This paper provides analysis of resource usage and requirements in this trace and is an attempt to give an insight about such kind of production trace similar to the ones in cloud environment. The major contributions of this paper include statistical profile of jobs based on resource usage, clustering of workload patterns and classification of jobs into different types based on k-means clustering. Though there have been earlier works for analysis of this trace, but our analysis provides several new findings such as jobs in a production trace are trimodal and there occurs symmetry in the tasks within a long job type.",https://ieeexplore.ieee.org/document/7982333,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 170}, {'database': 'IEEE', 'search_string': ""'classification' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 114}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 476}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 1210}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 21}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 406}, {'database': 'arxiv', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 33}, {'database': 'arxiv', 'search_string': ""'classification' AND ('cloud')"", 'index': 271}, {'database': 'arxiv', 'search_string': ""'classification' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 5}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 274}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud')"", 'index': 772}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 2}]",3.0,,,,,,2016 IEEE Intl Conference on Computational Science and Engineering (CSE) and IEEE Intl Conference on Embedded and Ubiquitous Computing (EUC) and 15th Intl Symposium on Distributed Computing and Applications for Business Engineering (DCABES),True,['resource-provisioning'],,,,,,,,,
284,A Budget Constrained Scheduling Algorithm for Hybrid Cloud Computing Systems Under Data Privacy,"['A. Rezaeian', ' H. Abrishami', ' S. Abrishami', ' M. Naghibzadeh']",2016,"In hybrid cloud model, organizations can keep their sensitive information and critical applications in the private cloud and move other data and applications to a public cloud, if necessary. To maintain data privacy in workflow applications, we present a budget constrained hybrid cloud scheduler (BCHCS) which is a static heuristic scheduling algorithm. It is able to make decisions about scheduling sensitive tasks on private cloud and uses public cloud's resources for non-sensitive tasks, such that the makespan is minimized, while the budget limitation imposed by the user is satisfied. Experimental results show that the proposed method guarantees the execution of sensitive tasks on private cloud while achieving at least 7 percent lower makespan and higher success rate in comparison to similar existing techniques.",https://ieeexplore.ieee.org/document/7484195,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 261}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 57}]",4.0,,,,,,2016 IEEE International Conference on Cloud Engineering (IC2E),True,['resource-provisioning'],,,,,,,,,
285,An LOF-Based Adaptive Anomaly Detection Scheme for Cloud Computing,"['T. Huang', ' Y. Zhu', ' Q. Zhang', ' Y. Zhu', ' D. Wang', ' M. Qiu', ' L. Liu']",2013,"One of the most attractive things about cloud computing from the perspective of business people is that it provides an effective means to outsource IT. The behaviors of business applications on cloud are constantly evolving due to technical upgrading, cloud migration as well as social outbreaks. These changes bring the challenge of detecting anomalies during the change of applications on cloud. LOF (Local Outlier Factor) algorithm has already been proven as the most promising outlier detection method for detecting network intrusions. To improve the performance of detection, LOF needs a complete set of normal behaviors of business applications, which is usually not available in cloud computing. We present an adaptive anomaly detection scheme for cloud computing based on LOF. Our scheme learns behaviors of applications both in training and detecting phase. It is adaptive to the change during detecting. The adaptability of our scheme reduces demand of efforts on collecting training data before detecting. It also enables the ability to detect contextual anomalies. Experimental results show that our scheme can effectively detect contextual anomalies with relatively low computational overhead.",https://ieeexplore.ieee.org/document/6605790,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 415}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 712}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 953}]",9.0,,,,['failure-detection'],,2013 IEEE 37th Annual Computer Software and Applications Conference Workshops,True,['failure-management'],,,,,,,,,
286,Workload characterization and prediction in the cloud: A multiple time series approach,"['A. Khan', ' X. Yan', ' S. Tao', ' N. Anerousis']",2012,"Cloud computing promises high scalability, flexibility and cost-effectiveness to satisfy emerging computing requirements. To efficiently provision computing resources in the cloud, system administrators need the capabilities of characterizing and predicting workload on the Virtual Machines (VMs). In this paper, we use data traces obtained from a real data center to develop such capabilities. First, we search for repeatable workload patterns by exploring cross-VM workload correlations resulted from the dependencies among applications running on different VMs. Treating workload data samples as time series, we develop a co-clustering technique to identify groups of VMs that frequently exhibit correlated workload patterns, and also the time periods in which these VM groups are active. Then, we introduce a method based on Hidden Markov Modeling (HMM) to characterize the temporal correlations in the discovered VM clusters and to predict variations of workload patterns. The experimental results show that our method can not only help better understand group-level workload characteristics, but also make more accurate predictions on workload changes in a cloud.",https://ieeexplore.ieee.org/document/6212065,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 1022}, {'database': 'IEEE', 'search_string': ""'clustering' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 49}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 510}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 1225}]",246.0,['markov-model'],['traces'],['novel-use'],['workload-prediction'],['vvm'],2012 IEEE Network Operations and Management Symposium,True,['resource-provisioning'],,,,,,,,,
287,Detecting Abnormal Machine Characteristics in Cloud Infrastructures,"['K. Bhaduri', ' K. Das', ' B. L. Matthews']",2011,"In the cloud computing environment resources are accessed as services rather than as a product. Monitoring this system for performance is crucial because of typical pay-per- use packages bought by the users for their jobs. With the huge number of machines currently in the cloud system, it is often extremely difficult for system administrators to keep track of all machines using distributed monitoring programs such as Ganglia1 which lacks system health assessment and summarization capabilities. To overcome this problem, we propose a technique for automated anomaly detection using machine performance data in the cloud. Our algorithm is entirely distributed and runs locally on each computing machine on the cloud in order to rank the machines in order of their anomalous behavior for given jobs. There is no need to centralize any of the performance data for the analysis and at the end of the analysis, our algorithm generates error reports, thereby allowing the system administrators to take corrective actions. Experiments performed on real data sets collected for different jobs validate the fact that our algorithm has a low overhead for tracking anomalous machines in a cloud infrastructure.",https://ieeexplore.ieee.org/document/6137372,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 1095}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 1053}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 452}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 968}]",15.0,,,,['failure-detection'],,2011 IEEE 11th International Conference on Data Mining Workshops,True,['failure-management'],,,,,,,,,
288,A penalty-based grouping genetic algorithm for multiple composite SaaS components clustering in Cloud,"['Z. I. M. Yusoh', ' M. Tang']",2012,"Software as a Service (SaaS) in Cloud is getting more and more significant among software users and providers recently. A SaaS that is delivered as composite application has many benefits including reduced delivery costs, flexible offers of the SaaS functions and decreased subscription cost for users. However, this approach has introduced a new problem in managing the resources allocated to the composite SaaS. The resource allocation that has been done at the initial stage may be overloaded or wasted due to the dynamic environment of a Cloud. A typical data center resource management usually triggers a placement reconfiguration for the SaaS in order to maintain its performance as well as to minimize the resource used. Existing approaches for this problem often ignore the underlying dependencies between SaaS components. In addition, the reconfiguration also has to comply with SaaS constraints in terms of its resource requirements, placement requirement as well as its SLA. To tackle the problem, this paper proposes a penalty-based Grouping Genetic Algorithm for multiple composite SaaS components clustering in Cloud. The main objective is to minimize the resource used by the SaaS by clustering its component without violating any constraint. Experimental results demonstrate the feasibility and the scalability of the proposed algorithm.",https://ieeexplore.ieee.org/document/6377929,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 1110}, {'database': 'IEEE', 'search_string': ""'clustering' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 58}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 1030}]",3.0,,,,,,"2012 IEEE International Conference on Systems, Man, and Cybernetics (SMC)",True,['resource-provisioning'],,,,,,,,,
289,Towards autonomic workload provisioning for enterprise Grids and clouds,"['A. Quiroz', ' H. Kim', ' M. Parashar', ' N. Gnanasambandam', ' N. Sharma']",2009,"This paper explores autonomic approaches for optimizing provisioning for heterogeneous workloads on enterprise grids and clouds. Specifically, this paper presents a decentralized, robust online clustering approach that addresses the distributed nature of these environments, and can be used to detect patterns and trends, and use this information to optimize provisioning of virtual (VM) resources. It then presents a model-based approach for estimating application service time using long-term application performance monitoring, to provide feedback about the appropriateness of requested resources as well as the system's ability to meet QoS constraints and SLAs. Specifically for high-performance computing workloads, the use of a quadratic response surface model (QRSM) is justified with respect to traditional models, demonstrating the need for application-specific modeling. The proposed approaches are evaluated using a real computing center workload trace and the results demonstrate both their effectiveness and cost-efficiency.",https://ieeexplore.ieee.org/document/5353066,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 1158}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 663}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('cloud')"", 'index': 1904}]",100.0,,,,['resource-consolidation'],,2009 10th IEEE/ACM International Conference on Grid Computing,True,['resource-provisioning'],,,,,,,,,
290,DEARS: A Deep Learning Based Elastic and Automatic Resource Scheduling Framework for Cloud Applications,"['M. Hassan', ' H. Chen', ' Y. Liu']",2018,"Cloud computing paradigm supports more enterprises to provide satisfactory web services to their clients. However, the bursty and fluctuation of requests challenge the traditional resource scheduling framework. Previous strategies manage the jobs in each virtual machines (VMs) according to the derived historical utilization patterns, where the misalignment on the utilization curves may cause the resource over-prediction and over-provisioning. To better reduce the service latency and the above mentioned problem, we propose DEARS, a Deep learning based Elastic and Automatic Resource Scheduling framework for cloud applications. It gives a proactive and reactive strategy, where the LSTM model is pro-applied to predict the future request demand based on historical workload. The corresponding VM allocation is separately managed by restriction assessment, VM provision, and dynamic consolidation modules. Then the SLAs feedback are iteratively applied to reactively improve the performance of resource allocation. Experiments based on real-life collected data shows the feasibility and efficiency of our framework. The high accuracy of prediction contributes to a more suitable allocation. And a better trade-off between QoS and SLAs in server side is achieved compared with the baselines.",https://ieeexplore.ieee.org/document/8672346,True,"[{'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 25}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 9}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 424}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 675}]",1.0,['multilayer-perceptron'],,,"['scheduling', 'resource-consolidation']",,"2018 IEEE Intl Conf on Parallel & Distributed Processing with Applications, Ubiquitous Computing & Communications, Big Data & Cloud Computing, Social Computing & Networking, Sustainable Computing & Communications (ISPA/IUCC/BDCloud/SocialCom/SustainCom)",True,['resource-provisioning'],,,,,,,,,
291,Anomaly Detection and Root Cause Localization in Virtual Network Functions,"['C. Sauvanaud', ' K. Lazri', ' M. Kaâniche', ' K. Kanoun']",2016,"The maturity of hardware virtualization has motivated Communication Service Providers (CSPs) to apply thisparadigm to network services. Virtual Network Functions (VNFs)result from this trend and raise new dependability challengesrelated to network softwarisation that are still not thoroughlyexplored. This paper describes a new approach to detect ServiceLevel Agreements (SLAs) violations and preliminary symptomsof SLAs violations. In particular, one other major objectiveof our approach is to help CSP administrators to identify theanomalous VM at the origin of the detected SLA violation, whichshould enable them to proactively plan for appropriate recoverystrategies. To this end, we make use of virtual machine (VM)monitoring data and perform both a per-VM and an ensembleanalysis. Our approach includes a supervised machine learningalgorithm as well as fault injection tools. The experimental testbedconsists of a virtual IP Multimedia Subsystem developed by theClearwater project. Experimental results show that our approachcan achieve high precision and recall, and low false alarm rateand can pinpoint the root anomalous VNF VM causing SLAviolations. It can also detect preliminary symptoms of highworkloads triggering SLA violations.",https://ieeexplore.ieee.org/document/7774520,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 7}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 622}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 24}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 455}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 33}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 625}]",20.0,,,,"['root-cause-analysis', 'failure-detection']",,2016 IEEE 27th International Symposium on Software Reliability Engineering (ISSRE),True,['failure-management'],,True,,,,,,,
292,Hardware Remediation at Scale,"['F. Lin', ' M. Beadon', ' H. D. Dixit', ' G. Vunnam', ' A. Desai', ' S. Sankar']",2018,"Large scale services have automated hardware remediation to maintain the infrastructure availability at a healthy level. In this paper, we share the current remediation flow at Facebook, and how it is being monitored. We discuss a class of hardware issues that are transient and typically have higher rates during heavy load. We describe how our remediation system was enhanced to be efficient in detecting this class of issues. As hardware and systems change in response to the advancement in technology and scale, we have also utilized machine learning frameworks for hardware remediation to handle the introduction of new hardware failure modes. We present an ML methodology that uses a set of predictive thresholds to monitor remediation efficiency over time. We also deploy a recommendation system based on natural language processing, which is used to recommend repair actions for efficient diagnosis and repair. We also describe current areas of research that will enable us to improve hardware availability further.",https://ieeexplore.ieee.org/document/8416200,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 3}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 807}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1363}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 10}]",2.0,"['autoregression', 'language-modeling', 'gmm']","['tickets', 'logs', 'metrics']",,"['failure-detection', 'remediation']",['hardware'],2018 48th Annual IEEE/IFIP International Conference on Dependable Systems and Networks Workshops (DSN-W),True,['failure-management'],True,True,"['solution-recommendation', 'triage', 'anomaly-detection']",True,91.0,,,,['transient']
293,Failure Prediction Based on Anomaly Detection for Complex Core Routers,"['S. Jin', ' Z. Zhang', ' K. Chakrabarty', ' X. Gu']",2018,"Data-driven prognostic health management is essential to ensure high reliability and rapid error recovery in commercial core router systems. The effectiveness of prognostic health management depends on whether failures can be accurately predicted with sufficient lead time. This paper describes how time-series analysis and machine-learning techniques can be used to detect anomalies and predict failures in complex core router systems. First both a feature-categorization-based hybrid method and a changepoint-based method have been developed to detect anomalies in time-varying features with different statistical characteristics. Next, a SVM-based failure predictor is developed to predict both categories and lead time of system failures from collected anomalies. A comprehensive set of experimental results is presented for data collected during 30 days of field operation from over 20 core routers deployed by customers of a major telecom company.",https://ieeexplore.ieee.org/document/8587729,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 461}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 85}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('remediation' OR 'recovery')"", 'index': 163}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 81}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('anomaly detection' OR 'outlier detection')"", 'index': 218}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 52}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 372}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 385}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 882}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 14}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('anomaly detection' OR 'outlier detection')"", 'index': 87}, {'database': 'ACM', 'search_string': ""'regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 65}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 22}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 14}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('remediation' OR 'recovery')"", 'index': 65}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 5}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 93}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prediction' OR 'failure prediction')"", 'index': 7}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 8}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 9}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 15}]",1.0,['support-vector-machine'],"['kpis', 'metrics']",,"['failure-prediction', 'failure-detection']",['router'],2018 IEEE/ACM International Conference on Computer-Aided Design (ICCAD),True,['failure-management'],,,,,,10.1145/3240765.3243476,Conference Paper,"['failure prediction', 'changepoint detection', 'anomaly detection']",
294,Which Crashes Should I Fix First?: Predicting Top Crashes at an Early Stage to Prioritize Debugging Efforts,"['D. Kim', ' X. Wang', ' S. Kim', ' A. Zeller', ' S. C. Cheung', ' S. Park']",2011,"Many popular software systems automatically report failures back to the vendors, allowing developers to focus on the most pressing problems. However, it takes a certain period of time to assess which failures occur most frequently. In an empirical investigation of the Firefox and Thunderbird crash report databases, we found that only 10 to 20 crashes account for the large majority of crash reports; predicting these “top crashes” thus could dramatically increase software quality. By training a machine learner on the features of top crashes of past releases, we can effectively predict the top crashes well before a new release. This allows for quick resolution of the most important crashes, leading to improved user experience and better allocation of maintenance efforts.",https://ieeexplore.ieee.org/document/5711013,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 496}]",41.0,,,,['failure-prediction'],,IEEE Transactions on Software Engineering,True,['failure-management'],,,,,,,,,
295,Health Status Assessment and Failure Prediction for Hard Drives with Recurrent Neural Networks,"['C. Xu', ' G. Wang', ' X. Liu', ' D. Guo', ' T. Liu']",2016,"Recently, in order to improve reactive fault tolerance techniques in large scale storage systems, researchers have proposed various statistical and machine learning methods based on SMART attributes. Most of these studies have focused on predicting failures of hard drives, i.e., labeling the status of a hard drive as “good” or not. However, in real-world storage systems, hard drives often deteriorate gradually rather than suddenly. Correspondingly, their SMART attributes change continuously towards failure. Inspired by this observation, we introduce a novel method based on Recurrent Neural Networks (RNN) to assess the health statuses of hard drives based on the gradually changing sequential SMART attributes. Compared to a simple failure prediction method, a health status assessment is more valuable in practice because it enables technicians to schedule the recovery of different hard drives according to the level of urgency. Experiments on real-world datasets for disks of different brands and scales demonstrate that our proposed method can not only achieve a reasonable accurate health status assessment, but also achieve better failure prediction performance than previous work.",https://ieeexplore.ieee.org/document/7426410,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 858}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 59}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 63}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 1737}]",43.0,['rnn'],['host-metrics'],['novel-use'],['failure-prediction'],['hard-drive'],IEEE Transactions on Computers,True,['failure-management'],True,,['hardware-failure-prediction'],True,31.0,,,,
296,Anomaly Detection with HMM Gauge Likelihood Analysis,"['B. Lorbeer', ' T. Deutsch', ' P. Ruppel', ' A. Küpper']",2019,"This paper describes a new method, HMM gauge likelihood analysis, or GLA, of detecting anomalies in discrete time series using Hidden Markov Models and clustering. At the center of the method lies the comparison of subsequences. To achieve this, they first get assigned to their Hidden Markov Models using the Baum-Welch algorithm. Next, those models are described by an approximating representation of the probability distributions they define. Finally, this representation is then analyzed with the help of some clustering technique or other outlier detection tool and anomalies are detected. Clearly, HMMs could be substituted by some other appropriate model, e.g. some other dynamic Bayesian network. Our learning algorithm is unsupervised, so it doesn't require the labeling of large amounts of data. The usability of this method is demonstrated by applying it to synthetic and real-world syslog data.",https://ieeexplore.ieee.org/document/8848216,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 25}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 8}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('anomaly detection' OR 'outlier detection')"", 'index': 12}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 22}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 74}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 4}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 13}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 8}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 4}, {'database': 'arxiv', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('anomaly detection' OR 'outlier detection')"", 'index': 4}]",1.0,['markov-model'],,,['failure-detection'],,2019 IEEE Fifth International Conference on Big Data Computing Service and Applications (BigDataService),True,['failure-management'],,True,['anomaly-detection'],,,,,,
297,Online and Scalable Unsupervised Network Anomaly Detection Method,"['J. Dromard', ' G. Roudière', ' P. Owezarski']",2017,"Nowadays, network intrusion detectors mainly rely on knowledge databases to detect suspicious traffic. These databases have to be continuously updated which requires important human resources and time. Unsupervised network anomaly detectors overcome this issue by using “intelligent” techniques to identify anomalies without any prior knowledge. However, these systems are often very complex as they need to explore the network traffic to identify flows patterns. Therefore, they are often unable to meet real-time requirements. In this paper, we present a new online and real-time unsupervised network anomaly detection algorithm (ORUNADA). Our solution relies on a discrete time-sliding window to update continuously the feature space and an incremental grid clustering to detect rapidly the anomalies. The evaluations showed that ORUNADA can process online large network traffic while ensuring a low detection delay and good detection performance. The experiments performed on the traffic of a core network of a Spanish intermediate Internet service provider demonstrated that ORUNADA detects in less than half a second an anomaly after its occurrence. Furthermore, the results highlight that our solution outperforms in terms of true positive rate and false positive rate existing techniques reported in the literature.",https://ieeexplore.ieee.org/document/7740019,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 347}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 164}]",20.0,,,,['failure-detection'],,IEEE Transactions on Network and Service Management,True,['failure-management'],,,,,,,,,
298,"A Scalable, Non-Parametric Method for Detecting Performance Anomaly in Large Scale Computing","['L. Yu', ' Z. Lan']",2016,"As computer systems continue to grow in scale and complexity, performance problems become common and a major concern for large-scale computing. Performance anomalies caused by application bugs, hardware or software faults, or resource contention can have great impact on system-wide performance and could lead to significant economic losses for service providers. While many detection methods have been presented in the past, the newly emerging challenges are detection scalability and practical use. In this paper, we propose a scalable, non-parametric method for effectively detecting performance anomalies in large-scale systems. The design is generic for anomaly detection in a variety of parallel and distributed systems exhibiting peer-comparable property. It adopts a divide-and-conquer approach to address the scalability challenge and explores the use of non-parametric clustering and two-phase majority voting to improve detection flexibility and accuracy. We derive probabilistic models to quantitatively evaluate our decentralized design. Experiments with a suite of applications on production systems demonstrate that this method outperforms existing methods in terms of detection accuracy with a negligible runtime overhead.",https://ieeexplore.ieee.org/document/7236888,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 798}]",8.0,['clustering'],,,['failure-detection'],,IEEE Transactions on Parallel and Distributed Systems,True,['failure-management'],,,,,,,,,
299,Machine learning anomaly detection in large systems,['J. Murphree'],2016,"We have a need for methods to efficiently determine the health of a system. Diagnostics and prognostics determine system heath through analysis of data from sensors. Anomalies in the data can help us determine if there is a failure or a pending failure. There are common statistical methods to detect anomalies in individual measurements. For systems with many measurements, the anomalies may occur as specific combinations of values. Large systems have various associated states and modes which define the valid measurements. The amount of data to analyze grows very quickly as the system becomes more complex. In recent years techniques have been developed to address large data analysis. Machine Learning encompasses a broad selection of tools to optimize a statistical model of the data. These tools include supervised learning techniques, such as linear regression and logistic regression, in which training data exists to tune the model. Unsupervised learning, such as clustering, is used to explore data which does not have a defined output label associated with inputs data. Standard approaches to training supervised learning systems require a large sample of positive and negative outcome data. Some uses of machine learning involve data where there are very few cases of negative outcomes. There are machine learning algorithms defined as Anomaly Detection which are designed to deal with this type of data. Simple algorithms include Gaussian Distribution Analysis, which assumes independence in distributions of data. Large Systems with anomalies defined in the dependent combinations of data require either a manual creation of combinations of independent variables, or Multivariate Gaussian Distribution Analysis, which does not scale well for large systems. A further complication is the mixture of linear and discrete data. Neural Networks are a type of learning system which has been applied to each of the individual needs addressed above. This paper describes an approach to anomaly detection using neural networks for the specific problems in large systems to efficiently determine system health.",https://ieeexplore.ieee.org/document/7589589,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 815}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 21}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 251}, {'database': 'IEEE', 'search_string': ""'regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 168}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('anomaly detection' OR 'outlier detection')"", 'index': 306}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 114}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 508}]",12.0,,,,['failure-detection'],,2016 IEEE AUTOTESTCON,True,['failure-management'],,True,,,,,,,
300,Supporting the DevOps Feedback Loop using Unsupervised Machine Learning,"['I. Figalist', ' A. Biesdorf', ' C. Brand', ' S. Feld', ' M. Kiermeier']",2019,"Nowadays, software systems and applications need to adapt rapidly to changing requirements and evolve in an agile manner. This creates the need for independent deployments, e.g. as part of DevOps. Due to the flexibility and fast release cycles comprised by DevOps, continuous monitoring and the generation of feedback is crucial to a system’s quality, especially if it is continuously developed. For this purpose, we propose a feedback system that combines operations data with development data in order to trace anomalies occurring in production back to their root cause by defining patterns, detecting anomalous behavior, and generating feedback that is transferred back into the development process. To this end, we utilize two different unsupervised machine learning techniques, the k-means clustering and the archetypal analysis, to describe the data set and use the results as a basis to characterize the behavior of new data points as either normal or anomalous. The feedback system was tested and evaluated using real data produced by an application that is currently developed within a large, industrial company and serves as a link to support the loop of continuous planning, development, deployment, monitoring, and feedback.",https://ieeexplore.ieee.org/document/8778283,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 1204}, {'database': 'IEEE', 'search_string': ""'classification' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 6}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 4}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 633}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 1085}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 6}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 745}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 13}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1090}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 2}]",0.0,['clustering'],"['kpis', 'host-metrics']",['novel-use'],['failure-detection'],,2019 IEEE International Symposium on INnovations in Intelligent SysTems and Applications (INISTA),True,['failure-management'],,True,['anomaly-detection'],,,,,,
301,AIOps: Real-World Challenges and Research Innovations,"['Y. Dang', ' Q. Lin', ' P. Huang']",2019,"AIOps is about empowering software and service engineers (e.g., developers, program managers, support engineers, site reliability engineers) to efficiently and effectively build and operate online services and Apps at scale with artificial intelligence (AI) and machine learning (ML) techniques. AIOps can help achieve higher service quality and customer satisfaction, engineering productivity boost, and cost reduction. In this technical briefing, we summarize the real-world challenges on building AIOps solutions based on our practice and experience in Microsoft, propose a roadmap of AIOps related research directions, and share a few successful AIOps solutions we have built for Microsoft service products.",https://ieeexplore.ieee.org/document/8802836,True,"[{'database': 'IEEE', 'search_string': ""'AIOps'"", 'index': 1}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 20}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 17}, {'database': 'ACM', 'search_string': ""'AIOps'"", 'index': 1}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 24}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 33}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 45}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 8}]",1.0,,,['discussion'],,,2019 IEEE/ACM 41st International Conference on Software Engineering: Companion Proceedings (ICSE-Companion),True,['aiops-general'],,True,,,,10.1109/ICSE-Companion.2019.00023,Conference Paper,"['DevOps', 'AIOps', 'software analytics']",
302,Autonomic Resource Management with Support Vector Machines,"['O. Niehorster', ' A. Krieger', ' J. Simon', ' A. Brinkmann']",2011,"The use of virtualization technology makes data centers more dynamic and easier to administrate. Today, cloud providers offer customers access to complex applications running on virtualized hardware. Nevertheless, big virtualized data centers become stochastic environments and the implification on the user side leads to many challenges for the provider. He has to find cost-efficient configurations and has to deal with dynamic environments to ensure service guarantees. In this paper, we introduce a software solution that reduces the degree of human intervention to manage cloud services. We present a multi-agent system located in the Software as a Service (SaaS) layer. Agents allocate resources, configure applications, check the feasibility of requests, and generate cost estimates. The agents learn behavior models of the services via Support Vector Machines (SVMs) and share their experiences via a global knowledge base. We evaluate our approach on real cloud systems with three different applications, a brokerage system, a high-performance computing software, and a web server.",https://ieeexplore.ieee.org/document/6076511,True,"[{'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 14}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 266}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 226}]",8.0,,,,,,2011 IEEE/ACM 12th International Conference on Grid Computing,True,['resource-provisioning'],,,,,,,,,
303,Deep Reinforcement Learning based Load Balancing Policy for balancing network traffic in datacenter environment,"['A. R. Doke', ' K. Sangeeta']",2018,Load balancer plays important role in handling a huge amount of network traffic by routing the request/traffic in such a way that clients get immediate response to their requests. But traffic management in this era of bigdata is becoming a challenging task and to maintain them with human support is becoming more expensive. We can address this challenge by applying Deep reinforcement learning for a network load balancer which will be both time and cost effective. Deep reinforcement learning understands and adjusts continuously with dynamic environment. Which can be used to optimize the performance of load balancer.,https://ieeexplore.ieee.org/document/8752969,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 5}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 9}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 25}]",0.0,,,,,,2018 Second International Conference on Green Computing and Internet of Things (ICGCIoT),True,['resource-provisioning'],,True,,,,,,,
304,Utility-Function-Driven Resource Allocation in Autonomic Systems,"['G. Tesauro', ' R. Das', ' W. E. Walsh', ' J. O. Kephart']",2005,We study autonomic resource allocation among multiple applications based on optimizing the sum of utility for each application. We compare two methodologies for estimating the utility of resources: a queuing-theoretic performance model and model-free reinforcement learning. We evaluate them empirically in a distributed prototype data center and highlight tradeoffs between the two methods,https://ieeexplore.ieee.org/document/1498088,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 26}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 106}]",26.0,,,,,,Second International Conference on Autonomic Computing (ICAC'05),True,['resource-provisioning'],,,,,,,,,
305,A Hybrid Reinforcement Learning Approach to Autonomic Resource Allocation,"['G. Tesauro', ' N. K. Jong', ' R. Das', ' M. N. Bennani']",2006,"Reinforcement Learning (RL) provides a promising new approach to systems performance management that differs radically from standard queuing-theoretic approaches making use of explicit system performance models. In principle, RL can automatically learn high-quality management policies without an explicit performance model or traffic model and with little or no built-in system specific knowledge. In our original work [1], [2], [3] we showed the feasibility of using online RL to learn resource valuation estimates (in lookup table form) which can be used to make high-quality server allocation decisions in a multi-application prototype Data Center scenario. The present work shows how to combine the strengths of both RL and queuing models in a hybrid approach in which RL trains offline on data collected while a queuing model policy controls the system. By training offline we avoid suffering potentially poor performance in live online training. We also now use RL to train nonlinear function approximators (e.g. multi-layer perceptrons) instead of lookup tables; this enables scaling to substantially larger state spaces. Our results now show that in both open-loop and closed-loop traffic, hybrid RL training can achieve significant performance improvements over a variety of initial model-based policies. We also find that, as expected, RL can deal effectively with both transients and switching delays, which lie outside the scope of traditional steady-state queuing theory.",https://ieeexplore.ieee.org/document/1662383,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 32}]",52.0,,,,['resource-consolidation'],,2006 IEEE International Conference on Autonomic Computing,True,['resource-provisioning'],,True,,,,,,,
306,Autonomic Provisioning with Self-Adaptive Neural Fuzzy Control for End-to-end Delay Guarantee,"['P. Lama', ' X. Zhou']",2010,"Autonomic server provisioning for performance assurance is a critical issue in data centers. It is important but challenging to guarantee an important performance metric, percentile-based end-to-end delay of requests flowing through a virtualized multi-tier server cluster. It is mainly due to dynamically varying workload and the lack of an accurate system performance model. In this paper, we propose a novel autonomic server allocation approach based on a model-independent and self-adaptive neural fuzzy control. There are model-independent fuzzy controllers that utilize heuristic knowledge in the form of rule base for performance assurance. Those controllers are designed manually on trial and error basis, often not effective in the face of highly dynamic workloads. We design the neural fuzzy controller as a hybrid of control theoretical and machine learning techniques. It is capable of self-constructing its structure and adapting its parameters through fast online learning. Unlike other supervised machine learning techniques, it does not require off-line training. We further enhance the neural fuzzy controller to compensate for the effect of server switching delays. Extensive simulations demonstrate the effectiveness of our new approach in achieving the percentile-based end-to-end delay guarantees. Compared to a rule-based fuzzy controller enabled server allocation approach, the new approach delivers superior performance in the face of highly dynamic workloads. It is robust to workload variation, change in delay target and server switching delays.",https://ieeexplore.ieee.org/document/5581598,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 42}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 81}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 150}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 44}]",32.0,,,,,,"2010 IEEE International Symposium on Modeling, Analysis and Simulation of Computer and Telecommunication Systems",True,['resource-provisioning'],,,,,,,,,
307,Log filtering and interpretation for root cause analysis,"['H. Zawawy', ' K. Kontogiannis', ' J. Mylopoulos']",2010,"Problem diagnosis in large software systems is a challenging and complex task. The sheer complexity and size of the logged data make it often difficult for human operators and administrators to perform problem diagnosis and root cause analysis. A challenge in this area is to provide the necessary means, tools, and techniques for the operators to focus their attention to specific parts of the logged data reducing thus the complexity of the diagnostic process. In this paper, we propose a framework for filtering logs according to specific analysis goals and diagnostic hypotheses set by the user or by an automated process. More specifically, the proposed framework uses annotated goal trees to model the constraints and the conditions by which the functionality of a particular system is being delivered. Next, a transformation process maps such constraints and conditions to a collection of queries that can be either applied to a relational database that stores the logged data or use Latent Semantic Indexing to identify the most relevant log entries for the given query. The results of such queries provide a subset of the logged data that is compliant with the goal tree and can be used by a diagnostic SAT-solver based algorithm. Experimental results show that the filtering process can reduce the time and complexity of the diagnosis when applied to multi-tier heterogeneous service oriented systems.",https://ieeexplore.ieee.org/document/5609556,True,"[{'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 4}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 141}]",8.0,['constraint-solving'],,,['root-cause-analysis'],,2010 IEEE International Conference on Software Maintenance,True,['failure-management'],,,,,,,,,
308,Data Mining Based Root-Cause Analysis of Performance Bottleneck for Big Data Workload,"['W. Qi', ' Y. Li', ' H. Zhou', ' W. Li', ' H. Yang']",2017,"Straggler task is commonly considered as the major bottleneck in parallel data processing. Previous work mainly focuses on the coarse-grained straggler detection and optimization such as speculative scheduling. However, fine-grained root-cause analysis of straggler tasks is rarely considered. In addition, existing work simply depends on empirical analysis, which lacks of useful guidance to performance optimization. In this paper, we propose a new methodology of fine-grained straggler root-cause analysis using machine learning. We collect raw metrics from Spark event log and hardware sampling tool, and refine them into high-level metrics for model learning. Then we present the root-cause analysis of stragglers through CART tree. A customized prune method is also applied to improve analysis accuracy. From the analysis, we derive several new findings beyond the well known causes of stragglers. Our work provides a new perspective on identifying and understanding the inefficiency in parallel data processing programs by applying machine learning techniques to fine-grained root-cause analysis of straggler tasks.",https://ieeexplore.ieee.org/document/8291936,True,"[{'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 6}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 832}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1483}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1082}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 5}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 7}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1506}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1529}]",4.0,,,,['root-cause-analysis'],,2017 IEEE 19th International Conference on High Performance Computing and Communications; IEEE 15th International Conference on Smart City; IEEE 3rd International Conference on Data Science and Systems (HPCC/SmartCity/DSS),True,['failure-management'],,,,,,,,,
309,Random forest and change point detection for root cause localization in large scale systems,"['D. V. Sagar', ' P. B. Sivakumar', ' R. V. Anand']",2014,"Identification of root causes of a performance problem is very difficult in case of large scale IT environment. A model which is scalable and reasonably accurate is required for such complex scenarios. This paper proposes a hybrid model using random forest and statistical change point detection, for root cause localization. Based on impurity measure and change in error rates, random forest identifies the features which can be a potential cause for the problem. Since it is a tree based approach, it does not require any prior information about the measured features. To reduce the number of false classifications, a second level of selection using change point analysis is done. The ability of random forest to work well on very large dataset makes the solution scalable and accurate. Proposed model is applied and verified by identifying the root causes for Service Level Objective Violations in enterprise IT systems.",https://ieeexplore.ieee.org/document/7238442,True,"[{'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 13}, {'database': 'IEEE', 'search_string': ""'classification' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 11}]",0.0,['random-forest'],,['novel-use'],['root-cause-analysis'],,2014 IEEE International Conference on Computational Intelligence and Computing Research,True,['failure-management'],,,['root-cause-diagnosis'],,,,,,
310,Mining unstructured log files for recurrent fault diagnosis,"['T. Reidemeister', ' Miao Jiang', ' P. A. S. Ward']",2011,"Enterprise software systems are large and complex with limited support for automated root-cause analysis. Avoiding system downtime and loss of revenue dictates a fast and efficient root-cause analysis process. Operator practice and academic research have shown that about 80% of failures in such systems have recurrent causes; therefore, significant efficiency gains can be achieved by automating their identification. In this paper, we present a novel approach to modelling features of log files. This model offers a compact representation of log data that can be efficiently extracted from large amounts of monitoring data. We also use decision-tree classifiers to learn and classify symptoms of recurrent faults. This representation enables automated fault matching and, in addition, enables human investigators to understand manifestations of failure easily. Our model does not require any access to application source code, a specification of log messages, or deep application knowledge. We evaluate our proposal using fault-injection experiments against other proposals in the field. First, we show that the features needed for symptom definition can be extracted more efficiently than does related work. Second, we show that these features enable an accurate classification of recurrent faults using only standard machine learning techniques. This enables us to identify accurately up to 78% of the faults in our evaluation data set.",https://ieeexplore.ieee.org/document/5990536,True,"[{'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 24}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 242}, {'database': 'IEEE', 'search_string': ""'classification' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 22}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 237}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 374}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 23}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 29}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 454}]",8.0,,,,['root-cause-analysis'],,12th IFIP/IEEE International Symposium on Integrated Network Management (IM 2011) and Workshops,True,['failure-management'],,,,,,,,,
311,POD-Diagnosis: Error Diagnosis of Sporadic Operations on Cloud Applications,"['X. Xu', ' L. Zhu', ' I. Weber', ' L. Bass', ' D. Sun']",2014,"Applications in the cloud are subject to sporadic changes due to operational activities such as upgrade, redeployment, and on-demand scaling. These operations are also subject to interferences from other simultaneous operations. Increasing the dependability of these sporadic operations is non-trivial, particularly since traditional anomaly-detection-based diagnosis techniques are less effective during sporadic operation periods. A wide range of legitimate changes confound anomaly diagnosis and make baseline establishment for ""normal"" operation difficult. The increasing frequency of these sporadic operations (e.g. due to continuous deployment) is exacerbating the problem. Diagnosing failures during sporadic operations relies heavily on logs, while log analysis challenges stemming from noisy, inconsistent and voluminous logs from multiple sources remain largely unsolved. In this paper, we propose Process Oriented Dependability (POD)-Diagnosis, an approach that explicitly models these sporadic operations as processes. These models allow us to (i) determine orderly execution of the process, and (ii) use the process context to filter logs, trigger assertion evaluations, visit fault trees and perform on-demand assertion evaluation for online error diagnosis and root cause analysis. We evaluated the approach on rolling upgrade operations in Amazon Web Services (A WS) while performing other simultaneous operations. During our evaluation, we correctly detected all of the 160 injected faults, as well as 46 interferences caused by concurrent operations. We did this with 91.95% precision. Of the correctly detected faults, the accuracy rate of error diagnosis is 96.55%.",https://ieeexplore.ieee.org/document/6903584,True,"[{'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 35}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 617}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 12}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 251}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 597}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 415}]",28.0,,,,['root-cause-analysis'],,2014 44th Annual IEEE/IFIP International Conference on Dependable Systems and Networks,True,['failure-management'],,,,,,,,,
312,Performance Metric Selection for Autonomic Anomaly Detection on Cloud Computing Systems,['S. Fu'],2011,"With ever-growing complexity and dynamicity of cloud computing systems, dependability assurance has become a major concern in system design and management. In this paper, we propose a framework for autonomic anomaly detection in the cloud. Mutual information is exploited to quantify the relevance and redundancy among the large number of performance metrics. An incremental search algorithm is presented for metric selection. We apply principal component analysis to further reduce the metric dimension, while keeping the variance in the health- related data as much as possible. A detection mechanism with semi-supervised decision tree classifiers works on the reduce metric dimensionality and identifies anomalies. We have implemented a prototype of our autonomic anomaly detection framework and evaluated its performance on an institute-wide cloud computing system.",https://ieeexplore.ieee.org/document/6134532,True,"[{'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 169}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 391}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 1206}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 731}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 156}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 564}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 989}]",11.0,,,,['failure-detection'],,2011 IEEE Global Telecommunications Conference - GLOBECOM 2011,True,['failure-management'],,,,,,,,,
313,Multi-objective optimization model of virtual resources scheduling under cloud computing and it's solution,"['J. Zhao', ' W. Zeng', ' M. Liu', ' G. Li']",2011,"It's an basic requirement in cloud computing that scheduling virtual resources to physical resources with balance load, however, the simple scheduling methods can not meet this requirement. This paper proposed a virtual resources scheduling model and solved it by advanced Non-dominated Sorting Genetic Algorithm II (NSGA II). This model was evaluated by balance load, virtual resources and physical resources were abstracted a lot of nodes with attributes based on analyzing the flow of virtual resources scheduling. NSGA II was employed to address this model and a new tree sorting algorithms was adopted to improve the efficiency of NSGA II. In experiment, verified the correctness of this model. Comparing with Random algorithm, Static algorithm and Rank algorithm by a lot of experiments, at least 1.06 and at most 40.25 speed-up of balance degree can be obtained by NSGA II.",https://ieeexplore.ieee.org/document/6138518,True,"[{'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 298}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 596}]",12.0,['genetic-programming'],,['novel-use'],['resource-consolidation'],,2011 International Conference on Cloud and Service Computing,True,['resource-provisioning'],,,,,,,,,
314,Semantics-Based Anomaly Detection of Processes in Linux Containers,"['H. Liang', ' Q. Hao', ' M. Li', ' Y. Zhang']",2016,"With the development of the cloud computing, Linux containers are playing an important role in industrial use, however, the containers are suffering more and more cyber-attacks. A novel semantics-based anomaly detection approach of processes in Linux containers is presented and implemented in this paper, which extracts the features of processes by using the system calls produced by container behaviors, finds the relations between the processes, and builds the features tree of the processes. Experiments show that the approach we proposed can identify the abnormal processes effectively in Linux containers.",https://ieeexplore.ieee.org/document/8281174,True,"[{'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 551}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 72}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 1234}]",0.0,,,,['failure-detection'],,"2016 International Conference on Identification, Information and Knowledge in the Internet of Things (IIKI)",True,['failure-management'],,,,,,,,,
315,Improved approach for software defect prediction using artificial neural networks,"['T. Sethi', ' Gagandeep']",2016,"Software defect prediction (SDP) is a most dynamic research area in software engineering. SDP is a process used to predict the deformities in the software. To identifying the defects before the arrival of item or aimed the software improvement, to make software dependable, defect prediction model is utilized. It is always desirable to predict the defects at early stages of life cycle. Hence to predict the defects before testing the SDP is done at end of each phase of SDLC. It helps to reduce the cost as well as time. To produce high quality software, the artificial neural network approach is applied to predict the defect. Nine metrics are applied to the multiple phases of SDLC and twenty genuine software projects are used. The software project data were collected from a team of organization and their responses were recorded in linguistic terms. For assessment of model the mean magnitude of relative error (MMRE) and balanced mean magnitude of relative error (BMMRE) measures are used. In this research work, the implementation of neural network based software defect prediction is compared with the results of fuzzy logic basic approach. In the proposed approach, it is found that the neural network based training model is providing better and effective results on multiple parameters.",https://ieeexplore.ieee.org/document/7785003,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 56}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 855}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 676}]",5.0,,,,['failure-prediction'],,"2016 5th International Conference on Reliability, Infocom Technologies and Optimization (Trends and Future Directions) (ICRITO)",True,['failure-management'],,,,,,,,,
316,End-to-end encrypted traffic classification with one-dimensional convolution neural networks,"['W. Wang', ' M. Zhu', ' J. Wang', ' X. Zeng', ' Z. Yang']",2017,"Traffic classification plays an important and basic role in network management and cyberspace security. With the widespread use of encryption techniques in network applications, encrypted traffic has recently become a great challenge for the traditional traffic classification methods. In this paper we proposed an end-to-end encrypted traffic classification method with one-dimensional convolution neural networks. This method integrates feature extraction, feature selection and classifier into a unified end-to-end framework, intending to automatically learning nonlinear relationship between raw input and expected output. To the best of our knowledge, it is the first time to apply an end-to-end method to the encrypted traffic classification domain. The method is validated with the public ISCX VPN-nonVPN traffic dataset. Among all of the four experiments, with the best traffic representation and the fine-tuned model, 11 of 12 evaluation metrics of the experiment results outperform the state-of-the-art method, which indicates the effectiveness of the proposed method.",https://ieeexplore.ieee.org/document/8004872,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 263}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1115}]",93.0,['cnn'],['packet-content'],['novel-use'],['failure-detection'],['network'],2017 IEEE International Conference on Intelligence and Security Informatics (ISI),True,['failure-management'],True,,['traffic-classification'],True,67.0,,,,['privacy']
317,Automated debugging of SLO violations in enterprise systems,"['M. Natu', ' S. Patil', ' V. Sadaphal', ' H. Vin']",2010,"A critical business requirement of today's enterprise applications is automated debugging of violation of Service Level Objectives (SLOs). However, the increasing scale and complexity of these systems present various challenges in building such solutions. This problem becomes challenging mainly because of two reasons: (1) availability of large number of metrics that can potentially be the causes, (2) availability of large number of data-points. The existing techniques are either highly compute-intensive and thus are not viable for use on large volumes of data or compromise on accuracy. To successfully balance these two objectives simultaneously, we propose to intelligently prune the search space. We apply feature selection to remove irrelevant and redundant metrics. We then identify temporal regions of interest to narrow down the analysis to a smaller set of data-points. We present a comparative study of the proposed approach with other existing approaches through experimental evaluation.",https://ieeexplore.ieee.org/document/5432008,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 692}]",4.0,,,,['failure-detection'],,2010 Second International Conference on COMmunication Systems and NETworks (COMSNETS 2010),True,['failure-management'],,,,,,,,,
318,A new proposal to provide estimation of QoS and QoE over WiMAX networks: An approach based on computational intelligence and discrete-event simulation,"['V. A. Machado', ' C. N. Silva', ' R. S. Oliveira', ' A. M. Melo', ' M. Silva', ' C. R. L. Francês', ' J. C. W. A. Costa', ' N. L. Vijaykumar', ' C. M. Hirata']",2011,This paper presents an estimation of Quality of Experience (QoE) metrics based on Quality of Service (QoS) metrics in WiMAX networks. Applications used to generate such estimations were EvalVid and Network Simulator 2 (NS-2). The QoE was estimated by employing a Multilayer Artificial Neural Network by means of the WEKA tool. The results show a very efficient estimation of metrics of QoE parameters with respect to QoS parameters.,https://ieeexplore.ieee.org/document/6107419,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1118}]",11.0,,,,,,2011 IEEE Third Latin-American Conference on Communications,True,['resource-provisioning'],,,,,,,,,
319,CoSL: A coordinated statistical learning approach to measuring the capacity of multi-tier websites,"['Jia Rao', ' Cheng-Zhong Xu']",2008,"Website capacity determination is crucial to measurement-based access control, because it determines when to turn away excessive client requests to guarantee consistent service quality under overloaded conditions. Conventional capacity measurement approaches based on high-level performance metrics like response time and throughput may result in either resource over-provisioning or lack of responsiveness. It is because a website may have different capacities in terms of the maximum concurrent level when the characteristic of workload changes. Moreover, bottleneck in a multi-tier website may shift among tiers as client access pattern changes. In this paper, we present an online robust measurement approach based on statistical machine learning techniques. It uses a Bayesian network to correlate low level instrumentation data like system and user cpu time, available memory size, and I/O status that are collected at run-time to high level system states in each tier. A decision tree is induced over a group of coordinated Bayesian models in different tiers to identify the bottleneck dynamically when the system is overloaded. Experimental results demonstrate its accuracy and robustness in different traffic loads.",https://ieeexplore.ieee.org/document/4536232,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1202}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1234}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1655}]",2.0,,,,,,2008 IEEE International Symposium on Parallel and Distributed Processing,True,['resource-provisioning'],,,,,,,,,
320,Issues in Bottleneck Detection in Multi-Tier Enterprise Applications,"['J. Parekh', ' G. Jung', ' G. Swint', ' C. Pu', ' A. Sahai']",2006,"In this work, the performance of various machine learning classifiers with regard to bottleneck detection in enterprise, multi-tier applications governed by service level objectives is described. Specifically, in this paper, it demonstrates the effectiveness of three classifiers, a tree-augmented Naive Bayesian network, a J48 decision tree, and LogitBoost, using our bottleneck detection process, which delves into a new area of performance analysis based on the trends of metrics (first order derivative) rather than the metric value itself. Furthermore, the efficiency of each classifier by measuring the convergence speed, or the number of staging trials required in order to provide positive results is illustrated. Finally, the effectiveness of the classifiers used in the bottleneck detection process as each classifier strongly identifies the enterprise system bottleneck",https://ieeexplore.ieee.org/document/4015772,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1270}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 473}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 825}]",7.0,,,,['failure-detection'],,200614th IEEE International Workshop on Quality of Service,True,['failure-management'],,,,,,,,,
321,Ergodic Continuous Hidden Markov Models for workload characterization,"['A. Moro', ' E. Mumolo', ' M. Nolich']",2009,"In this paper we present a novel approach for accurate characterization of the execution workload run by a computer. Usually, workload characterization is performed by measuring the type and amount of resources requested during a program execution (for instance the usage of CPU, I/O, network, etc.). The sequence of measures is then treated as a stochastic process and analyzed with statistical techniques. The novelty of our approach is that we instead use directly the sequence of memory references generated during the execution of a program. The sequences of memory references are treated as sequences of floating point numbers, and analyzed with signal processing techniques. In the feature extraction phase we use spectral analysis while in the pattern matching phase we use ergodic continuous hidden Markov models (ECHMMs). The ECHMM models estimated in an initial training phase can be used both for online workload classification of a running process and for synthetic traces generation. Several processes of the same workload are necessary to obtain an HMM model of the workload. The proposed algorithms is evaluated via trace driven simulations using the SPEC 2000 workloads. We show that ECHMMs describe address memory sequences; average classification accuracy is about 76% with eight different workloads.",https://ieeexplore.ieee.org/document/5297771,True,"[{'database': 'IEEE', 'search_string': ""'classification' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 427}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 40}]",6.0,,,,,,2009 Proceedings of 6th International Symposium on Image and Signal Processing and Analysis,True,['resource-provisioning'],,,,,,,,,
322,Anomaly Detection Using Inter-Arrival Curves for Real-Time Systems,"['M. Salem', ' M. Crowley', ' S. Fischmeister']",2016,"Real-time embedded systems are a significant class of applications, poised to grow even further as automated vehicles and the Internet of Things become a reality. An important problem for these systems is to detect anomalies during operation. Anomaly detection is a form of classification, which can be driven by data collected from the system at execution time. We propose inter-arrival curves as a novel analytic modelling technique for discrete event traces. Our approach relates to the existing technique of arrival curves and expands the technique to anomaly detection. Inter-arrival curves analyze the behaviour of events within a trace by providing upper and lower bounds to their inter-arrival occurrence. We exploit inter-arrival curves in a classification framework that detects deviations within these bounds for anomaly detection. Also, we show how inter-arrival curves act as good features to extract recurrent behaviour that these systems often exhibit. We demonstrate the feasibility and viability of the fully implemented approach with an industrial automotive case study (CAN traces) as well as a deployed aerospace case study (RTOS kernel traces).",https://ieeexplore.ieee.org/document/7557872,True,"[{'database': 'IEEE', 'search_string': ""'classification' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 485}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 513}]",6.0,,,,['failure-detection'],,2016 28th Euromicro Conference on Real-Time Systems (ECRTS),True,['failure-management'],,,,,,,,,
323,Failure Prediction of Jobs in Compute Clouds: A Google Cluster Case Study,"['X. Chen', ' C. Lu', ' K. Pattabiraman']",2014,"Most cloud computing clusters are built from unreliable, commercial off-the-shelf components. The high failure rates in their hardware and software components result in frequent node and application failures. Therefore, it is important to predict application failures before they occur to avoid resource wastage. In this paper, we investigate how to identify application failures based on resource usage measurements from the Google cluster traces. We apply recurrent neural networks to the resource usage measures, and generate features to categorize the input resource usage time series into different classes. Our results show that the model is able to predict failures of batch applications, which are the dominant jobs in the Google cluster. Moreover, we explore early classification to identify failures, and find that the prediction algorithm provides the cloud system enough time to take proactive actions much earlier than the termination of applications, with an average 6% to 10% of resource savings.",https://ieeexplore.ieee.org/document/6983864,True,"[{'database': 'IEEE', 'search_string': ""'classification' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1322}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 77}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 1452}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 1960}]",36.0,['rnn'],['host-metrics'],['novel-use'],['failure-prediction'],['job'],2014 IEEE International Symposium on Software Reliability Engineering Workshops,True,['failure-management'],,True,['system-failure-prediction'],,,,,,
324,Metrics selection for fault-proneness prediction of software modules,"['Luo Yunfeng', ' Ben Kerong']",2010,"It would be valuable to use metrics to identify the fault-proneness of software modules. It is important to select the most appropriate particular metric subset for fault-proneness prediction. We proposed an approach of metrics selection, which firstly utilized the correlation analysis to eliminate the high the correlation metrics and then ranked the remaining metrics based on the gray relational analysis. Three classifiers, that were logistic regression model, NaiveBayes, and J48, were utilized to empirically investigate the usefulness of selected metrics. Our results, based on a public domain NASA data set, indicate that 1) proposed method for metrics selection is effective, and 2) using 3-4 metrics gets the balanced performance for fault-proneness prediction of software modules.",https://ieeexplore.ieee.org/document/5541206,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 5}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1}]",1.0,"['naive-bayes', 'logistic-regression', 'decision-tree']",['software-metrics'],,['failure-prediction'],['software'],2010 International Conference On Computer Design and Applications,True,['failure-management'],,,,,,,,,
325,Exploratory study of a UML metric for fault prediction,['A. E. Camargo Cruz'],2010,"This paper describes the use of a UML metric, an approximation of the CK-RFC metric, for predicting faulty classes before their implementation. We built a code-based prediction model of faulty classes using Logistic Regression. Then, we tested it in different projects, using on the one hand their UML metrics, and on the other hand their code metrics. To decrease the difference of values between UML and code measures, we normalized them using Linear Scaling to Unit Variance. Our results indicate that the proposed UML RFC metric can predict faulty code as well as its corresponding code metric does. Moreover, the normalization procedure used was of great utility, not just for enabling our UML metric to predict faulty code, using a code-based prediction model, but also for improving the prediction results across different packages and projects, using the same model.",https://ieeexplore.ieee.org/document/6062214,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 50}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 2}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 16}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 10}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 58}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 9}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 31}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 23}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 31}]",1.0,['logistic-regression'],['software-metrics'],,['failure-prediction'],['software'],2010 ACM/IEEE 32nd International Conference on Software Engineering,True,['failure-management'],,,,,,10.1145/1810295.1810393,Conference Paper,"['UML', 'CK metrics', 'fault-prone code', 'fault-proneness prediction', 'logistic regression']",
326,Using regression trees to classify fault-prone software modules,"['T. M. Khoshgoftaar', ' E. B. Allen', ' Jianyu Deng']",2002,"Software faults are defects in software modules that might cause failures. Software developers tend to focus on faults, because they are closely related to the amount of rework necessary to prevent future operational software failures. The goal of this paper is to predict which modules are fault-prone and to do it early enough in the life cycle to be useful to developers. A regression tree is an algorithm represented by an abstract tree, where the response variable is a real quantity. Software modules are classified as fault-prone or not, by comparing the predicted value to a threshold. A classification rule is proposed that allows one to choose a preferred balance between the two types of misclassification rates. A case study of a very large telecommunications systems considered software modules to be fault-prone, if any faults were discovered by customers. Our research shows that classifying fault-prone modules with regression trees and the using the classification rule in this paper, resulted in predictions with satisfactory accuracy and robustness.",https://ieeexplore.ieee.org/document/1044344,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 145}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 384}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 71}]",69.0,,,,['failure-prediction'],,IEEE Transactions on Reliability,True,['failure-management'],,,,,,,,,
327,Tree-based software quality estimation models for fault prediction,"['T. M. Khoshgoftaar', ' N. Seliya']",2002,"Complex high-assurance software systems depend highly on reliability of their underlying software applications. Early identification of high-risk modules can assist in directing quality enhancement efforts to modules that are likely to have a high number of faults. Regression tree models are simple and effective as software quality prediction models, and timely predictions from such models can be used to achieve high software reliability. This paper presents a case study from our comprehensive evaluation (with several large case studies) of currently available regression tree algorithms for software fault prediction. These are, CART-LS (least squares), S-PLUS, and CART-LAD (least absolute deviation). The case study presented comprises of software design metrics collected from a large network telecommunications system consisting of almost 13 million lines of code. Tree models using design metrics are built to predict the number of faults in modules. The algorithms are also compared based on the structure and complexity of their tree models. Performance metrics, average absolute and average relative errors are used to evaluate fault prediction accuracy.",https://ieeexplore.ieee.org/document/1011339,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 243}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 54}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 3}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 41}]",51.0,,,,['failure-prediction'],,Proceedings Eighth IEEE Symposium on Software Metrics,True,['failure-management'],,,,,,,,,
328,A system for online power prediction in virtualized environments using gaussian mixture models,"['G. Dhiman', ' K. Mihic', ' T. Rosing']",2010,"In this paper we present a system for online power prediction in vir-tualized environments. It is based on Gaussian mixture models that use architectural metrics of the physical and virtual machines (VM) collected dynamically by our system to predict both the physical machine and per VM level power consumption. A real implementation of our system shows that it can achieve average prediction error of less than 10%, outperforming state of the art regression based approaches at negligible runtime overhead.",https://ieeexplore.ieee.org/document/5523620,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 988}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 9}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 100}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 33}]",51.0,,,,"['power-management', 'workload-prediction']",,Design Automation Conference,True,['resource-provisioning'],,True,,,,10.1145/1837274.1837478,Conference Paper,"['virtualization', 'workload characterization', 'power', 'regression', 'Gaussian mixture models']",
329,Real Time Failure Prediction of Load Balancers and Firewalls,"['T. Ghosh', ' D. Sarkar', ' T. Sharma', ' A. Desai', ' R. Bali']",2016,"Load balancers and firewalls are entry points to many key websites in any network infrastructure. Any downtime due to these devices, even for a few minutes, may result in major business impact. There are many proactive event management systems/products which can monitor their health and alert on device failure. However, predicting failure with sufficient lead time still remains a challenging problem. Event management systems take system log messages and simple network management protocol (SNMP) traps as inputs and typically use rule based framework to generate network events. In this paper we discuss how we can use event data alone to build predictive model for device failures. Here all event data is generated by an event management tool. Event data volume is substantially less compared to full system log data. Also, these network devices fail rarely. Lower volume of event data enables us collect historical data for longer durations and thus collect decent number of failure samples over time which we can use for training and validation. We present a prediction model for failure events based on event sequence data. We have introduced new stratified sampling techniques along with a new feature engineering method using sliding time windows on event data. Experiments show that for rare device failure events like the load balancers it suffices to use event data to model device failures instead of using raw system log data. We have evaluated binary classification algorithms like support vector machines (SVM) and logistic regression (LR). Experimental results show that with the proper noise cleanup technique and model tuning we can achieve precision of 77% and recall of 67% for the failure predictions of network devices only from event data.",https://ieeexplore.ieee.org/document/7917199,True,"[{'database': 'IEEE', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 11}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 249}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 68}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1005}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 43}, {'database': 'IEEE', 'search_string': ""'logistic regression' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 71}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 78}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1477}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 38}, {'database': 'IEEE', 'search_string': ""'regression' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 437}]",2.0,,,,['failure-prediction'],,"2016 IEEE International Conference on Internet of Things (iThings) and IEEE Green Computing and Communications (GreenCom) and IEEE Cyber, Physical and Social Computing (CPSCom) and IEEE Smart Data (SmartData)",True,['failure-management'],,True,,,,,,,
330,Disk failure prediction in heterogeneous environments,"['C. A. C. Rincón', ' J. Pâris', ' R. Vilalta', ' A. M. K. Cheng', ' D. D. E. Long']",2017,"Recent studies have shown the benefits of using SMART attributes to predict disk failures in homogeneous populations of disks from the same make and model. We address here the case of data centers with more heterogeneous disk populations, such as the ones described in the BackBlaze datasets, and propose to build global disk failure predictors that would apply to disks of all makes and models. Our first challenge was the large number of SMART parameters that were missing for most makes and models in many disk instances of our dataset. As a result, we had to discard the SMART attributes that were missing in at least 90 percent of the disks, which left us with 21 SMART attributes. We then applied a Reverse Arrangement Test to these attributes to select the strongest disk failure indicators. We investigated three different machine learning models (Decision Trees, Neural Networks, and Logistic Regression) using the 2015 BackBlaze data to train and validate our predictors. Our best model was a decision tree that identified t rue failure events among the disks that tested positive for at least one of our failure indicators. We then used the 2016 BackBlaze data to evaluate its performance. Our results show that our decision tree identifies at least 52 percent of all disk failures and makes nearly all its predictions several days ahead: no more than 2.45 percent of the predicted failures occur within one day or two of the prediction. Finally, we compared the performance of our predictor with those of the RAIDShield and the original BackBlaze predictor. We found out that RAIDShield could predict at most 18 percent of disk failures, that is, 34 percent fewer failures than our decision tree while the BackBlaze predictor predicted 60 percent of disk failures but generated 4 to 5 false alarms per correct prediction.",https://ieeexplore.ieee.org/document/8046776,True,"[{'database': 'IEEE', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 15}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 84}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 75}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 19}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 53}]",3.0,,,,['failure-prediction'],,2017 International Symposium on Performance Evaluation of Computer and Telecommunication Systems (SPECTS),True,['failure-management'],,,,,,,,,
331,A Convolutional Auto-Encoder Method for Anomaly Detection on System Logs,"['Y. Cui', ' Y. Sun', ' J. Hu', ' G. Sheng']",2018,"Anomaly detection on system logs is to report system failures with utilization of console logs collected from devices, which ensures the reliability of systems. Most previous researches split logs into sequential time windows and regarded each window as an independent instance for classification using popular machine learning methods like support vector machine(SVM), however, neglected the time patterns under logs. Those approaches also suffer from information loss due to the vector representation, and high dimensionality if there is a large number of log events. To make up these deficiencies, unlike most traditional methods that used a vector to represent a period behavior at the macro level, we construct a 2D matrix to reveal more detailed system behaviors in the time period by dividing each window into sequential subwindows. To provide a more efficient representation, we further use the ant colony optimization algorithm to find a highly-coupled event template as the horizontal index of the 2D window matrix to replace the disordered one. To capture time dependencies, a multi-module convolutional auto-encoder is configured as that different paralleled modules scan among different time intervals to extract information respectively. These features are then concatenated in latent space as the final input, which contains diversified time information, for classification by SVM. The experiments on Blue Gene/L log dataset showed that our proposed method outperforms the state-of-art SVM method.",https://ieeexplore.ieee.org/document/8616515,True,"[{'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 27}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 27}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 296}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 35}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('anomaly detection' OR 'outlier detection')"", 'index': 108}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 62}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 147}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 225}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 314}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 139}]",0.0,,,,['failure-detection'],,"2018 IEEE International Conference on Systems, Man, and Cybernetics (SMC)",True,['failure-management'],,True,,,,,,,
332,Failure Prediction in IBM BlueGene/L Event Logs,"['Y. Liang', ' Y. Zhang', ' H. Xiong', ' R. Sahoo']",2007,"Frequent failures are becoming a serious concern to the community of high-end computing, especially when the applications and the underlying systems rapidly grow in size and complexity. In order to develop effective fault-tolerant strategies, there is a critical need to predict failure events. To this end, we have collected detailed event logs from IBM BlueGene/L, which has 128 K processors, and is currently the fastest supercomputer in the world. In this study, we first show how the event records can be converted into a data set that is appropriate for running classification techniques. Then we apply classifiers on the data, including RIPPER (a rule-based classifier), Support Vector Machines (SVMs), a traditional Nearest Neighbor method, and a customized Nearest Neighbor method. We show that the customized nearest neighbor approach can outperform RIPPER and SVMs in terms of both coverage and precision. The results suggest that the customized nearest neighbor approach can be used to alleviate the impact of failures.",https://ieeexplore.ieee.org/document/4470294,True,"[{'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 94}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 95}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 52}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 72}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 298}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 108}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 252}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 322}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 40}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 47}]",143.0,"['rule-mining', 'support-vector-machine', 'similarity-matching']",['logs'],['novel-use'],['failure-prediction'],['hpc'],Seventh IEEE International Conference on Data Mining (ICDM 2007),True,['failure-management'],True,,['system-failure-prediction'],True,42.0,,,,
333,Online failure prediction in cloud datacenters by real-time message pattern learning,"['Y. Watanabe', ' H. Otsuka', ' M. Sonoda', ' S. Kikuchi', ' Y. Matsumoto']",2012,"Once failures occur in a cloud datacenter accommodating a large number of virtual resources, they tend to spread rapidly and widely, impacting on many cloud users (tenant owners). One of the best ways to prevent a failure from spreading in the system is identifying signs of the failure before its occurrence and deal with it proactively before it causes serious problems. Although several approaches have been proposed to predict failures by analyzing past system message logs and identifying the relationship between the messages and the failures, it is still difficult to automatically predict the failure for several reasons such as various types of log message formats or time gaps between message pattern learning and application of the identified patterns in real systems. Based on this understanding, we propose a new failure prediction method in this paper which learns message patterns as the signs of failure automatically by classifying messages by their similarity without depending on their format and re-Iearning of message patterns in frequently-changed configurations. We implemented our failure prediction method and evaluated it by using system log data recorded in an actual cloud datacenter. The experimental result shows that our approach predicted failures with 80% precision and covered 90% of failure occurrences.",https://ieeexplore.ieee.org/document/6427566,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 74}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 25}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 33}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 57}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 72}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 63}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 480}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 152}, {'database': 'IEEE', 'search_string': ""'classification' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 80}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 898}]",38.0,,['logs'],,['failure-detection'],,4th IEEE International Conference on Cloud Computing Technology and Science Proceedings,True,['failure-management'],,True,['anomaly-detection'],,,,,,
334,Resource Scheduling in Web Servers in Cloud Computing Using Multiple Artificial Neural Networks,"['F. F. d. Almeida', ' A. d. A. Neto', ' M. M. Teixeira']",2015,"This work presents a solution based on Multiple Artificial Neural Networks for resource scheduling problem in cloud computing. In these computer systems there are still major challenges to achieve a high level of efficiency. One of these challenges is the computational resource scheduling, necessary for an application more efficient of available resources. Thus, with this work it is intentional to show the usage of Artificial Neural Networks to improve the scheduling of such resources in web servers in cloud computing, in order to search an optimization for this process. Experimental results show the usage of Multiple Artificial Neural Networks can obtain, on average, a better response time compared to traditional scheduling algorithms as RoundRobin (22%) and Greedy (8%).",https://ieeexplore.ieee.org/document/7429434,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 239}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 966}]",1.0,,,,,,2015 Fourteenth Mexican International Conference on Artificial Intelligence (MICAI),True,['resource-provisioning'],,,,,,,,,
335,Heterogeneous resource allocation in Cloud Management,"['S. Kadioglu', ' M. Colena', ' S. Sebbah']",2016,"This paper introduces a combinatorial problem arising from real-world business requirements as part of resource allocation in Cloud Management. In particular, we focus on the allocation of a set of heterogeneous resources serving multiple tenants with different service level agreements. There exist certain business rules that govern the application stemming from privacy, performance, and capacity requirements. We show how to formulate the problem as constrained optimization and then solve it efficiently using Artificial Intelligence based constraint propagation. Our approach stands out as a high-level, declarative solution that is efficient and easy to maintain and update.",https://ieeexplore.ieee.org/document/7778589,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 270}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 151}]",2.0,,,,,,2016 IEEE 15th International Symposium on Network Computing and Applications (NCA),True,['resource-provisioning'],,,,,,,,,
336,A set-based discrete PSO for cloud workflow scheduling with user-defined QoS constraints,"['W. Chen', ' J. Zhang']",2012,"Cloud computing has emerged as a powerful computing paradigm that enables users to access computing services anywhere on demand. It provides a flexible way to implement computation-intensive workflow applications on a pay-per-use basis. Since users are more concerned on the satisfaction of Quality of Service (QoS) in cloud systems, the cloud workflow scheduling problem that addresses different QoS requirements of users has become an important and challenging problem for workflow management in cloud computing. In this paper, we tackle a cloud workflow scheduling problem which enables users to define various QoS constraints like the deadline constraint, the budget constraint, and the reliability constraint. It also enables users to specify one preferred QoS parameter as the optimization objective. A set-based PSO (S-PSO) approach is proposed for this scheduling problem. As the allocation of service instances can be regarded as the selection problem from a set of service instances, it is found the set-based representation scheme in S-PSO is natural for the considered problem. In addition, the S-PSO provides an effective way to take advantage of problem-based heuristics to further accelerate search. We define penalty-based fitness functions to address the multiple QoS constraints and integrate the S-PSO with seven heuristics. A discrete version of the comprehensive learning PSO (CLPSO) algorithm based on the S-PSO method is implemented. Experimental results show that the proposed approach is very competitive especially on the instances with tight QoS constraints.",https://ieeexplore.ieee.org/document/6377821,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 758}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 724}]",23.0,['particle-swarm'],,,"['scheduling', 'resource-consolidation']",,"2012 IEEE International Conference on Systems, Man, and Cybernetics (SMC)",True,['resource-provisioning'],,,,,,,,,
337,A self-evolving anomaly detection framework for developing highly dependable utility clouds,"['H. S. Pannu', ' Jianguo Liu', ' S. Fu']",2012,"Utility clouds continue to grow in scale and in the complexity of their components and interactions, which introduces a key challenge to failure and resource management for highly dependable cloud computing. Autonomic anomaly detection is a crucial technique for understanding emergent, cloud-wide phenomena and self-managing cloud resources for system-level dependability assurance. To identify anomalies, we need to monitor the system execution and collect health-related runtime performance data. These data are usually unlabeled and a prior failure history is not always available in production systems, especially for newly deployed or managed utility clouds. In this paper, we present a self-evolving anomaly detection framework with mechanisms for dependability assurance in utility clouds. No prior failure history is required. The detector self-evolves by recursively exploring newly generated verified detection results for future anomaly identification. Statistical learning technologies are exploited in detector determination and working dataset selection. Experimental results in an institute-wide cloud computing system show that the detection accuracy improves as it evolves. With self-evolvement, the detector can achieve 92.1% detection sensitivity and 83.8% detection specificity, which makes it well suitable for building highly dependable utility clouds.",https://ieeexplore.ieee.org/document/6503343,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 801}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1469}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 588}]",13.0,,,,['failure-detection'],,2012 IEEE Global Communications Conference (GLOBECOM),True,['failure-management'],,,,,,,,,
338,Cloud Resource Auto-scaling System Based on Hidden Markov Model (HMM),"['A. Y. Nikravesh', ' S. A. Ajila', ' C. Lung']",2014,"The elasticity characteristic of cloud computing enables clients to acquire and release resources on demand. This characteristic reduces clients' cost by making them pay for the resources they actually have used. On the other hand, clients are obligated to maintain Service Level Agreement (SLA) with their users. One approach to deal with this cost-performance trade-off is employing an auto-scaling system which automatically adjusts application's resources based on its load. In this paper we have proposed an auto-scaling system based on Hidden Markov Model (HMM). We have conducted an experiment on Amazon EC2 infrastructure to evaluate our model. Our results show HMM can generate correct scaling actions in 97% of time. CPU utilization, throughput, and response time are being considered as performance metrics in our experiment.",https://ieeexplore.ieee.org/document/6882012,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 811}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 16}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1006}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 15}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 19}]",4.0,['markov-model'],,['novel-use'],['resource-consolidation'],,2014 IEEE International Conference on Semantic Computing,True,['resource-provisioning'],,True,,,,,,,
339,PREPARE: Predictive Performance Anomaly Prevention for Virtualized Cloud Systems,"['Y. Tan', ' H. Nguyen', ' Z. Shen', ' X. Gu', ' C. Venkatramani', ' D. Rajan']",2012,"Virtualized cloud systems are prone to performance anomalies due to various reasons such as resource contentions, software bugs, and hardware failures. In this paper, we present a novel Predictive Performance Anomaly Prevention (PREPARE) system that provides automatic performance anomaly prevention for virtualized cloud computing infrastructures. PREPARE integrates online anomaly prediction, learning-based cause inference, and predictive prevention actuation to minimize the performance anomaly penalty without human intervention. We have implemented PREPARE on top of the Xen platform and tested it on the NCSU's Virtual Computing Lab using a commercial data stream processing system (IBM System S) and an online auction benchmark (RUBiS). The experimental results show that PREPARE can effectively prevent performance anomalies while imposing low overhead to the cloud infrastructure.",https://ieeexplore.ieee.org/document/6258001,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 843}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 440}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 902}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 386}]",71.0,,,,['failure-detection'],['vm'],2012 IEEE 32nd International Conference on Distributed Computing Systems,True,['failure-management'],,,['anomaly-detection'],,,,,,
340,Virtual Machine Failure Prediction Method Based on AdaBoost-Hidden Markov Model,"['Z. Li', ' L. Liu', ' D. Kong']",2019,"The failure prediction method of virtual machines (VM) guarantees reliability to cloud platforms. However, the uncertainty of VM security state will affect the reliability and task processing capabilities of the entire cloud platform. In this Study, a failure prediction method of VM based on AdaBoost-Hidden Markov Model was proposed to improve the reliability of VMs and overall performance of cloud platforms. This method analyzed the deep relationship between the observation state and the hidden state of the VM through the hidden Markov model, proved the influence of the AdaBoost algorithm on the hidden Markov model (HMM), and realized the prediction of the VM failure state. Results show that the proposed method adapts to the complex dynamic cloud platform environment, can effectively predict the failure state of VMs, and improve the predictive ability of VM security state.",https://ieeexplore.ieee.org/document/8669506,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1090}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 7}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 3}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 1404}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 3}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 2}]",2.0,['markov-model'],,,['failure-prediction'],['vm'],"2019 International Conference on Intelligent Transportation, Big Data & Smart City (ICITBS)",True,['failure-management'],,True,['system-failure-prediction'],,,,,,
341,Resource Management Using Virtual Ontologies for Scientific Applications,"['C. Hur', ' H. Yoo', ' S. Kim', ' Y. Kim']",2009,"Cloud computing has recently paid attention as a way to share the resources to provide scalable and on-demand services. Across various authorities, cloud resources should be configured to a virtual organization (VO) according to user's requirements. Ontology-based representation of cloud computing environment would be able to conceptualize common attributes among cloud resources and to represent semantic relations among them. However, mutual compatibility among different VOs is limited because a method applying ontology to cloud is in progress. We propose to introduce a resource virtualization method using virtual ontology. A new virtual ontology (VOn) is configured dynamically based on requirement of users, and the VOn is mapped to actual cloud resources. Our service uses a map/reduce model for rapid and efficient merging a number of ontology. The execution environment is orchestrated with selected resources mapped to the VOn, which is generated by ontology merge engine.",https://ieeexplore.ieee.org/document/5405670,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 1129}]",,,,,,,Proceedings of the 4th International Conference on Ubiquitous Information Technologies & Applications,True,['resource-provisioning'],,,,,,,,,
342,vNMF: Distributed fault detection using clustering approach for network function virtualization,"['M. Miyazawa', ' M. Hayashi', ' R. Stadler']",2015,"Network function virtualization introduces additional complexity for network management through the use of virtualization environments. The amount of managed data and the operational complexity increases, which makes service assurance and failure recovery harder to realize. In response to this challenge, the paper proposes a distributed management function, called virtualized network management function (vNMF), to detect failures related to virtualized services. vNMF detects the failures by monitoring physical-layer statistics that are processed with a self-organizing map algorithm. Experimental results show that memory leaks and network congestion failures can be successfully detected and that and the accuracy of failure detection can be significantly improved compared to common k-means clustering.",https://ieeexplore.ieee.org/document/7140349,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 29}, {'database': 'IEEE', 'search_string': ""'clustering' AND ('remediation' OR 'recovery')"", 'index': 245}]",15.0,,,,['failure-detection'],,2015 IFIP/IEEE International Symposium on Integrated Network Management (IM),True,['failure-management'],,,,,,,,,
343,Inductive learning for fault diagnosis,"['H. R. Berenji', ' J. Ametha', ' D. Vengerov']",2003,"There is a steadily increasing need for autonomous systems that must be able to function with minimal human intervention to detect and isolate faults, and recover from such faults. In this paper we present a novel hybrid Model based and Data Clustering (MDC) architecture for fault monitoring and diagnosis, which is suitable for complex dynamic systems with continuous and discrete variables. The MDC approach allows for adaptation of both structure and parameters of identified models using supervised and reinforcement learning techniques. The MDC approach will be illustrated using the model and data from the Hybrid Combustion Facility (HCF) at the NASA Ames Research Center.",https://ieeexplore.ieee.org/document/1209453,True,"[{'database': 'IEEE', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 185}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 47}]",14.0,,,,['root-cause-analysis'],,"The 12th IEEE International Conference on Fuzzy Systems, 2003. FUZZ '03.",True,['failure-management'],,,,,,,,,
344,Towards an Autonomic Auto-scaling Prediction System for Cloud Resource Provisioning,"['A. Y. Nikravesh', ' S. A. Ajila', ' C. Lung']",2015,"This paper investigates the accuracy of predictive auto-scaling systems in the Infrastructure as a Service (IaaS) layer of cloud computing. The hypothesis in this research is that prediction accuracy of auto-scaling systems can be increased by choosing an appropriate time-series prediction algorithm based on the performance pattern over time. To prove this hypothesis, an experiment has been conducted to compare the accuracy of time-series prediction algorithms for different performance patterns. In the experiment, workload was considered as the performance metric, and Support Vector Machine (SVM) and Neural Networks (NN) were utilized as time-series prediction techniques. In addition, we used Amazon EC2 as the experimental infrastructure and TPC-W as the benchmark to generate different workload patterns. The results of the experiment show that prediction accuracy of SVM and NN depends on the incoming workload pattern of the system under study. Specifically, the results show that SVM has better prediction accuracy in the environments with periodic and growing workload patterns, while NN outperforms SVM in forecasting unpredicted workload pattern. Based on these experimental results, this paper proposes an architecture for a self-adaptive prediction suite using an autonomic system approach. This suite can choose the most suitable prediction technique based on the performance pattern, which leads to more accurate prediction results.",https://ieeexplore.ieee.org/document/7194655,True,"[{'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 61}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 44}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 49}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 70}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 21}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 19}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 11}]",20.0,,,,"['resource-consolidation', 'workload-prediction']",,2015 IEEE/ACM 10th International Symposium on Software Engineering for Adaptive and Self-Managing Systems,True,['resource-provisioning'],,,,,,,Conference Paper,"['autonomic', 'workload pattern', 'cloud computing', 'support vector machine', 'resource provisioning', 'neural networks', 'auto-scaling']",
345,Discriminative Model for Google Host Load Prediction with Rich Feature Set,"['P. Huang', ' D. Ye', ' Z. Fan', ' P. Huang', ' X. Li']",2015,"Host load prediction is one of the key research issues in Cloud computing. However, due to the drastic fluctuation of the host load in the Cloud, accurately predicting the host load remains a challenge. In this paper, a discriminative model (SVM) is employed to improve upon the accuracy of host load prediction in a Cloud data center. A rich set of features are generated by function based methods and incorporated into discriminative modelling. The performance of our proposed method is empirically evaluated using a one-month trace of a Google data center with over 12000 heterogeneous hosts. The results show that the proposed method achieves a better prediction performance than some state-of-the-art methods.",https://ieeexplore.ieee.org/document/7152619,True,"[{'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 134}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 81}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 116}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 67}, {'database': 'ACM', 'search_string': ""'clustering' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 65}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 26}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 33}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 7}, {'database': 'ACM', 'search_string': ""'regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 57}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 25}]",1.0,['support-vector-machine'],,['novel-use'],['workload-prediction'],,"2015 15th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing",True,['resource-provisioning'],,,,,,10.1109/CCGrid.2015.99,Conference Paper,"['discriminative model', 'Google workload', 'support vector machine', 'host load prediction']",
346,Feature Selection Based on the Kullback-Leibler Distance and its Application on Fault Diagnosis,"['Y. Xue', ' L. Zhang', ' B. Wang', ' F. Li']",2019,"The core concept of pattern recognition is that digs inner mode between data in the same class. The within-class data has a similar distribution, while between-class data has some distinction in different forms. Feature selection utilizes the difference between two-class data to reduce the number of features in the training models. A large amount of feature selection methods have widely used in different fields. This paper proposes a novel feature selection method based on the Kullback-Leibler distance which measures the distance of distribution between two features. For fault diagnosis, the proposed feature selection method is combined with support vector machine to improve its performance. Experimental results validate the effectiveness and superior of the proposed feature selection method, and the proposed diagnosis model can increase the detection rate in chemistry process.",https://ieeexplore.ieee.org/document/8916388,True,"[{'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 366}]",0.0,"['support-vector-machine', 'entropy-selection']",,['novel-use'],['root-cause-analysis'],,2019 Seventh International Conference on Advanced Cloud and Big Data (CBD),True,['failure-management'],,True,['root-cause-diagnosis'],,,,,,
347,Toward comprehensible software defect prediction models using fuzzy logic,['H. A. Al-Jamimi'],2016,"Software defect prediction is a discipline that predicts the defects proneness of future modules. Software metrics are used for this kind of predication. However, the predication metrics are associated with uncertainty, thus the metrics need to be expressed in linguistic terms to overcome ambiguity and uncertainty. Two types of knowledge are utilized as input to the prediction models: software metrics and expert's opinions. This paper proposes a framework for developing fuzzy logic-based software predication model using different set of software metrics. It aims to provide a generic set of metrics to be used for software defects prediction. The performance of the proposed Fuzzy-based models has been validated using real software projects data where Takagi-Sugeno fuzzy inference engine is used to predict software defects. Validation results are satisfactory.",https://ieeexplore.ieee.org/document/7883031,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 2}]",7.0,['fuzzy-logic'],,,['failure-prevention'],['source-code'],2016 7th IEEE International Conference on Software Engineering and Service Science (ICSESS),True,['failure-management'],,,['software-defect-prediction'],,,,,,
348,Predicting Resource Allocation and Costs for Business Processes in the Cloud,"['T. Mastelic', ' W. Fdhila', ' I. Brandic', ' S. Rinderle-Ma']",2015,"By moving business processes into the cloud, business partners can benefit from lower costs, more flexibility and greater scalability in terms of resources offered by the cloud providers. In order to execute a process or a part of it, a business process owner selects and leases feasible resources while considering different constraints, e.g., Optimizing resource requirements and minimizing their costs. In this context, utilizing information about the process models or the dependencies between tasks can help the owner to better manage leased resources. In this paper, we propose a novel resource allocation technique based on the execution path of the process, used to assist the business process owner in efficiently leasing computing resources. The technique comprises three phases, namely process execution prediction, resource allocation and cost estimation. The first exploits the business process model metrics and attributes in order to predict the process execution and the requires resources, while the second utilizes this prediction for efficient allocation of the cloud resources. The final phase estimates and optimizes costs of leased resources by combining different pricing models offered by the provider.",https://ieeexplore.ieee.org/document/7196503,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1914}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 1559}]",9.0,,,,,,2015 IEEE World Congress on Services,True,['resource-provisioning'],,,,,,,,,
349,REPD: Source Code Defect Prediction As Anomaly Detection,"['P. Afric', ' L. Sikic', ' A. S. Kurdija', ' G. Delac', ' M. Silic']",2019,"In this paper, we present a novel approach to defect prediction within project source code. Since defect prediction datasets are typically imbalanced, and there are few defective examples, we treat defect prediction as anomaly detection. We present our Reconstruction Error Probability Distribution (REPD) model and compare it on five different datasets to five standardly used models: Gaussian Naive Bayes, Logistic regression, k-nearest-neighbors, decision tree, and SVM. For the main performance results we use F1-scores. Using statistical means, we show that our model produces significantly better results, improving F1-score up to 10.11%.",https://ieeexplore.ieee.org/document/8859503,True,"[{'database': 'IEEE', 'search_string': ""'logistic regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 22}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 534}, {'database': 'IEEE', 'search_string': ""'regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 186}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('anomaly detection' OR 'outlier detection')"", 'index': 294}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 229}]",0.0,"['naive-bayes', 'logistic-regression', 'similarity-matching', 'support-vector-machine', 'decision-tree']",['source-code'],,['failure-prevention'],['source-code'],"2019 IEEE 19th International Conference on Software Quality, Reliability and Security Companion (QRS-C)",True,['failure-management'],,,['software-defect-prediction'],,,,,,['imbalance']
350,Analyzing Massive Machine Maintenance Data in a Computing Cloud,"['A. Bahga', ' V. K. Madisetti']",2012,"We present a novel framework, CloudView, for storage, processing and analysis of massive machine maintenance data, collected from a large number of sensors embedded in industrial machines, in a cloud computing environment. This paper describes the architecture, design, and implementation of CloudView, and how the proposed framework leverages the parallel computing capability of a computing cloud based on a large-scale distributed batch processing infrastructure that is built of commodity hardware. A case-based reasoning (CBR) approach is adopted for machine fault prediction, where the past cases of failure from a large number of machines are collected in a cloud. A case-base of past cases of failure is created using the global information obtained from a large number of machines. CloudView facilitates organization of sensor data and creation of case-base with global information. Case-base creation jobs are formulated using the MapReduce parallel data processing model. CloudView captures the failure cases across a large number of machines and shares the failure information with a number of local nodes in the form of case-base updates that occur in a time scale of every few hours. At local nodes, the real-time sensor data from a group of machines in the same facility/plant is continuously matched to the cases from the case-base for predicting the incipient faults-this local processing takes a much shorter time of a few seconds. The case-base is updated regularly (in the time scale of a few hours) on the cloud to include new cases of failure, and these case-base updates are pushed from CloudView to the local nodes. Experimental measurements show that fault predictions can be done in real-time (on a timescale of seconds) at the local nodes and massive machine data analysis for case-base creation and updating can be done on a timescale of minutes in the cloud. Our approach, in addition to being the first reported use of the cloud architecture for maintenance data storage, processing and analysis, also evaluates several possible cloud-based architectures that leverage the advantages of the parallel computing capabilities of the cloud to make local decisions with global information efficiently, while avoiding potential data bottlenecks that can occur in getting the maintenance data in and out of the cloud.",https://ieeexplore.ieee.org/document/6104038,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 214}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prediction' OR 'failure prediction')"", 'index': 24}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 208}]",92.0,"['case-based-reasoning', 'similarity-matching']",,['new-method'],['failure-prediction'],['cloud'],IEEE Transactions on Parallel and Distributed Systems,True,['failure-management'],,,['system-failure-prediction'],,,,,,
351,LOGAN: Problem Diagnosis in the Cloud Using Log-Based Reference Models,"['B. C. Tak', ' S. Tao', ' L. Yang', ' C. Zhu', ' Y. Ruan']",2016,"Problem diagnosis is one crucial aspect in the cloud operation that is becoming increasingly challenging. On the one hand, the volume of logs generated in today's cloud is overwhelmingly large. On the other hand, cloud architecture becomes more distributed and complex, which makes it more difficult to troubleshoot failures. In order to address these challenges, we have developed a tool, called LOGAN, that enables operators to quickly identify the log entries that potentially lead to the root cause of a problem. It constructs behavioral reference models from logs that represent the normal patterns. When problem occurs, our tool enables operators to inspect the divergence of current logs from the reference model and highlight logs likely to contain the hints to the root cause. To support these capabilities we have designed and developed several mechanisms. First, we developed log correlation algorithms using various IDs embedded in logs to help identify and isolate log entries that belong to the failed request. Second, we provide efficient log comparison to help understand the differences between different executions. Finally we designed mechanisms to highlight critical log entries that are likely to contain information pertaining to the root cause of the problem. We have implemented the proposed approach in a popular cloud management system, OpenStack, and through case studies, we demonstrate this tool can help operators perform problem diagnosis quickly and effectively.",https://ieeexplore.ieee.org/document/7484164,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 783}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 182}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 1685}]",12.0,['pattern-matching'],['logs'],,"['root-cause-analysis', 'failure-detection']",['openstack'],2016 IEEE International Conference on Cloud Engineering (IC2E),True,['failure-management'],,,['anomaly-detection'],,,,,,
352,PerfCompass: Online Performance Anomaly Fault Localization and Inference in Infrastructure-as-a-Service Clouds,"['D. J. Dean', ' H. Nguyen', ' P. Wang', ' X. Gu', ' A. Sailer', ' A. Kochut']",2016,"Infrastructure-as-a-service clouds are becoming widely adopted. However, resource sharing and multi-tenancy have made performance anomalies a top concern for users. Timely debugging those anomalies is paramount for minimizing the performance penalty for users. Unfortunately, this debugging often takes a long time due to the inherent complexity and sharing nature of cloud infrastructures. When an application experiences a performance anomaly, it is important to distinguish between faults with a global impact and faults with a local impact as the diagnosis and recovery steps forfaults with a global impact or local impact are quite different. In this paper, we present PerfCompass, an online performance anomaly fault debugging tool that can quantify whether a production-run performance anomaly has a global impact or local impact. PerfCompass can use this information to suggest the root cause as either an external fault (e.g., environment-based) or an internal fault (e.g., software bugs). Furthermore, PerfCompass can identify top affected system calls to provide useful diagnostic hints for detailed performance debugging. PerfCompass does not require source code or runtime application instrumentation, which makes it practical for production systems. We have tested PerfCompass by running five common open source systems (e.g., Apache, MySQL, Tomcat, Hadoop, Cassandra) inside a virtualized cloud testbed. Our experiments use a range of common infrastructure sharing issues and real software bugs. The results show that PerfCompass accurately classifies 23 out of the 24 tested cases without calibration and achieves 100 percent accuracy with calibration. PerfCompass provides useful diagnosis hints within several minutes and imposes negligible runtime overhead to the production system during normal execution time.",https://ieeexplore.ieee.org/document/7127024,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 869}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 46}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 602}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('remediation' OR 'recovery')"", 'index': 561}]",8.0,,,,['root-cause-analysis'],,IEEE Transactions on Parallel and Distributed Systems,True,['failure-management'],,,,,,,,,
353,Integrated QoS-aware Resource Provisioning for Parallel and Distributed Applications,"['Z. Li', ' L. Wang', ' Y. Zhang', ' T. Truong-Huu', ' E. S. Lim', ' P. M. Mohan', ' S. Chen', ' S. Ren', ' M. Gurusamy', ' Z. Qin', ' R. S. M. Goh']",2015,"With more parallel and distributed applications moving to Cloud and data centers, it is challenging to provide predictable and controllable resources to multiple tenants, and thus guarantee application performance. In this paper, we propose an integrated QoS-aware resource provisioning platform based on virtualization technology for computing, storage and network resources. Coarse-grained CPU mapping and fine-grained CPU scheduling mechanisms are proposed to enable adjustable computing power. A hierarchical distributed scheduling mechanism is implemented on a scalable storage system to guarantee I/O throughput for individual tenants and applications. A network manager has also been developed to guarantee the data transmission rate. Web-based interface enables users to monitor real time resource utilization and to adjust resource QoS levels on the fly. According to our experimental results, the resource cost can be saved up to 45% without degrading the performance of a distributed data processing benchmark, and the performance of a parallel agent-based simulation can be improved by 91% using the same amount of resources.",https://ieeexplore.ieee.org/document/7395932,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 1981}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 1454}]",60.0,,,,['resource-consolidation'],,2015 IEEE/ACM 19th International Symposium on Distributed Simulation and Real Time Applications (DS-RT),True,['resource-provisioning'],,,,,,,,,
354,Study on network failure prediction based on alarm logs,"['J. Zhong', ' W. Guo', ' Z. Wang']",2016,"To avoid unpredictable losses because of network failure, the reliability of the network needs to be evaluated in some application scenarios. This paper start the network failure prediction research upon 14 months' network alarm logs we collected. The logs are of one Metropolitan area network. The research method is shown as below: firstly, construct features to represent network characteristics by the means of the feature construction method which is based on two levels time windows; secondly, select optimal parameter combination to create the feature files through multiple experiments; thirdly, design and build adaptive failure prediction model according to classification learning methods. Numbers of experiments show that accuracy of predicting whether the network failure takes place in 6 hours is up to 70%, is better than the prediction result of Weibull distribution model obviously; the results of classification prediction for network equipment failure are slightly better than the prediction method on the basis of Weibull distribution. Preliminary research results show that most network failures can be predicted through analyzing previous network running logs and the method proposed in this paper is verified to be with good prediction effect. This method can detect failures in practical application on early stage and reduce unnecessary economic losses.",https://ieeexplore.ieee.org/document/7460337,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 27}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 14}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 253}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 295}]",6.0,,,,['failure-prediction'],,2016 3rd MEC International Conference on Big Data and Smart City (ICBDSC),True,['failure-management'],,,,,,,,,
355,Code Failure Prediction and Pattern Extraction Using LSTM Networks,"['M. Hajiaghayi', ' E. Vahedi']",2019,"In this paper, we use a well-known Deep Learning technique called Long Short Term Memory (LSTM) recurrent neural networks to find sessions that are prone to code failure in applications that rely on telemetry data for system health monitoring. We also use LSTM networks to extract telemetry patterns that lead to a specific code failure. For code failure prediction, we treat the telemetry events, sequence of telemetry events and the outcome of each sequence as words, sentence and sentiment in the context of sentiment analysis, respectively. Our proposed method is able to process a large set of data and can automatically handle edge cases in code failure prediction. We take advantage of Bayesian optimization technique to find the optimal hyper parameters as well as the type of LSTM cells that leads to the best prediction performance. We then introduce the Contributors and Blockers concepts. In this paper, contributors are the set of events that casue a code failure, while blockers are the set of events that each of them individually prevents a code failure from happening, even in presence of one or multiple contributor(s). Once the proposed LSTM model is trained, we use a greedy approach to find the contributors and blockers. To develop and test our proposed method, we use synthetic (simulated) data in the first step. The synthetic data is generated using a number of rules for code failures, as well as a number of rules for preventing a code failure from happening. The trained LSTM model shows over 99% accuracy for detecting code failures in the synthetic data. The results from the proposed method outperform the classical learning models such as Decision Tree and Random Forest. Using the proposed greedy method, we are able to find the contributors and blockers in the synthetic data in more than 90% of the cases, with a performance better than sequential rule and pattern mining algorithms. In the next step, we train and test our proposed LSTM method on real data that we collected from sequences of activities performed by millions of Microsoft Office customers.",https://ieeexplore.ieee.org/document/8848220,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 29}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 3}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 25}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 2}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 3}, {'database': 'arxiv', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 5}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 28}, {'database': 'arxiv', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 1}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 2}]",1.0,"['rule-mining', 'search', 'rnn']",['events'],"['new-method', 'comparison', 'novel-use']",['failure-prediction'],,2019 IEEE Fifth International Conference on Big Data Computing Service and Applications (BigDataService),True,['failure-management'],,,['system-failure-prediction'],,,,,,
356,System-level hardware failure prediction using deep learning,"['X. Sun', ' K. Chakrabarty', ' R. Huang', ' Y. Chen', ' B. Zhao', ' H. Cao', ' Y. Han', ' X. Liang', ' L. Jiang']",2019,"Disk and memory faults are the leading causes of server breakdown. A proactive solution is to predict such hardware failure at the runtime and then isolate the hardware at risk and backup the data. However, the current model-based predictors are incapable of using the discrete time-series data, such as the values of device attributes, which conveys high-level information of the device behavior. In this paper, we propose a novel deep-learning based prediction scheme for system-level hardware failure prediction. We normalize the distribution of samples' attributes from different vendors to make use of diverse training sets. We propose a temporal Convolution Neural Network based model that is insensitive to the noise in the time dimension. Finally, we design a loss function to train the model with extremely imbalanced samples effectively. Experimental results from an open S.M.A.R.T data set and an industrial data set show the effectiveness of the proposed scheme.",https://ieeexplore.ieee.org/document/8806998,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 31}, {'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 9}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 153}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault detection' OR 'failure detection')"", 'index': 74}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 23}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 90}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 13}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 16}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 55}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 9}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 20}]",1.0,['cnn'],['host-metrics'],,['failure-prediction'],"['memory', 'hard-drive']",2019 56th ACM/IEEE Design Automation Conference (DAC),True,['failure-management'],,True,['hardware-failure-prediction'],,,10.1145/3316781.3317918,Conference Paper,"['transfer learning', 'Hardware failure', 'temporal CNN']",['imbalance']
357,A machine learning approach to database failure prediction,"['İ. Karakurt', ' S. Özer', ' T. Ulusinan', ' M. C. Ganiz']",2017,"In this study, we apply machine learning algorithms to predict technical failures that can be encountered in Oracle databases and related services. In order to train machine learning algorithms, data from log files are collected hourly from Oracle database systems and labeled with two classes; normal or abnormal. We use several data science approaches to preprocess and transform the input data from raw format to the format, which can be feed to the algorithms. After the preprocessing, several different machine learning classifiers are trained and evaluated on our datasets. Our results show that warnings that lead to failures which is dubbed as abnormal events can be predicted using supervised machine learning algorithms, in particular, the Random Forest algorithm, with a relatively satisfactory Recall (75.7%) and Precision (84.9%) which is visibly higher than the other classifiers.",https://ieeexplore.ieee.org/document/8093426,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 48}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 24}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 122}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 139}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 722}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 1020}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1207}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 32}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 6}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 992}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 368}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 289}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1307}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1447}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 14}]",1.0,,,,['failure-prediction'],,2017 International Conference on Computer Science and Engineering (UBMK),True,['failure-management'],,,,,,,,,
358,Toward Comprehensible Software Fault Prediction Models Using Bayesian Network Classifiers,"['K. Dejaeger', ' T. Verbraken', ' B. Baesens']",2013,"Software testing is a crucial activity during software development and fault prediction models assist practitioners herein by providing an upfront identification of faulty software code by drawing upon the machine learning literature. While especially the Naive Bayes classifier is often applied in this regard, citing predictive performance and comprehensibility as its major strengths, a number of alternative Bayesian algorithms that boost the possibility of constructing simpler networks with fewer nodes and arcs remain unexplored. This study contributes to the literature by considering 15 different Bayesian Network (BN) classifiers and comparing them to other popular machine learning techniques. Furthermore, the applicability of the Markov blanket principle for feature selection, which is a natural extension to BN theory, is investigated. The results, both in terms of the AUC and the recently introduced H-measure, are rigorously tested using the statistical framework of Demšar. It is concluded that simple and comprehensible networks with less nodes can be constructed using BN classifiers other than the Naive Bayes classifier. Furthermore, it is found that the aspects of comprehensibility and predictive performance need to be balanced out, and also the development context is an item which should be taken into account during model selection.",https://ieeexplore.ieee.org/document/6175912,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 67}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 37}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 39}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 44}]",163.0,"['bayesian-network', 'logistic-regression', 'naive-bayes', 'random-forest']",['code-metrics'],"['comparison', 'novel-use']",['failure-prevention'],,IEEE Transactions on Software Engineering,True,['failure-management'],True,,['software-defect-prediction'],True,4.0,,,,
359,Predicting hardware failure using machine learning,"['A. Chigurupati', ' R. Thibaux', ' N. Lassar']",2016,"The Weibull distribution has historically been the Reliability Engineer's best tool for describing the probability of failures over time [1]. While this technique is very accurate at describing failure distributions for large populations of components, it works very poorly at predicting the time until failure of an individual component. The mean time until failure is often used to predict times until failure of individual components, but this value may vary greatly with actual times until failure. With the advent of machine learning techniques, the ability to learn from past behavior in order to predict future behavior makes it possible to predict an individual component's time until failure much more accurately. In this paper, we explore the predictive abilities of a machine learning technique to improve upon our ability to predict individual component times until failure in advance of actual failure. Once failure is predicted, an impending problem can be fixed before it actually occurs. This paper brings to light a machine learning approach for predicting individual component times until failure that we will show is far more accurate than the traditional MTBF approach. The algorithm built was able to monitor the health of 14 hardware samples and notify us of an impending failure well ahead of actual failure, providing adequate time to fix the problem before actual failure occurred.",https://ieeexplore.ieee.org/document/7448033,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 101}, {'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 86}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 7}]",9.0,['support-vector-machine'],['logs'],['novel-use'],['failure-prediction'],[],2016 Annual Reliability and Maintainability Symposium (RAMS),True,['failure-management'],True,,['hardware-failure-prediction'],,,,,,
360,A combined Bayesian network method for predicting drive failure times from SMART attributes,"['S. Pang', ' Y. Jia', ' R. Stones', ' G. Wang', ' X. Liu']",2016,"Statistical and machine learning methods have been proposed to predict hard drive failure based on SMART attributes, and many achieve good performance. However, these models do not give a good indication as to when a drive will fail, only predicting that it will fail. To this end, we propose a new notion of a drive's health degree based on the remaining working time of hard drive before actual failure occurs. An ensemble learning method is implemented to predict these health degrees: four popular individual classifiers are individually trained and used in a Combined Bayesian Network (CBN). Experiments show that the CBN model can give a health assessment under the proposed definition where drives are predicted to fail no later than their actual failure time 70% or more of the time, while maintaining prediction performance standards at least approximately as good as the individual classifiers.",https://ieeexplore.ieee.org/document/7727837,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 156}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 43}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 152}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 71}]",9.0,['bayesian-network'],['host-metrics'],,['failure-prediction'],['hard-drive'],2016 International Joint Conference on Neural Networks (IJCNN),True,['failure-management'],,,['hardware-failure-prediction'],,,,,,
361,Data Mining-Based Analysis of HPC Center Operations,"['J. Klinkenberg', ' C. Terboven', ' S. Lankes', ' M. S. Müller']",2017,"Size and complexity of contemporary High Performance Computing (HPC) systems increases permanently. While the reliability of a single component and compute node is high, the huge amount of components comprising these systems results in the fact that defects happen regularly. This drives the need to manage failure situations. Common issues are component failures or node soft lock-ups that typically lead to crashes of the user jobs that are scheduled on the affected node, and may cause undesired downtime. One approach to mitigate the impact of such problems is to predict node failures with a sufficient lead time in order to take proactive measures. However, accurate prediction is a challenging task.The literature describes several approaches that focus on gathering and analyzing system event logs in order to create prediction models. In this paper, we present a different approach by using descriptive statistics and supervised machine learning to create a prediction model from monitoring data. Our approach is based on the assumption, that features of a certain time frame before a critical event (i. e., a failure or soft lock-up) can serve as an indicator. Consequently, our model is trained with monitoring data from critical and healthy time frames. The evaluation with standard monitoring data collected from the HPC systems at RWTH Aachen University shows that our classifier is able to locate potentially failing nodes with a 10-fold cross precision of 98% and recall of 91 %.",https://ieeexplore.ieee.org/document/8049014,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 199}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 803}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 304}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1511}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 157}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 45}]",4.0,"['random-forest', 'logistic-regression', 'support-vector-machine', 'multilayer-perceptron', 'decision-tree']","['kpis', 'host-metrics', 'network-metrics']","['comparison', 'novel-use']",['failure-prediction'],,2017 IEEE International Conference on Cluster Computing (CLUSTER),True,['failure-management'],,,['system-failure-prediction'],,,,,,
362,Analyzing the effect of bagged ensemble approach for software fault prediction in class level and package level metrics,"['A. Shanthini', ' R. M. Chandrasekaran']",2014,"Faults in a module tend to cause failure of the software product. These defective modules in the software pose considerable risk by increasing the developing cost and decreasing the customer satisfaction. Hence in a software development life cycle it is very important to predict the faulty modules in the software product. Prediction of the defective modules should be done as early as possible so as to improve software developers' ability to identify the defect-prone modules and focus quality assurance activities such as testing and inspections on those defective modules. For quality assurance activity, it is important to concentrate on the software metrics. Software metrics play a vital role in measuring the quality of software. Many researchers focused on classification algorithm for predicting the software defect. On the other hand, classifiers ensemble can effectively improve classification performance when compared with a single classifier. This paper mainly addresses using ensemble approach of Support Vector Machine (SVM) for fault prediction. Ensemble classifier was examined for Eclipse Package level dataset and NASA KC1 dataset. We showed that proposed ensemble of Support Vector Machine is superior to individual approach for software fault prediction in terms of classification rate through Root Mean Square Error Rate (RMSE), AUC-ROC, ROC curves.",https://ieeexplore.ieee.org/document/7033809,True,"[{'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 11}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 6}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 38}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 18}, {'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 3}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 2}]",3.0,,,,['failure-prediction'],,International Conference on Information Communication and Embedded Systems (ICICES2014),True,['failure-management'],,True,,,,,,,
363,Hard Drive Failure Prediction Using Classification and Regression Trees,"['J. Li', ' X. Ji', ' Y. Jia', ' B. Zhu', ' G. Wang', ' Z. Li', ' X. Liu']",2014,"Some statistical and machine learning methods have been proposed to build hard drive prediction models based on the SMART attributes, and have achieved good prediction performance. However, these models were not evaluated in the way as they are used in real-world data centers. Moreover, the hard drives deteriorate gradually, but these models can not describe this gradual change precisely. This paper proposes new hard drive failure prediction models based on Classification and Regression Trees, which perform better in prediction performance as well as stability and interpretability compared with the state-of the-art model, the Back propagation artificial neural network model. Experiments demonstrate that the Classification Tree (CT) model predicts over 95% of failures at a false alarm rate (FAR) under 0.1% on a real-world dataset containing 25,792 drives. Aiming at the practical application of prediction models, we test them with different drive families, with fewer number of drives, and with different model updating strategies. The CT model still shows steady and good performance. We propose a health degree model based on Regression Tree (RT) as well, which can give the drive a health assessment rather than a simple classification result. Therefore, the approach can deal with warnings raised by the prediction model in order of their health degrees. We implement a reliability model for RAID-6 systems with proactive fault tolerance and show that our CT model can significantly improve the reliability and/or reduce construction and maintenance cost of large-scale storage systems.",https://ieeexplore.ieee.org/document/6903596,True,"[{'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 13}, {'database': 'IEEE', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 12}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 30}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 7}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 71}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 137}]",72.0,"['regression-tree', 'decision-tree']",,['novel-use'],['failure-prediction'],['hard-drive'],2014 44th Annual IEEE/IFIP International Conference on Dependable Systems and Networks,True,['failure-management'],True,,['hardware-failure-prediction'],True,32.0,,,,
364,Predicting Faults in High Assurance Software,"['N. Seliya', ' T. M. Khoshgoftaar', ' J. V. Hulse']",2010,"Reducing the number of latent software defects is a development goal that is particularly applicable to high assurance software systems. For such systems, the software measurement and defect data is highly skewed toward the not-fault-prone program modules, i.e., the number of fault-prone modules is relatively very small. The skewed data problem, also known as class imbalance, poses a unique challenge when training a software quality estimation model. However, practitioners and researchers often build defect prediction models without regard to the skewed data problem. In high assurance systems, the class imbalance problem must be addressed when building defect predictors. This study investigates the roughly balanced bagging (RBBag) algorithm for building software quality models with data sets that suffer from class imbalance. The algorithm combines bagging and data sampling into one technique. A case study of 15 software measurement data sets from different real-world high assurance systems is used in our investigation of the RBBag algorithm. Two commonly used classification algorithms in the software engineering domain, Naive Bayes and C4.5 decision tree, are combined with RBBag for building the software quality models. The results demonstrate that defect prediction models based on the RBBag algorithm significantly outperform models built without any bagging or data sampling. The RBBag algorithm provides the analyst with a tool for effectively addressing class imbalance when training defect predictors during high assurance software development.",https://ieeexplore.ieee.org/document/5634306,True,"[{'database': 'IEEE', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 153}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 90}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1398}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 907}]",16.0,,,,['failure-prediction'],,2010 IEEE 12th International Symposium on High Assurance Systems Engineering,True,['failure-management'],,,,,,,,,
365,LogLens: A Real-Time Log Analysis System,"['B. Debnath', ' M. Solaimani', ' M. A. G. Gulzar', ' N. Arora', ' C. Lumezanu', ' J. Xu', ' B. Zong', ' H. Zhang', ' G. Jiang', ' L. Khan']",2018,"Administrators of most user-facing systems depend on periodic log data to get an idea of the health and status of production applications. Logs report information, which is crucial to diagnose the root cause of complex problems. In this paper, we present a real-time log analysis system called LogLens that automates the process of anomaly detection from logs with no (or minimal) target system knowledge and user specification. In LogLens, we employ unsupervised machine learning based techniques to discover patterns in application logs, and then leverage these patterns along with the real-time log parsing for designing advanced log analytics applications. Compared to the existing systems which are primarily limited to log indexing and search capabilities, LogLens presents an extensible system for supporting both stateless and stateful log analysis applications. Currently, LogLens is running at the core of a commercial log analysis solution handling millions of logs generated from the large-scale industrial environments and reported up to 12096x man-hours reduction in troubleshooting operational problems compared to the manual approach.",https://ieeexplore.ieee.org/document/8416368,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1}, {'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 488}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 661}]",4.0,,,,['failure-detection'],,2018 IEEE 38th International Conference on Distributed Computing Systems (ICDCS),True,['failure-management'],,True,,,,,,,
366,Mining Logs Files for Computing System Management,"['Wei Peng', ' Tao Li', ' Sheng Ma']",2005,"With advancement in science and technology, computing systems become increasingly more difficult to monitor, manage and maintain. Traditional approaches to system management have been largely based on domain experts through a knowledge acquisition process to translate domain knowledge into operating rules and policies. This has been experienced as a cumbersome, labor intensive, and error prone process. There is thus a pressing need for automatic and efficient approaches to monitor and manage complex computing systems. A popular approach to system management is based on analyzing system log files. However, several new aspects of the system log data have been less emphasized in existing analysis methods and posed several challenges. The aspects include disparate formats and relatively short text messages in data reporting, asynchronous data collection, and temporal characteristics in data representation. First, a typical computing system contains different devices with different software components, possibly from different providers. These various components have multiple ways to report events, conditions, errors and alerts. The heterogeneity and inconsistency of log formats make it difficult to automate problem determination. To perform automated analysis, we need to categorize the text messages with disparate formats into common situations. Second, text messages in the log files are relatively short with a large vocabulary size. Third, each text message usually contains a timestamp. The temporal characteristics provide additional context information of the messages and can be used to facilitate data analysis. In this paper, we apply text mining to automatically categorize the messages into a set of common categories, and propose two approaches of incorporating temporal information to improve the categorization performance",https://ieeexplore.ieee.org/document/1498077,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 70}]",9.0,,,,['failure-detection'],,Second International Conference on Autonomic Computing (ICAC'05),True,['failure-management'],,,,,,,,,
367,Experience Report: Log Mining Using Natural Language Processing and Application to Anomaly Detection,"['C. Bertero', ' M. Roy', ' C. Sauvanaud', ' G. Tredan']",2017,"Event logging is a key source of information on a system state. Reading logs provides insights on its activity, assess its correct state and allows to diagnose problems. However, reading does not scale: with the number of machines increasingly rising, and the complexification of systems, the task of auditing systems' health based on logfiles is becoming overwhelming for system administrators. This observation led to many proposals automating the processing of logs. However, most of these proposal still require some human intervention, for instance by tagging logs, parsing the source files generating the logs, etc. In this work, we target minimal human intervention for logfile processing and propose a new approach that considers logs as regular text (as opposed to related works that seek to exploit at best the little structure imposed by log formatting). This approach allows to leverage modern techniques from natural language processing. More specifically, we first apply a word embedding technique based on Google's word2vec algorithm: logfiles' words are mapped to a high dimensional metric space, that we then exploit as a feature space using standard classifiers. The resulting pipeline is very generic, computationally efficient, and requires very little intervention. We validate our approach by seeking stress patterns on an experimental platform. Results show a strong predictive performance (≈ 90% accuracy) using three out-of-the-box classifiers.",https://ieeexplore.ieee.org/document/8109100,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 186}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 614}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 307}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 559}]",8.0,,,,['failure-detection'],,2017 IEEE 28th International Symposium on Software Reliability Engineering (ISSRE),True,['failure-management'],,,,,,,,,
368,Learning to Log: Helping Developers Make Informed Logging Decisions,"['J. Zhu', ' P. He', ' Q. Fu', ' H. Zhang', ' M. R. Lyu', ' D. Zhang']",2015,"Logging is a common programming practice of practical importance to collect system runtime information for postmortem analysis. Strategic logging placement is desired to cover necessary runtime information without incurring unintended consequences (e.g., Performance overhead, trivial logs). However, in current practice, there is a lack of rigorous specifications for developers to govern their logging behaviours. Logging has become an important yet tough decision which mostly depends on the domain knowledge of developers. To reduce the effort on making logging decisions, in this paper, we propose a ""learning to log"" framework, which aims to provide informative guidance on logging during development. As a proof of concept, we provide the design and implementation of a logging suggestion tool, Log Advisor, which automatically learns the common logging practices on where to log from existing logging instances and further leverages them for actionable suggestions to developers. Specifically, we identify the important factors for determining where to log and extract them as structural features, textual features, and syntactic features. Then, by applying machine learning techniques (e.g., Feature selection and classifier learning) and noise handling techniques, we achieve high accuracy of logging suggestions. We evaluate Log Advisor on two industrial software systems from Microsoft and two open-source software systems from Git Hub (totally 19.1M LOC and 100.6K logging statements). The encouraging experimental results, as well as a user study, demonstrate the feasibility and effectiveness of our logging suggestion tool. We believe our work can serve as an important first step towards the goal of ""learning to log"".",https://ieeexplore.ieee.org/document/7194593,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 201}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 185}, {'database': 'ACM', 'search_string': ""'regression' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 95}]",106.0,,,,['failure-detection'],,2015 IEEE/ACM 37th IEEE International Conference on Software Engineering,True,['failure-management'],True,,['log-enhancement'],True,68.0,,Conference Paper,,
369,Improving backfilling by using machine learning to predict running times,"['E. Gaussier', ' D. Glesser', ' V. Reis', ' D. Trystram']",2015,"The job management system is the HPC middleware responsible for distributing computing power to applications. While such systems generate an ever increasing amount of data, they are characterized by uncertainties on some parameters like the job running times. The question raised in this work is: To what extent is it possible/useful to take into account predictions on the job running times for improving the global scheduling? We present a comprehensive study for answering this question assuming the popular EASY backfilling policy. More precisely, we rely on some classical methods in machine learning and propose new cost functions well-adapted to the problem. Then, we assess our proposed solutions through intensive simulations using several production logs. Finally, we propose a new scheduling algorithm that outperforms the popular EASY backfilling algorithm by 28% considering the average bounded slowdown objective.",https://ieeexplore.ieee.org/document/7832838,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 379}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 953}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 85}]",14.0,,,,,,"SC '15: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis",True,['resource-provisioning'],,,,,,10.1145/2807591.2807646,Conference Paper,"['scheduling', 'machine learning', 'running time estimation', 'high performance computing']",
370,Doomsday: Predicting Which Node Will Fail When on Supercomputers,"['A. Das', ' F. Mueller', ' P. Hargrove', ' E. Roman', ' S. Baden']",2018,"Predicting which node will fail and how soon remains a challenge for HPC resilience, yet may pave the way to exploiting proactive remedies before jobs fail. Not only for increasing scalability up to exascale systems but even for contemporary supercomputer architectures does it require substantial efforts to distill anomalous events from noisy raw logs. To this end, we propose a novel phrase extraction mechanism called TBP (time-based phrases) to pin-point node failures, which is unprecedented. Our study, based on real system data and statistical machine learning, demonstrates the feasibility to predict which specific node will fail in Cray systems. TBP achieves no less than 83% recall rates with lead times as high as 2 minutes. This opens up the door for enhancing prediction lead times for supercomputing systems in general, thereby facilitating efficient usage of both computing capacity and power in large scale production systems.",https://ieeexplore.ieee.org/document/8665792,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 805}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1513}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 73}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 74}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault localization' OR 'failure localization')"", 'index': 30}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault localization' OR 'failure localization')"", 'index': 31}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 77}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 78}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 85}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 86}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault localization' OR 'failure localization')"", 'index': 60}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault localization' OR 'failure localization')"", 'index': 62}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 83}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 85}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 37}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 38}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 52}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 53}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 66}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 67}]",1.0,['language-modeling'],"['messages', 'host-metrics', 'logs', 'network-metrics']",,['failure-prediction'],,"SC18: International Conference for High Performance Computing, Networking, Storage and Analysis",True,['failure-management'],,,['system-failure-prediction'],,,10.1109/SC.2018.00012,Conference Paper,"['HPC', 'machine learning', 'failure analysis']",
371,Anomaly Detection Based on Job Monitoring Metrics in Distributed System,"['M. Ding', ' Z. Xiong', ' J. Yu']",2018,"In distributed systems, application delays caused by stragglers become a common problem. And interference by competing the resources can make more stragglers. Previous works mostly focus on straggler detection using statistical analysis methods based on the data extracted from logs. These methods cannot provide fine-grained insights to help users optimize their programs. In this paper, we propose an anomaly detection approach using classification method in machine learning based on job monitoring resource metrics. Due to interference, the change of metrics may vary randomly as the job progresses. In order to compare the metrics in different situation, we extract the job, stage and task information from the logs. From the point of system resource utilization, there are three kinds of anomalies we detect, which are the stragglers(tasks), the abnormal jobs and the interfered nodes. We prove that in most situation, more stragglers happen under interference, and the task time for defining stragglers is longer than that in the similar stage time, as well as the node that the abnormal jobs lived is the interfered nodes.We use the task time in the same stage to label the data for training the adaptive boosting classifier model solely with the resource features. In this way, the model can detect straggles, abnormal jobs and interfered nodes in real-time. Additionally, Experiments show that the accuracy of anomaly detection reaches 92%. Case studies show that our framework is effective in detecting abnormal jobs and interfered nodes.",https://ieeexplore.ieee.org/document/8965641,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1204}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 576}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 825}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1619}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 728}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1057}]",0.0,,,,['failure-detection'],,"2018 International Conference on Security, Pattern Analysis, and Cybernetics (SPAC)",True,['failure-management'],,,,,,,,,
372,Multi-dimensional Knowledge Integration for Efficient Incident Management in a Services Cloud,"['R. Gupta', ' K. H. Prasad', ' L. Luan', ' D. Rosu', ' C. Ward']",2009,"The increasing complexity and dynamics in IT infrastructure and the emerging Cloud services present challenges to timely incident/problem diagnosis and resolution. In this paper we present a problem determination platform with multi-dimensional knowledge integration (e.g. configuration data, system vital data, log data, related tickets) and enablement for efficient incident and problem management of the enterprise. Three features of the platform are discussed: automated ticket classification, the automated association of resource with tickets based on integration with configuration database, and the collection of the system vitals relevant to the ticket through integration with monitoring systems. In response to the emerging Cloud services and their highly dynamic service operation context, we identify the need for a proactive service management approach which incorporates configurations and deployment of incident management tools, policies, and templates throughout the service life cycle in order to enable effective and efficient incident management in service operation.",https://ieeexplore.ieee.org/document/5284021,True,"[{'database': 'IEEE', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 1627}, {'database': 'IEEE', 'search_string': ""'classification' AND ('cloud')"", 'index': 1378}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1592}]",15.0,,,,['remediation'],,2009 IEEE International Conference on Services Computing,True,['failure-management'],,,,,,,,,
373,Cloud Resource Scaling for Big Data Streaming Applications Using a Layered Multi-dimensional Hidden Markov Model,"['O. Runsewe', ' N. Samaan']",2017,"Recent advancements in technology have led to a deluge of data that require real-time analysis with strict latency constraints. A major challenge, however, is determining the amount of resources required by big data stream processing applications in response to heterogeneous data sources, streaming events, unpredictable data volume and velocity changes. Over-provisioning of resources for peak loads can be wasteful while under-provisioning can have a huge impact on the performance of the streaming applications. The majority of research efforts on resource scaling in the cloud are investigated from the cloud provider's perspective, they focus on web applications and do not consider multiple resource bottlenecks. We aim at analyzing the resource scaling problem from a big data streaming application provider's point of view such that efficient scaling decisions can be made for future resource utilization. This paper proposes a Layered Multi-dimensional Hidden Markov Model (LMD-HMM) for facilitating the management of resource auto-scaling for big data streaming applications in the cloud. Our detailed experimental evaluation shows that LMD-HMM performs best with an accuracy of 98%, outperforming the single-layer hidden markov model.",https://ieeexplore.ieee.org/document/7973790,True,"[{'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 12}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 5}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 13}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 80}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 30}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 1}]",2.0,,,,,,"2017 17th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID)",True,['resource-provisioning'],,True,,,,10.1109/CCGRID.2017.147,Conference Paper,"['Cloud Computing', 'Big Data', 'Stream Processing', 'Resource Scaling', 'Resource Prediction', 'Layered Hidden Markov Model']",
374,Workload Analysis for the Scope of User Demand Prediction Model Evaluations in Cloud Environments,"['J. Panneerselvam', ' L. Liu', ' N. Antonopoulos', ' Y. Bo']",2014,"Alongside the healthy development of the Cloud-based technologies across various application deployments, their associated energy consumptions incurred by the excess usage of Information and Communication Technology (ICT) resources, is one of the serious concerns demanding effective solutions with immediate effect. Effective auto scaling of the Cloud resources in accordance to the incoming user demand and thereby reducing the idle resources is one optimum solution which not only reduces the excess energy consumptions but also helps maintaining the Quality of Service (QoS). Whilst achieving such tasks, estimating the user demand in advance with reliable level of accuracy has become an integral and vital component. With this in mind, this research work is aimed at analyzing the Cloud workloads and further evaluating the performances of two widely used prediction techniques such as Markov modelling and Bayesian modelling with 7 hours of Google cluster data. An important outcome of this research work is the categorization and characterization of the Cloud workloads which will assist leading into the user demand prediction parameter modelling.",https://ieeexplore.ieee.org/document/7027611,True,"[{'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 56}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 65}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 19}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 16}]",15.0,,,,,,2014 IEEE/ACM 7th International Conference on Utility and Cloud Computing,True,['resource-provisioning'],,True,,,,10.1109/UCC.2014.144,Conference Paper,"['modelling', 'workloads', 'pattern', 'prediction']",
375,Root cause analysis using artificial intelligence,"['A. Chigurupati', ' N. Lassar']",2017,"A complex hardware system experiences several different types of failure modes when deployed in its field of use. Although each failure mode merits its very own root cause analysis, experience has shown us that a majority of root cause analyses require us to answer a very similar set of questions. These recurring questions often use very similar type of field data to be answered. For example, usage parameters such as hours of operation, total power cycles, vibration fatigue, etc. are used almost always while trying to figure out the failure mechanism. By taking advantage of this, we can automate the process of finding the most likely failure mechanism for a given failure mode. In the current work, we start by computing what we believe are parameters that are most relevant across all different types of failure modes. Once these parameters are computed, we will build a Bayesian Network to model the cause-effect relationship between the degradation parameters (cause) and failure modes(effect) that occur on the field. Bayesian Network is a probabilistic graphical model which concisely describes the relationship between many random variables and their conditional independence via an acyclic directed graph. Once the topology of the Bayesian network is laid out by a domain expert, the conditional probabilities can be learnt from existing data, thereby completely describing the system using the notion of joint probabilities. Two real life field issues are used as examples to demonstrate the accuracy of the network once it is modelled and built. This paper demonstrates that accurately modelling the hardware system as a Bayesian Network substantially accelerates the process of root cause analysis.",https://ieeexplore.ieee.org/document/7889651,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 2}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 1}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 1}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 1}]",3.0,,,,['root-cause-analysis'],,2017 Annual Reliability and Maintainability Symposium (RAMS),True,['failure-management'],,,,,,,,,
376,G-RCA: A Generic Root Cause Analysis Platform for Service Quality Management in Large IP Networks,"['H. Yan', ' L. Breslau', ' Z. Ge', ' D. Massey', ' D. Pei', ' J. Yates']",2012,"An increasingly diverse set of applications, such as Internet games, streaming videos, e-commerce, online banking, and even mission-critical emergency call services, all relies on IP networks. In such an environment, best-effort service is no longer acceptable. This requires a transformation in network management from detecting and replacing individual faulty network elements to managing the end-to-end service quality as a whole. In this paper, we describe the design and development of a Generic Root Cause Analysis platform (G-RCA) for service quality management (SQM) in large IP networks. G-RCA contains a comprehensive service dependency model that incorporates topological and cross-layer relationships, protocol interactions, and control plane dependencies. G-RCA abstracts the root cause analysis process into signature identification for symptom and diagnostic events, temporal and spatial event correlation, and reasoning and inference logic. G-RCA provides a flexible rule specification language that allows operators to quickly customize G-RCA and provide different root cause analysis tools as new problems need to be investigated. G-RCA is also integrated with data trending, manual data exploration, and statistical correlation mining capabilities. G-RCA has proven to be a highly effective SQM platform in several different applications, and we present results regarding BGP flaps, PIM flaps in Multicast VPN service, and end-to-end throughput degradation in content delivery network (CDN) service.",https://ieeexplore.ieee.org/document/6170907,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 13}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 5}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 9}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 9}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 34}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 1}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 2}, {'database': 'ACM', 'search_string': ""'classification' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 4}, {'database': 'ACM', 'search_string': ""'classification' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 7}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 18}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 24}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 9}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 12}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 4}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 6}]",18.0,,,,['root-cause-analysis'],,IEEE/ACM Transactions on Networking,True,['failure-management'],,True,,,,10.1109/TNET.2012.2188837,Journal Article,"['service quality management (SQM)', 'network management', 'root cause analysis (RCA)']",
377,Optimization of fault diagnosis based on the combination of Bayesian Networks and Case-Based Reasoning,"['L. Bennacer', ' L. Ciavaglia', ' A. Chibani', ' Y. Amirat', ' A. Mellouk']",2012,"Fault diagnosis is one of the most important tasks in fault management. The main objective of the fault management system is to detect and localize failures as soon as they occur to minimize their effects on the network performance and therefore on the service quality perceived by users. In this paper, we present a new hybrid approach that combines Bayesian Networks and Case-Based Reasoning to overcome the usual limits of fault diagnosis techniques and reduce human intervention in this process. The proposed mechanism allows identifying the root cause failure with a finer precision and high reliability while reducing the process computation time and taking into account the network dynamicity.",https://ieeexplore.ieee.org/document/6211970,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 20}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 38}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault localization' OR 'failure localization')"", 'index': 16}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 579}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault detection' OR 'failure detection')"", 'index': 557}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 10}]",21.0,['bayesian-network'],,,"['root-cause-analysis', 'failure-detection']",['network'],2012 IEEE Network Operations and Management Symposium,True,['failure-management'],,,,,,,,,
378,Mining Causality of Network Events in Log Data,"['S. Kobayashi', ' K. Otomo', ' K. Fukuda', ' H. Esaki']",2018,"Network log messages (e.g., syslog) are expected to be valuable and useful information to detect unexpected or anomalous behavior in large scale networks. However, because of the huge amount of system log data collected in daily operation, it is not easy to extract pinpoint system failures or to identify their causes. In this paper, we propose a method for extracting the pinpoint failures and identifying their causes from network syslog data. The methodology proposed in this paper relies on causal inference that reconstructs causality of network events from a set of time series of events. Causal inference can filter out accidentally correlated events, thus it outputs more plausible causal events than traditional cross-correlation-based approaches can. We apply our method to 15 months' worth of network syslog data obtained from a nationwide academic network in Japan. The proposed method significantly reduces the number of pseudo correlated events compared with the traditional methods. Also, through three case studies and comparison with trouble ticket data, we demonstrate the effectiveness of the proposed method for practical network operation.",https://ieeexplore.ieee.org/document/8122062,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 25}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 32}]",5.0,,,,['root-cause-analysis'],,IEEE Transactions on Network and Service Management,True,['failure-management'],,,,,,,,,
379,On the Use of Fuzzy Modeling in Virtualized Data Center Management,"['J. Xu', ' M. Zhao', ' J. Fortes', ' R. Carpenter', ' M. Yousif']",2007,"One of the most important goals of data-center management is to reduce cost through efficient use of resources. Virtualization techniques provide the opportunity of carving individual physical servers into multiple virtual containers that can be run and managed separately. A key challenge that comes with virtualization is the simultaneous on-demand provisioning of shared resources to virtual containers and the management of their capacities to meet service quality targets at the least cost. This paper proposes a two-level resource management system with local controllers at the virtual-container level and a global controller at the resource-pool level. Autonomic resource allocation is realized through the interaction of the local and global controllers. A novelty of the controller designs is their use of fuzzy logic to efficiently and robustly deal with the complexity of the virtualized data center and the uncertainties of the dynamically changing workloads. Experimental results obtained through a prototype implementation demonstrate that, for the scenarios under consideration, the proposed resource management system can significantly reduce resource consumption while still achieving application performance targets.",https://ieeexplore.ieee.org/document/4273119,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 12}]",89.0,,,,['resource-consolidation'],,Fourth International Conference on Autonomic Computing (ICAC'07),True,['resource-provisioning'],,,,,,,,,
380,Semi-supervised network traffic classification using deep generative models,"['T. Li', ' S. Chen', ' Z. Yao', ' X. Chen', ' J. Yang']",2018,"Network traffic classification plays a fundamental role in area of network management and security. Recent days, machine learning techniques have been used to classify network traffic. In particular, semi-supervised learning is very fit for practical scenarios, where pre-labelled training flows are hard to obtain. In this paper, a semi-supervised classification scheme is proposed for network traffic classification by using deep generative models. Specifically, the feature extractor module aims to automatically find representation features of raw traffic data in a lower dimensional feature space. Subsequently, using these representation features, a separated classifier is trained by the semi-supervised classification module. The method is verified with three different levels of datasets: Anomaly detection-level, protocol-level and application-level. The results show that our scheme can not only detect malware traffic, but also classify the traffic according to their protocol, application, and attack types. Using less than 20% labelled flows of the whole dataset, we can achieve the accuracy of over 95% which is a satisfying value compared with supervised learning method.",https://ieeexplore.ieee.org/document/8686880,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 191}, {'database': 'IEEE', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 461}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 992}]",1.0,,,,['failure-detection'],,"2018 14th International Conference on Natural Computation, Fuzzy Systems and Knowledge Discovery (ICNC-FSKD)",True,['failure-management'],,True,"['traffic-classification', 'anomaly-detection']",,,,,,
381,Toward Automated Anomaly Identification in Large-Scale Systems,"['Z. Lan', ' Z. Zheng', ' Y. Li']",2010,"When a system fails to function properly, health-related data are collected for troubleshooting. However, it is challenging to effectively identify anomalies from the voluminous amount of noisy, high-dimensional data. The traditional manual approach is time-consuming, error-prone, and even worse, not scalable. In this paper, we present an automated mechanism for node-level anomaly identification in large-scale systems. A set of techniques is presented to automatically analyze collected data: data transformation to construct a uniform data format for data analysis, feature extraction to reduce data size, and unsupervised learning to detect the nodes acting differently from others. Moreover, we compare two techniques, principal component analysis (PCA) and independent component analysis (ICA), for feature extraction. We evaluate our prototype implementation by injecting a variety of faults into a production system at NCSA. The results show that our mechanism, in particular, the one using ICA-based feature extraction, can effectively identify faulty nodes with high accuracy and low computation overhead.",https://ieeexplore.ieee.org/document/4815224,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 366}]",58.0,,,,['failure-detection'],,IEEE Transactions on Parallel and Distributed Systems,True,['failure-management'],,True,,,,,,,
382,Exploring machine learning techniques for fault localization,"['L. C. Ascari', ' L. Y. Araki', ' A. R. T. Pozo', ' S. R. Vergilio']",2009,"Debugging is the most important task related to the testing activity. It has the goal of locating and removing a fault after a failure occurred during test. However, it is not a trivial task and generally consumes effort and time. Debugging techniques generally use testing information but usually they are very specific for certain domains, languages and development paradigms. Because of this, a neural network (NN) approach has been investigated with this goal. It is independent of the context and presented promising results for procedural code. However it was not validated in the context of object-oriented (OO) applications. In addition to this, the use of other machine learning techniques is also interesting, because they can be more efficient. With this in mind, the present work adapts the NN approach to the OO context and also explores the use of support vector machines (SVMs). Results from the use of both techniques are presented and analysed. They show that their use contributes for easing the fault localization task.",https://ieeexplore.ieee.org/document/4813783,True,"[{'database': 'IEEE', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault localization' OR 'failure localization')"", 'index': 8}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 36}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault localization' OR 'failure localization')"", 'index': 38}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 12}]",8.0,,,,['root-cause-analysis'],,2009 10th Latin American Test Workshop,True,['failure-management'],,False,,,,,,,
383,Multiple Fault Localization Using Constraint Programming and Pattern Mining,"['N. Aribi', ' M. Maamar', ' N. Lazaar', ' Y. Lebbah', ' S. Loudni']",2017,"Fault localization problem is one of the most difficult processes in software debugging. The current constraint-based approaches draw strength from declarative data mining and allow to consider the dependencies between statements with the notion of patterns. Tackling large faulty programs is clearly a challenging issue for Constraint Programming (CP) approaches. Programs with multiple faults raise numerous issues due to complex dependencies between faults, making the localization quite complex for all of the current localization approaches. In this paper, we provide a new CP model with a global constraint to speed-up the resolution and we improve the localization to be able to tackle multiple faults. Finally, we give an experimental evaluation that shows that our approach improves on CP and standard approaches.",https://ieeexplore.ieee.org/document/8372037,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 10}]",1.0,,,,['root-cause-analysis'],,2017 IEEE 29th International Conference on Tools with Artificial Intelligence (ICTAI),True,['failure-management'],,,,,,,,,
384,Effective Software Fault Localization Using an RBF Neural Network,"['W. E. Wong', ' V. Debroy', ' R. Golden', ' X. Xu', ' B. Thuraisingham']",2012,"We propose the application of a modified radial basis function neural network in the context of software fault localization, to assist programmers in locating bugs effectively. This neural network is trained to learn the relationship between the statement coverage information of a test case and its corresponding execution result, success or failure. The trained network is then given as input a set of virtual test cases, each covering a single statement. The output of the network, for each virtual test case, is considered to be the suspiciousness of the corresponding covered statement. A statement with a higher suspiciousness has a higher likelihood of containing a bug, and thus statements can be ranked in descending order of their suspiciousness. The ranking can then be examined one by one, starting from the top, until a bug is located. Case studies on 15 different programs were conducted, and the results clearly show that our proposed technique is more effective than several other popular, state of the art fault localization techniques. Further studies investigate the robustness of the proposed technique, and illustrate how it can easily be applied to programs with multiple bugs as well.",https://ieeexplore.ieee.org/document/6058639,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 24}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault localization' OR 'failure localization')"", 'index': 9}]",106.0,,,,['root-cause-analysis'],['software'],IEEE Transactions on Reliability,True,['failure-management'],,,['fault-localization'],,,,,,
385,Using Hidden Semi-Markov Models for Effective Online Failure Prediction,"['F. Salfner', ' M. Malek']",2007,"A proactive handling of faults requires that the risk of upcoming failures is continuously assessed. One of the promising approaches is online failure prediction, which means that the current state of the system is evaluated in order to predict the occurrence of failures in the near future. More specifically, we focus on methods that use event-driven sources such as errors. We use hidden semi-Markov models (HSMMs)for this purpose and demonstrate effectiveness based on field data of a commercial telecommunication system. For comparative analysis we selected three well-known failure prediction techniques: a straightforward method that is based on a reliability model, dispersion frame technique by Lin and Siewiorek and the eventset-based method introduced by Vilalta et al. We assess and compare the methods in terms of precision, recall, F-measure, false-positive rate, and computing time. The experiments suggest that our HSMM approach is very effective with respect to online failure prediction.",https://ieeexplore.ieee.org/document/4365693,True,"[{'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 15}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 115}]",165.0,['markov-model'],['logs'],,['failure-prediction'],"['software', 'node']",2007 26th IEEE International Symposium on Reliable Distributed Systems (SRDS 2007),True,['failure-management'],True,,['system-failure-prediction'],True,44.0,,,,
386,Hard Drive Failure Prediction Using Big Data,"['W. Yang', ' D. Hu', ' Y. Liu', ' S. Wang', ' T. Jiang']",2015,"We design a general framework named Hdoctor for hard drive failure prediction. Hdoctor leverages the power of big data to achieve a significant improvement comparing to all previous researches that used sophisticated machine learning algorithms. Hdoctor exhibits a series of engineering innovations: (1) constructing time dependent features to characterize the Self-Monitoring, Analysis and Reporting Technology (SMART) value transitions during disk failures, (2) combining features to enable the model to learn the correlation among different SMART attributes, (3) regarding circumstance data such as cluster workload, temperature, humidity, location as related features. Meanwhile, Hdoctor collects/labels samples and updates model automatically, and works well for all kinds of disk failure prediction in our intelligent data center. In this work, we use Hdoctor to collect 74,477,717 training records from our clusters involving 220,022 disks. By training a simple and scalable model, our system achieves a detection rate of 97.82%, with a false alarm rate (FAR) of 0.3%, which hugely outperforms all previous algorithms. In addition, Hdoctor is an excellent indicator for how to predict different hardware failures efficiently under various circumstances.",https://ieeexplore.ieee.org/document/7371435,True,"[{'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 40}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 116}]",5.0,,,,['failure-prediction'],,2015 IEEE 34th Symposium on Reliable Distributed Systems Workshop (SRDSW),True,['failure-management'],,,,,,,,,
387,MFL: Method-Level Fault Localization with Causal Inference,"['G. Shu', ' B. Sun', ' A. Podgurski', ' F. Cao']",2013,"Recent studies have shown that use of causal inference techniques for reducing confounding bias improves the effectiveness of statistical fault localization (SFL) at the level of program statements. However, with very large programs and test suites, the overhead of statement-level causal SFL may be excessive. Moreover cost evaluations of statement-level SFL techniques generally are based on a questionable assumption-that software developers can consistently recognize faults when examining statements in isolation. To address these issues, we propose and evaluate a novel method-level SFL technique called MFL, which is based on causal inference methodology. In addition to reframing SFL at the method level, our technique incorporates a new algorithm for selecting covariates to use in adjusting for confounding bias. This algorithm attempts to ensure that such covariates satisfy the conditional exchangeability and positivity properties required for identifying causal effects with observational data. We present empirical results indicating that our approach is more effective than four method-level versions of well-known SFL techniques and that our confounder selection algorithm is superior to two alternatives.",https://ieeexplore.ieee.org/document/6569724,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 8}]",34.0,,,,['root-cause-analysis'],['software'],"2013 IEEE Sixth International Conference on Software Testing, Verification and Validation",True,['failure-management'],,,['fault-localization'],,,,,,
388,Towards automatic network fault localization in real time using probabilistic inference,"['A. Johnsson', ' C. Meirosu']",2013,"This paper describes the foundation of a novel network fault localization algorithm based on active network measurements and probabilistic inference. A fault condition could be an unacceptable large delay or packet loss rate. The solution is computationally efficient, autonomic in nature and provides the operator with a probability mass distribution that indicates the fault location. The probabilistic inference is based on a discrete state-space particle filter. We present results from a first feasibility study performed in a simulated environment. The precise location of single faults is determined rapidly when measurement paths are partly overlapping.",https://ieeexplore.ieee.org/document/6573198,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 13}]",1.0,['markov-model'],,,['root-cause-analysis'],['network'],2013 IFIP/IEEE International Symposium on Integrated Network Management (IM 2013),True,['failure-management'],,,,,,,,,
389,Probabilistic fault localization in communication systems using belief networks,"['M. Steinder', ' A. S. Sethi']",2004,"We apply Bayesian reasoning techniques to perform fault localization in complex communication systems while using dynamic, ambiguous, uncertain, or incorrect information about the system structure and state. We introduce adaptations of two Bayesian reasoning techniques for polytrees, iterative belief updating, and iterative most probable explanation. We show that these approximate schemes can be applied to belief networks of arbitrary shape and overcome the inherent exponential complexity associated with exact Bayesian reasoning. We show through simulation that our approximate schemes are almost optimally accurate, can identify multiple simultaneous faults in an event driven manner, and incorporate both positive and negative information into the reasoning process. We show that fault localization through iterative belief updating is resilient to noise in the observed symptoms and prove that Bayesian reasoning can now be used in practice to provide effective fault localization.",https://ieeexplore.ieee.org/document/1344005,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 61}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault localization' OR 'failure localization')"", 'index': 39}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 1629}, {'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault detection' OR 'failure detection')"", 'index': 1984}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 9}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault localization' OR 'failure localization')"", 'index': 10}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 21}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault localization' OR 'failure localization')"", 'index': 14}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault localization' OR 'failure localization')"", 'index': 45}]",115.0,['bayesian-network'],,,['root-cause-analysis'],,IEEE/ACM Transactions on Networking,True,['failure-management'],,,['fault-localization'],,,10.1109/TNET.2004.836121,Journal Article,"['root cause diagnosis', 'probabilistic inference', 'fault localization']",
390,The Probabilistic Program Dependence Graph and Its Application to Fault Diagnosis,"['G. K. Baah', ' A. Podgurski', ' M. J. Harrold']",2010,"This paper presents an innovative model of a program's internal behavior over a set of test inputs, called the probabilistic program dependence graph (PPDG), which facilitates probabilistic analysis and reasoning about uncertain program behavior, particularly that associated with faults. The PPDG construction augments the structural dependences represented by a program dependence graph with estimates of statistical dependences between node states, which are computed from the test set. The PPDG is based on the established framework of probabilistic graphical models, which are used widely in a variety of applications. This paper presents algorithms for constructing PPDGs and applying them to fault diagnosis. The paper also presents preliminary evidence indicating that a PPDG-based fault localization technique compares favorably with existing techniques. The paper also presents evidence indicating that PPDGs can be useful for fault comprehension.",https://ieeexplore.ieee.org/document/5374423,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 119}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 40}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault localization' OR 'failure localization')"", 'index': 31}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 55}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault localization' OR 'failure localization')"", 'index': 44}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault localization' OR 'failure localization')"", 'index': 84}]",48.0,,,,['root-cause-analysis'],,IEEE Transactions on Software Engineering,True,['failure-management'],,True,,,,10.1145/1390630.1390654,Conference Paper,"['probabilistic graphical models', 'program analysis', 'fault diagnosis', 'machine learning']",
391,Request success rate of multipathing I/O with a paired storage controller,"['G. Enagandula', ' V. Apte', ' B. Raj']",2013,"The success probability of I/O requests in presence of failures is increased by a combination of failover mechanisms built into the storage server, multiple access paths from I/O clients to the server, and timeout-retry mechanisms at the client itself. We define and evaluate a unified availability metric, request failures per million (RFPM), which quantifies request failure probability while taking into account client-side as well as server-side mechanisms. We calculate this metric using a two-level model of I/O service - a probability tree that captures the I/O driver behaviour, and a set of CTMC (Continuous Time Markov Chain) models that capture failover mechanisms at the server. The I/O driver model captures detailed timeout-retry mechanisms including retries at multiple ports (“multipathing”). The server model captures transient phenomena such as failure detection, takeover and emulation behaviour of a paired storage controller. The model shows that client retry mechanisms provide significant improvement in request success probability. The model is then used to study the sensitivity of RFPMs to parameters such as timeouts, reboot time and failure detection delay. The results show that the model can help in answering several what-if questions related to how system parameters impact request success rate.",https://ieeexplore.ieee.org/document/6698906,True,"[{'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault detection' OR 'failure detection')"", 'index': 517}]",52.0,,,,['failure-detection'],,2013 IEEE 24th International Symposium on Software Reliability Engineering (ISSRE),True,['failure-management'],,,['anomaly-detection'],,,,,,['transient']
392,Failure Prediction of Data Centers Using Time Series and Fault Tree Analysis,"['T. Chalermarrewong', ' T. Achalakul', ' S. C. W. See']",2012,"This paper proposes a framework for online failure prediction of data centers. A data center often has a high failure rate as it features a number of servers and components. Moreover, long running applications and intensive workloads are common in such facilities. Performance of the system depends on the availability of the machines, which can be easily compromised if failure cannot be handled gracefully. The main idea of this paper is to create an effective prediction model focusing on hardware failure. Accurate prediction may enhance the overall system performance. In this work, we employ two methods, namely, ARMA (Auto Regressive Moving Average) and Fault Tree Analysis. Experiments were then performed on a simulated cluster built based on Simi's platform. The results show prediction accuracy of 97%, which is very high. We thus believe that our framework is practical and can be adapted to use in data centers in the future.",https://ieeexplore.ieee.org/document/6413603,True,"[{'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 6}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 81}]",33.0,"['fault-tree', 'autoregression']",,['novel-use'],['failure-prediction'],['hardware'],2012 IEEE 18th International Conference on Parallel and Distributed Systems,True,['failure-management'],True,True,['hardware-failure-prediction'],True,45.0,,,,
393,Predicting Node Failure in High Performance Computing Systems from Failure and Usage Logs,"['N. Nakka', ' A. Agrawal', ' A. Choudhary']",2011,"In this paper, we apply data mining classification schemes to predict failures in a high performance computer system. Failure and Usage data logs collected on supercomputing clusters at Los Alamos National Laboratory (LANL) were used to extract instances of failure information. For each failure instance, past and future failure information is accumulated -- time of usage, system idle time, time of unavailability, time since last failure, time to next failure. We performed two separate analyses, with and without classifying the failures based on their root cause. Based on this data, we applied some popular decision tree classifiers to predict if a failure would occur within 1 hour. Our experiments show that our prediction system predicts failures with a high-degree of precision up to 73% and recall of about 80%. We also observed that employing the usage data along with the failure data has improved the accuracy of prediction.",https://ieeexplore.ieee.org/document/6009015,True,"[{'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('remediation' OR 'recovery')"", 'index': 779}, {'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 291}, {'database': 'IEEE', 'search_string': ""'classification' AND ('remediation' OR 'recovery')"", 'index': 548}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 446}]",14.0,,,,['failure-prediction'],,2011 IEEE International Symposium on Parallel and Distributed Processing Workshops and Phd Forum,True,['failure-management'],,,,,,,,,
394,Autonomic performance and power control for co-located Web applications on virtualized servers,"['P. Lama', ' Y. Guo', ' X. Zhou']",2013,"In a data center, various components of Web applications co-located on virtualized servers exhibit complex time-varying interactions and interference. It has a significant impact on the user perceived performance and power consumption of the underlying system. We propose and develop APPLEware, an autonomic middleware for joint performance and power control of co-located Web applications. It features a distributed control structure that provides performance assurance and energy efficiency for large complex systems. It applies machine learning based self-adaptive modeling to capture the complex and time-varying relationship between the application performance and allocation of resources to various application components, in the presence of highly dynamic and bursty workloads and inter-application performance interference. The distributed controllers perform coordinated resource allocation to meet the service level agreements of applications in an agile and energy-efficient manner. Experimental results based on a testbed implementation with benchmark applications demonstrate APPLEware's effectiveness and energy efficiency.",https://ieeexplore.ieee.org/document/6550266,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 89}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 165}]",16.0,,,,,,2013 IEEE/ACM 21st International Symposium on Quality of Service (IWQoS),True,['resource-provisioning'],,,,,,,,,
395,Integrating VM selection criteria in distributed dynamic VM consolidation using Fuzzy Q-Learning,"['S. S. Masoumzadeh', ' H. Hlavacs']",2013,"Distributed dynamic VM consolidation can be an effective strategy to improve energy efficiency in cloud environments. In general, this strategy can be decomposed into four decision-making tasks: (1) Host overloading detection, (2) VM selection, (3) Host underloading detection, and (4) VM placement. The goal is to consolidate virtual machines dynamically in a way that optimizes the energy-performance tradeoff online. In fact, this goal is achieved when each of the aforementioned decisions are made in an optimized fashion. In this paper we concentrate on the VM selection task and propose a Fuzzy Q-Learning (FQL) technique so as to make optimal decisions to select virtual machines for migration. We validate our approach with the CloudSim toolkit using real world PlanetLab workload. Experimental results show that using FQL yields far better results w.r.t. the energy-performance trade-off in cloud data centers in comparison to state of the art algorithms.",https://ieeexplore.ieee.org/document/6727854,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 123}]",15.0,,,,,,Proceedings of the 9th International Conference on Network and Service Management (CNSM 2013),True,['resource-provisioning'],,,,,,,,,
396,Software fault detection for reliability using recurrent neural network modeling,['A. Jomeiri'],2010,"Software fault detection is an important factor for quantitatively characterizing software quality. One of the proposed methods for software fault detection is neural networks. Fault detection is actually a pattern recognition task. Faulty and fault free data are different patterns which must be recognized. In this paper we propose a new framework for modeling software testing and fault detection in applications. Recurrent neural network architecture is used to improve performance of the system. Based on the experiments performed on the software reliability data obtained from middle-sized application software, it is observed that the non-linear RNN can be effective and efficient for software faults detection.",https://ieeexplore.ieee.org/document/5608831,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 26}]",0.0,['rnn'],,['novel-use'],['failure-prevention'],,2010 2nd International Conference on Software Technology and Engineering,True,['failure-management'],,True,['software-defect-prediction'],,,,,,
397,"""The Tail Wags the Dog"": A Study of Anomaly Detection in Commercial Application Performance","['R. Gow', ' S. Venugopal', ' P. K. Ray']",2013,"The IT industry needs systems management models that leverage available application information to detect quality of service, scalability and health of service. Ideally this technique would be common for varying application types with different n-tier architectures under normal production conditions of varying load, user session traffic, transaction type, transaction mix, and hosting environment. This paper shows that a whole of service measurement paradigm utilizing a black box M/M/1 queuing model and auto regression curve fitting of the associated CDF are an accurate model to characterize system performance signatures. This modeling method is used to detect application slow down events. The method did not rely on customizations specific to the n-tier architecture of the systems being analyzed and so the performance anomaly detection technique was shown to be platform and configuration agnostic.",https://ieeexplore.ieee.org/document/6730786,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 107}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 106}]",2.0,,,,['failure-detection'],,"2013 IEEE 21st International Symposium on Modelling, Analysis and Simulation of Computer and Telecommunication Systems",True,['failure-management'],,True,,,,,,,
398,Anomaly? application change? or workload change? towards automated detection of application performance anomaly and change,"['L. Cherkasova', ' K. Ozonat', ' Ningfang Mi', ' J. Symons', ' E. Smirni']",2008,"Automated tools for understanding application behavior and its changes during the application life-cycle are essential for many performance analysis and debugging tasks. Application performance issues have an immediate impact on customer experience and satisfaction. A sudden slowdown of enterprise-wide application can effect a large population of customers, lead to delayed projects and ultimately can result in company financial loss. We believe that online performance modeling should be a part of routine application monitoring. Early, informative warnings on significant changes in application performance should help service providers to timely identify and prevent performance problems and their negative impact on the service. We propose a novel framework for automated anomaly detection and application change analysis. It is based on integration of two complementary techniques: i) a regression-based transaction model that reflects a resource consumption model of the application, and ii) an application performance signature that provides a compact model of run-time behavior of the application. The proposed integrated framework provides a simple and powerful solution for anomaly detection and analysis of essential performance changes in application behavior. An additional benefit of the proposed approach is its simplicity: it is not intrusive and is based on monitoring data that is typically available in enterprise production environments.",https://ieeexplore.ieee.org/document/4630116,True,"[{'database': 'IEEE', 'search_string': ""'regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 249}]",53.0,,,,['failure-detection'],,2008 IEEE International Conference on Dependable Systems and Networks With FTCS and DCC (DSN),True,['failure-management'],,,,,,,,,
399,Mining causes of network events in log data with causal inference,"['S. Kobayashi', ' K. Fukuda', ' H. Esaki']",2017,"Network log message (e.g., syslog) is valuable information to detect unexpected or anomalous behavior in a large scale network. However, pinpointing failures and their causes is not an easy problem because of a huge amount of system log data in daily operation. In this study, we propose a method extracting failures and their causes from network syslog data. The main idea of the method relies on causal inference that reconstructs causality of network events from a set of the time series of events. Causal inference allows us to reduce the number of correlated events by chance, thus it outputs more plausible causal events than a traditional cross-correlation based approach. We apply our method to 15 months network syslog data obtained in a nation-wide academic network in Japan. Our method significantly reduces the number of pseudo correlated events compared with the traditional method. Also, through two case studies and comparison with trouble ticket data, we demonstrate the effectiveness of our method for network operation.",https://ieeexplore.ieee.org/document/7987263,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 40}]",4.0,"['correlation', 'graph-mining']",['logs'],,['root-cause-analysis'],,2017 IFIP/IEEE Symposium on Integrated Network and Service Management (IM),True,['failure-management'],,,"['root-cause-diagnosis', 'rca-others']",,,,,,
400,Contextual analysis of program logs for understanding system behaviors,"['Q. Fu', ' J. Lou', ' Q. Lin', ' R. Ding', ' D. Zhang', ' T. Xie']",2013,"Understanding the behaviors of a software system is very important for performing daily system maintenance tasks. In practice, one way to gain knowledge about the runtime behavior of a system is to manually analyze system logs collected during the system executions. With the increasing scale and complexity of software systems, it has become challenging for system operators to manually analyze system logs. To address these challenges, in this paper, we propose a new approach for contextual analysis of system logs for understanding a system's behaviors. In particular, we first use execution patterns to represent execution structures reflected by a sequence of system logs, and propose an algorithm to mine execution patterns from the program logs. The mined execution patterns correspond to different execution paths of the system. Based on these execution patterns, our approach further learns essential contextual factors (e.g., the occurrences of specific program logs with specific parameter values) that cause a specific branch or path to be executed by the system. The mining and learning results can help system operators to understand a software system's runtime execution logic and behaviors during various tasks such as system problem diagnosis. We demonstrate the feasibility of our approach upon two real-world software systems (Hadoop and Ethereal).",https://ieeexplore.ieee.org/document/6624054,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 99}, {'database': 'IEEE', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 230}]",3.0,,,,['root-cause-analysis'],,2013 10th Working Conference on Mining Software Repositories (MSR),True,['failure-management'],,True,,,,,,,
401,"Next Stop ""NoOps"": Enabling Cross-System Diagnostics Through Graph-Based Composition of Logs and Metrics","['M. Zasadziński', ' M. Solé', ' A. Brandon', ' V. Muntés-Mulero', ' D. Carrera']",2018,"Performing diagnostics in IT systems is an increasingly complicated task, and it is not doable in satisfactory time by even the most skillful operators. Systems and their architecture change very rapidly in response to business and user demand. Many organizations see value in the maintenance and management model of NoOps that stands for No Operations. One of the implementations of this model is a system that is maintained automatically without any human intervention. The path to NoOps involves not only precise and fast diagnostics but also reusing as much knowledge as possible after the system is reconfigured or changed. The biggest challenge is to leverage knowledge on one IT system and reuse this knowledge for diagnostics of another, different system. We propose a framework of weighted graphs which can transfer knowledge, and perform high-quality diagnostics of IT systems. We encode all possible data in a graph representation of a system state and automatically calculate weights of these graphs. Then, thanks to the evaluation of similarity between graphs, we transfer knowledge about failures from one system to another and use it for diagnostics. We successfully evaluate the proposed approach on Spark, Hadoop, Kafka and Cassandra systems.",https://ieeexplore.ieee.org/document/8514882,True,"[{'database': 'IEEE', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 484}, {'database': 'IEEE', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1492}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1865}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 832}]",1.0,['graph-mining'],,['new-method'],['root-cause-analysis'],,2018 IEEE International Conference on Cluster Computing (CLUSTER),True,['failure-management'],,,['root-cause-diagnosis'],,,,,,
402,Tracking Probabilistic Correlation of Monitoring Data for Fault Detection in Complex Systems,"['Zhen Guo', ' Guofei Jiang', ' Haifeng Chen', ' K. Yoshihira']",2006,"Due to their growing complexity, it becomes extremely difficult to detect and isolate faults in complex systems. While large amount of monitoring data can be collected from such systems for fault analysis, one challenge is how to correlate the data effectively across distributed systems and observation time. Much of the internal monitoring data reacts to the volume of user requests accordingly when user requests flow through distributed systems. In this paper, we use Gaussian mixture models to characterize probabilistic correlation between flow-intensities measured at multiple points. A novel algorithm derived from expectation-maximization (EM) algorithm is proposed to learn the ""likely"" boundary of normal data relationship, which is further used as an oracle in anomaly detection. Our recursive algorithm can adaptively estimate the boundary of dynamic data relationship and detect faults in real time. Our approach is tested in a real system with injected faults and the results demonstrate its feasibility",https://ieeexplore.ieee.org/document/1633515,True,"[{'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('anomaly detection' OR 'outlier detection')"", 'index': 334}, {'database': 'IEEE', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 118}]",19.0,,,,['failure-detection'],,International Conference on Dependable Systems and Networks (DSN'06),True,['failure-management'],,,,,,,,,
403,A comprehensive model for software rejuvenation,"['K. Vaidyanathan', ' K. S. Trivedi']",2005,"Recently, the phenomenon of software aging, one in which the state of the software system degrades with time, has been reported. This phenomenon, which may eventually lead to system performance degradation and/or crash/hang failure, is the result of exhaustion of operating system resources, data corruption, and numerical error accumulation. To counteract software aging, a technique called software rejuvenation has been proposed, which essentially involves occasionally terminating an application or a system, cleaning its internal state and/or its environment, and restarting it. Since rejuvenation incurs an overhead, an important research issue is to determine optimal times to initiate this action. In this paper, we first describe how to include faults attributed to software aging in the framework of Gray's software fault classification (deterministic and transient), and study the treatment and recovery strategies for each of the fault classes. We then construct a semi-Markov reward model based on workload and resource usage data collected from the UNIX operating system. We identify different workload states using statistical cluster analysis, estimate transition probabilities, and sojourn time distributions from the data. Corresponding to each resource, a reward function is then defined for the model based on the rate of resource depletion in each state. The model is then solved to obtain estimated times to exhaustion for each resource. The result from the semi-Markov reward model are then fed into a higher-level availability model that accounts for failure followed by reactive recovery, as well as proactive recovery. This comprehensive model is then used to derive optimal rejuvenation schedules that maximize availability or minimize downtime cost.",https://ieeexplore.ieee.org/document/1453531,True,"[{'database': 'IEEE', 'search_string': ""'classification' AND ('remediation' OR 'recovery')"", 'index': 1080}]",270.0,['markov-model'],['host-metrics'],['new-method'],['failure-prevention'],,IEEE Transactions on Dependable and Secure Computing,True,['failure-management'],True,True,['rejuvenation'],True,21.0,,,,
404,TRACON: Interference-aware scheduling for data-intensive applications in virtualized environments,"['R. C. Chiang', ' H. H. Huang']",2011,"Large-scale data centers leverage virtualization technology to achieve excellent resource utilization, scalability, and high availability. Ideally, the performance of an application running inside a virtual machine (VM) shall be independent of co-located applications and VMs that share the physical machine. However, adverse interference effects exist and are especially severe for data-intensive applications in such virtualized environments. In this work, we present TRACON, a novel Task and Resource Allocation CONtrol framework that mitigates the interference effects from concurrent data-intensive applications and greatly improves the overall system performance. TRACON utilizes modeling and control techniques from statistical machine learning and consists of three major components: the interference prediction model that infers application performance from resource consumption observed from different VMs, the interference-aware scheduler that is designed to utilize the model for effective resource management, and the task and resource monitor that collects application characteristics at the runtime for model adaption. We simulate TRACON with a wide variety of data-intensive applications including bioinformatics, data mining, video processing, email and web servers, etc. The evaluation results show that TRACON can achieve up to 50% improvement on application runtime, and up to 80% on I/O throughput for data-intensive applications in virtualized data centers.",https://ieeexplore.ieee.org/document/6114409,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 1341}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1365}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 49}]",53.0,,,,['scheduling'],,"SC '11: Proceedings of 2011 International Conference for High Performance Computing, Networking, Storage and Analysis",True,['resource-provisioning'],,True,,,,10.1145/2063384.2063447,Conference Paper,,
405,Probabilistic reasoning for streaming anomaly detection,"['K. M. Carter', ' W. W. Streilein']",2012,"In many applications it is necessary to determine whether an observation from an incoming high-volume data stream matches expectations or is anomalous. A common method for performing this task is to use an Exponentially Weighted Moving Average (EWMA), which smooths out the minor variations of the data stream. While EWMA is efficient at processing high-rate streams, it can be very volatile to abrupt transient changes in the data, losing utility for appropriately detecting anomalies. In this paper we present a probabilistic approach to EWMA which dynamically adapts the weighting based on the observation probability. This results in robustness to data anomalies yet quick adaptability to distributional data shifts.",https://ieeexplore.ieee.org/document/6319708,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 7}]",13.0,,,,['failure-detection'],,2012 IEEE Statistical Signal Processing Workshop (SSP),True,['failure-management'],,,,,,,,,
406,Toward decentralized probabilistic management,"['A. G. Prieto', ' D. Gillblad', ' R. Steinert', ' A. Miron']",2011,"In recent years, data communication networks have grown to immense size and have been diversified by the mobile revolution. Existing management solutions are based on a centralized deterministic paradigm, which is appropriate for networks of moderate size operating in relatively stable conditions. However, it is becoming increasingly apparent that these management solutions are not able to cope with the large dynamic networks that are emerging. In this article, we argue that the adoption of a decentralized and probabilistic paradigm for network management will be crucial to meet the challenges of future networks, such as efficient resource usage, scalability, robustness, and adaptability. We discuss the potential of decentralized probabilistic management and its impact on management operations, and illustrate the paradigm by three example solutions for real-time monitoring and anomaly detection.",https://ieeexplore.ieee.org/document/5936159,True,"[{'database': 'IEEE', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 502}]",17.0,,,,['failure-detection'],,IEEE Communications Magazine,True,['failure-management'],,,,,,,,,
407,A dynamic checkpointing scheme based on reinforcement learning,"['H. Okamura', ' Y. Nishimura', ' T. Dohi']",2004,"We develop a new checkpointing scheme for a uniprocess application. First, we model the checkpointing scheme by a semiMarkov decision process, and apply the reinforcement learning algorithm to estimate statistically the optimal checkpointing policy. More specifically, the representative reinforcement learning algorithm, called the Q-learning algorithm, is used to develop an adaptive checkpointing scheme. In simulation experiments, we examine the asymptotic behavior of the system overhead with adaptive checkpointing and show quantitatively that the proposed dynamic checkpoint algorithm is useful and robust under an incomplete knowledge on the failure time distribution.",https://ieeexplore.ieee.org/document/1276566,True,"[{'database': 'IEEE', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('remediation' OR 'recovery')"", 'index': 79}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 971}]",23.0,['reinforcement-learning'],['configuration'],['novel-use'],['failure-prevention'],"['node', 'cluster', 'software']","10th IEEE Pacific Rim International Symposium on Dependable Computing, 2004. Proceedings.",True,['failure-management'],True,,['checkpointing'],True,22.0,,,,
408,Mining Health Models for Performance Monitoring of Services,"['M. Acharya', ' V. Kommineni']",2009,"Online services such as search and live applications rely on large infrastructures in data centers, consisting of both stateless servers (e.g., web servers) and stateful servers (e.g., database servers). Acceptable performance of such infrastructures, and hence the availability of online services, rely on a very large number of parameters such as per-process resources and configurable system/application parameters. These parameters are available for collection as performance counters distributed across various machines, but services have had a hard time determining which performance counters to monitor and what thresholds to use for performance alarms in a production environment. In this paper, we present a novel framework called PerfAnalyzer, a storage-efficient and pro-active performance monitoring framework for correlating service health with performance counters. PerfAnalyzer automatically infers and builds health models for any service by running the standard suite of predeployment tests for the service and data mining the resulting performance counter data-set. A filtered set of performance counters and thresholds of alarms are produced by our framework. The health model inferred by our framework can then be used to detect performance degradation and collect detailed data for root-cause analysis in a production environment. We have applied PerfAnalyzer on five simple stress scenarios - CPU, memory, I/O, disk, and network, and two real system - Microsoft's SQL Server 2005 and IIS 7.0 Web Server, with promising results.",https://ieeexplore.ieee.org/document/5431755,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 32}, {'database': 'ACM', 'search_string': ""'regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 21}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 17}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 48}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 16}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 81}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 55}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 4}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 29}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault localization' OR 'failure localization')"", 'index': 25}]",1.0,,[],[],['failure-detection'],,2009 IEEE/ACM International Conference on Automated Software Engineering,True,['failure-management'],,,['anomaly-detection'],,,10.1109/ASE.2009.95,Conference Paper,"['data mining', 'performance monitoring', 'performance counters', 'machine learning', 'service health model']",
409,iFeedback: Exploiting User Feedback for Real-Time Issue Detection in Large-Scale Online Service Systems,"['W. Zheng', ' H. Lu', ' Y. Zhou', ' J. Liang', ' H. Zheng', ' Y. Deng']",2019,"Large-scale online systems are complex, fast-evolving, and hardly bug-free despite the testing efforts. Backend system monitoring cannot detect many types of issues, such as UI related bugs, bugs with small impact on backend system indicators, or errors from third-party co-operating systems, etc. However, users are good informers of such issues: They will provide their feedback for any types of issues. This experience paper discusses our design of iFeedback, a tool to perform real-time issue detection based on user feedback texts. Unlike traditional approaches that analyze user feedback with computation-intensive natural language processing algorithms, iFeedback is focusing on fast issue detection, which can serve as a system life-condition monitor. In particular, iFeedback extracts word combination-based indicators from feedback texts. This allows iFeedback to perform fast system anomaly detection with sophisticated machine learning algorithms. iFeedback then further summarizes the texts with an aim to effectively present the anomaly to the developers for root cause analysis. We present our representative experiences in successfully applying iFeedback in tens of large-scale production online service systems in ten months.",https://ieeexplore.ieee.org/document/8952229,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 42}, {'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1140}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 48}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 1523}, {'database': 'ACM', 'search_string': ""'regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 22}, {'database': 'ACM', 'search_string': ""'AIOps'"", 'index': 3}, {'database': 'ACM', 'search_string': ""'regression' AND ('IT operations')"", 'index': 7}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 27}, {'database': 'ACM', 'search_string': ""'classification' AND ('IT operations')"", 'index': 9}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 12}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('IT operations')"", 'index': 8}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('IT operations')"", 'index': 9}, {'database': 'ACM', 'search_string': ""'clustering' AND ('IT operations')"", 'index': 8}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 12}, {'database': 'ACM', 'search_string': ""'classification' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 34}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('IT operations')"", 'index': 88}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 56}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('IT operations')"", 'index': 40}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 30}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('IT operations')"", 'index': 10}]",0.0,"['clustering', 'decision-tree']",,,['failure-detection'],,2019 34th IEEE/ACM International Conference on Automated Software Engineering (ASE),True,['failure-management'],,True,['anomaly-detection'],,,10.1109/ASE.2019.00041,Conference Paper,,
410,PRACTISE: Robust prediction of data center time series,"['J. Xue', ' F. Yan', ' R. Birke', ' L. Y. Chen', ' T. Scherer', ' E. Smirni']",2015,"We analyze workload traces from production data centers and focus on their VM usage patterns of CPU, memory, disk, and network bandwidth. Burstiness is a clear characteristic of many of these time series: there exist peak loads within clear periodic patterns but also within patterns that do not have clear periodicity. We present PRACTISE, a neural network based framework that can efficiently and accurately predict future loads, peak loads, and their timing. Extensive experimentation using traces from IBM data centers illustrates PRACTISE's superiority when compared to ARIMA and baseline neural network models, with average prediction errors that are significantly smaller. Its robustness is also illustrated with respect to the prediction window that can be short-term (i.e., hours) or long-term (i.e., a week).",https://ieeexplore.ieee.org/document/7367348,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 12}, {'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 371}]",28.0,,,,['workload-prediction'],,2015 11th International Conference on Network and Service Management (CNSM),True,['resource-provisioning'],,,,,,,,,
411,Continuous Incident Triage for Large-Scale Online Service Systems,"['J. Chen', ' X. He', ' Q. Lin', ' H. Zhang', ' D. Hao', ' F. Gao', ' Z. Xu', ' Y. Dang', ' D. Zhang']",2019,"In recent years, online service systems have become increasingly popular. Incidents of these systems could cause significant economic loss and customer dissatisfaction. Incident triage, which is the process of assigning a new incident to the responsible team, is vitally important for quick recovery of the affected service. Our industry experience shows that in practice, incident triage is not conducted only once in the beginning, but is a continuous process, in which engineers from different teams have to discuss intensively among themselves about an incident, and continuously refine the incident-triage result until the correct assignment is reached. In particular, our empirical study on 8 real online service systems shows that the percentage of incidents that were reassigned ranges from 5.43% to 68.26% and the number of discussion items before achieving the correct assignment is up to 11.32 on average. To improve the existing incident triage process, in this paper, we propose DeepCT, a Deep learning based approach to automated Continuous incident Triage. DeepCT incorporates a novel GRU-based (Gated Recurrent Unit) model with an attention-based mask strategy and a revised loss function, which can incrementally learn knowledge from discussions and update incident-triage results. Using DeepCT, the correct incident assignment can be achieved with fewer discussions. We conducted an extensive evaluation of DeepCT on 14 large-scale online service systems in Microsoft. The results show that DeepCT is able to achieve more accurate and efficient incident triage, e.g., the average accuracy identifying the responsible team precisely is 0.641~0.729 with the number of discussion items increasing from 1 to 5. Also, DeepCT statistically significantly outperforms the state-of-the-art bug triage approach.",https://ieeexplore.ieee.org/document/8952483,True,"[{'database': 'IEEE', 'search_string': ""('DL' OR 'deep learning') AND ('remediation' OR 'recovery')"", 'index': 116}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 881}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('remediation' OR 'recovery')"", 'index': 56}]",1.0,['rnn'],,['novel-use'],['remediation'],,2019 34th IEEE/ACM International Conference on Automated Software Engineering (ASE),True,['failure-management'],,,['triage'],,,10.1109/ASE.2019.00042,Conference Paper,"['online service systems', 'incident triage', 'deep learning']",
412,Anomaly detection in network traffic using extreme learning machine,"['Y. Imamverdiyev', ' L. Sukhostat']",2016,"Intrusion detection systems are one of the most relevant security features against network attacks. Machine learning methods are used to analyze network traffic parameters for the presence of an attack signs. In this paper, extreme learning machine method is considered for intrusion detection in network traffic. The experimental results lead to the conclusion of practical significance of the proposed approach for attacks detection in network traffic.",https://ieeexplore.ieee.org/document/7991732,True,"[{'database': 'IEEE', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 220}, {'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 200}]",4.0,,,,['failure-detection'],,2016 IEEE 10th International Conference on Application of Information and Communication Technologies (AICT),True,['failure-management'],,,,,,,,,
413,Guided Problem Diagnosis through Active Learning,"['S. Duan', ' S. Babu']",2008,"There is widespread interest today in developing tools that can diagnose the cause of a system failure accurately and efficiently based on monitoring data collected from the system. Over time, the system monitoring data will contain two types of failure data: (i) annotated failure data L, which is monitoring data collected from failure states of the system, where the cause of failure has been diagnosed and attached as annotations with the data; and (ii) unannotated failure data U. Previous work on wholly- or partially-automated diagnosis focused on L or U in isolation. In this paper, we argue that it is important to consider both L and U together to improve the overall accuracy of diagnosis; and in particular, to proactively move instances from U to L. However, such movement requires manual diagnosis effort from system administrators. Since manual diagnosis is expensive and time-consuming, we propose an algorithm to make the best use of manual effort while maximizing the benefit gained from newly diagnosed instances. We report an experimental evaluation of our algorithm using data from a variety of failures - both single failures and multiple correlated failures - injected in a testbed, as well as with synthetic data.",https://ieeexplore.ieee.org/document/4550826,True,"[{'database': 'IEEE', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 962}]",19.0,,,,['root-cause-analysis'],,2008 International Conference on Autonomic Computing,True,['failure-management'],,,,,,,,,
414,Localizing transient faults using dynamic bayesian networks,"['S. Jha', ' W. Li', ' S. A. Seshia']",2009,"Transient faults are a major concern in today's deep sub-micron semiconductor technology. These faults are rare but they have been known to cause catastrophic system-level failures. Transient errors often occur due to physical effects on deployed systems and hence, diagnosis of transient errors must be performed over manufactured chips or systems assembled from black-box components where arbitrary instrumentation of the system is not possible and hence, the system state is only partially observable. Further, these systems are often composed of components that are third party IP which further adds opaqueness to the system. In this paper, we propose a probabilistic approach to localize transient faults in space and time for such partially observable systems. From a set of correct traces and a failure trace, we seek to locate the faulty component and the cycle of operation at which the fault occurred. Our technique uses correct system traces over monitored components of the system to learn a dynamic Bayesian network (DBN) summarizing the temporal dependencies across the monitored components. This DBN is augmented with different error hypotheses allowed by the fault model. The most probable explanation (MPE) among these hypotheses corresponds to the most likely location of the error. We evaluated the effectiveness of our technique on a set of ISCAS89 benchmarks and a router design used in on-chip networks in a multi-core design.",https://ieeexplore.ieee.org/document/5340170,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 367}]",7.0,,,,['root-cause-analysis'],,2009 IEEE International High Level Design Validation and Test Workshop,True,['failure-management'],,True,,,,,,,
415,Improving accuracy of host load predictions on computational grids by artificial neural networks,"['Truong Vinh Truong Duy', ' Yukinori Sato', ' Y. Inoguchi']",2009,"The capability to predict the host load of a system is significant for computational grids to make efficient use of shared resources. This paper attempts to improve the accuracy of host load predictions by applying a neural network predictor to reach the goal of best performance and load balance. We describe feasibility of the proposed predictor in a dynamic environment, and perform experimental evaluation using collected load traces. The results show that the neural network achieves a consistent performance improvement with surprisingly low overhead. Compared with the best previously proposed method, the typical 20:10:1 network reduces the mean and standard deviation of the prediction errors by approximately 60% and 70%, respectively. The training and testing time is extremely low, as this network needs only a couple of seconds to be trained with more than 100,000 samples in order to make tens of thousands of accurate predictions within just a second.",https://ieeexplore.ieee.org/document/5160878,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 373}]",2.0,,,,,,2009 IEEE International Symposium on Parallel & Distributed Processing,True,['resource-provisioning'],,,,,,,,,
416,Unsupervised Anomaly Detection for Intricate KPIs via Adversarial Training of VAE,"['W. Chen', ' H. Xu', ' Z. Li', ' D. Pei', ' J. Chen', ' H. Qiao', ' Y. Feng', ' Z. Wang']",2019,"To ensure the reliability of the Internet-based application services, KPIs (Key Performance Monitors) are closely monitored in real time and the anomalies presented in the KPIs must be discovered in time. While anomaly detection for the seasonal smooth service-level KPIs (e.g., number of transactions per minute) have been solved reasonably well in the literature, the intricate KPIs at the machine level (e.g., the number of I/O requests on a server monitored per second) has been little studied. These intricate KPIs are prevalent and important, but exhibit non-Gaussian noises and complex data distribution that are hard to model. In this paper, we propose an adversarial training method in the Bayesian network based on partition analysis with solid theoretical proof. Based on it, we propose the first unsupervised anomaly detection algorithmBuzz for intricate KPIs with high performance. Its best F-scores on the data from a global Internet company range from 0.92 to 0.99, significantly outperforming a state-of-art VAE-based unsupervised approach without adversarial training and a state-of-art supervised approach.",https://ieeexplore.ieee.org/document/8737430,True,"[{'database': 'IEEE', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 373}]",2.0,"['multilayer-perceptron', 'autoencoder']",['kpis'],,['failure-detection'],,IEEE INFOCOM 2019 - IEEE Conference on Computer Communications,True,['failure-management'],,,['anomaly-detection'],,,,,,
417,ExplainIt! -- A Declarative Root-cause Analysis Engine for Time Series Data,"['Vimalkumar Jeyakumar', 'Omid Madani', 'Ali Parandeh', 'Ashutosh Kulshreshtha', 'Weifei Zeng', 'Navindra Yadav']",2019,"We present \sys, a declarative, unsupervised root-cause analysis engine that uses time series monitoring data from large complex systems such as data centres. \sys empowers operators to succinctly specify a large number of causal hypotheses to search for causes of interesting events. \sys then ranks these hypotheses, reducing the number of causal dependencies from hundreds of thousands to a handful for human understanding. We show how a declarative language, such as SQL, can be effective in declaratively enumerating hypotheses that probe the structure of an unknown probabilistic graphical causal model of the underlying system. Our thesis is that databases are in a unique position to enable users to rapidly explore the possible causal mechanisms in data collected from diverse sources. We empirically demonstrate how \sys had helped us resolve over 30~performance issues in a commercial product since late 2014, of which we discuss a few cases in detail.",https://doi.org/10.1145/3299869.3314048,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 2}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 3}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 5}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 2}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 8}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 5}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 3}]",1.0,['bayesian-network'],,,['root-cause-analysis'],['database'],Proceedings of the 2019 International Conference on Management of Data,True,['failure-management'],,,['root-cause-diagnosis'],,,10.1145/3299869.3314048,Conference Paper,"['causal inference', 'causal models', 'diagnosis', 'root-cause analysis', 'multivariate correlation analysis', 'sql', 'time-series', 'big-data']",
418,Performance Monitoring and Root Cause Analysis for Cloud-hosted Web Applications,"['Hiranya Jayathilaka', 'Chandra Krintz', 'Rich Wolski']",2017,"In this paper, we describe Roots - a system for automatically identifying the ""root cause"" of performance anomalies in web applications deployed in Platform-as-a-Service (PaaS) clouds. Roots does not require application-level instrumentation. Instead, it tracks events within the PaaS cloud that are triggered by application requests using a combination of metadata injection and platform-level instrumentation. We describe the extensible architecture of Roots, a prototype implementation of the system, and a statistical methodology for performance anomaly detection and diagnosis. We evaluate the efficacy of Roots using a set of PaaS-hosted web applications, and detail the performance overhead and scalability of the implementation.",https://doi.org/10.1145/3038912.3052649,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 4}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 11}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 8}]",14.0,,,,['root-cause-analysis'],,Proceedings of the 26th International Conference on World Wide Web,True,['failure-management'],,,,,,10.1145/3038912.3052649,Conference Paper,"['platform-as-a-service', 'web services', 'root cause analysis', 'application performance monitoring', 'cloud computing']",
419,Automated root cause isolation of performance regressions during software development,"['Christoph Heger', 'Jens Happe', 'Roozbeh Farahbod']",2013,"Performance is crucial for the success of an application. To build responsive and cost efficient applications, software engineers must be able to detect and fix performance problems early in the development process. Existing approaches are either relying on a high level of abstraction such that critical problems cannot be detected or require high manual effort. In this paper, we present a novel approach that integrates performance regression root cause analysis into the existing development infrastructure using performance-aware unit tests and the revision history. Our approach is easy to use and provides software engineers immediate insights with automated root cause analysis. In a realistic case study based on the change history of Apache Commons Math, we demonstrate that our approach can automatically detect and identify the root cause of a major performance regression.",https://doi.org/10.1145/2479871.2479879,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 7}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 11}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 10}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 6}, {'database': 'ACM', 'search_string': ""'classification' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 8}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 25}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 13}]",26.0,,"['source-code', 'tests']",,"['failure-prevention', 'root-cause-analysis']",['software'],Proceedings of the 4th ACM/SPEC International Conference on Performance Engineering,True,['failure-management'],,,"['software-defect-prediction', 'fault-localization']",,,10.1145/2479871.2479879,Conference Paper,"['root cause analysis', 'performance regression']",
420,Griffon: Reasoning about Job Anomalies with Unlabeled Data in Cloud-based Platforms,"['%A Liqun Shao', 'Yiwen Zhu', 'Siqi Liu', 'Abhiram Eswaran', 'Kristin Lieber', 'Janhavi Mahajan', 'Minsoo Thigpen', 'Sudhir Darbha', 'Subru Krishnan', 'Soundar Srinivasan et al.']",2019,"Microsoft's internal big data analytics platform is comprised of hundreds of thousands of machines, serving over half a million jobs daily, from thousands of users. The majority of these jobs are recurring and are crucial for the company's operation. Although administrators spend significant effort tuning system performance, some jobs inevitably experience slowdowns, i.e., their execution time degrades over previous runs. Currently, the investigation of such slowdowns is a labor-intensive and error-prone process, which costs Microsoft significant human and machine resources, and negatively impacts several lines of businesses. In this work, we present Griffin, a system we built and have deployed in production last year to automatically discover the root cause of job slowdowns. Existing solutions either rely on labeled data (i.e., resolved incidents with labeled reasons for job slowdowns), which is in most cases non-existent or non-trivial to acquire, or on time-series analysis of individual metrics that do not target specific jobs holistically. In contrast, in Griffin we cast the problem to a corresponding regression one that predicts the runtime of a job, and show how the relative contributions of the features used to train our interpretable model can be exploited to rank the potential causes of job slowdowns. Evaluated over historical incidents, we show that Griffin discovers slowdown causes that are consistent with the ones validated by domain-expert engineers, in a fraction of the time required by them.",https://doi.org/10.1145/3357223.3362716,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 12}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 15}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 21}, {'database': 'ACM', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 100}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 9}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 62}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 22}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 5}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 4}, {'database': 'ACM', 'search_string': ""'classification' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 19}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 52}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 23}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 3}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 10}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 14}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 97}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1510}, {'database': 'arxiv', 'search_string': ""'regression' AND ('cloud')"", 'index': 124}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 622}, {'database': 'arxiv', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 1528}, {'database': 'arxiv', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1287}]",0.0,"['multilayer-perceptron', 'random-forest', 'linear-regression', 'regression-tree']",,['new-method'],['root-cause-analysis'],,Proceedings of the ACM Symposium on Cloud Computing,True,['failure-management'],,,"['root-cause-diagnosis', 'workload-prediction']",,,10.1145/3357223.3362716,Conference Paper,"['analytics job', 'unlabeled data', 'anomaly detection', 'big data analytics cluster', 'reasoning and diagnostics', 'job slowdown', 'Job anomaly', 'root cause analysis']",
421,Anomaly-based fault detection in pervasive computing system,"['Byoung Uk Kim', 'Youssif Al-Nashif', 'Samer Fayssal', 'Salim Hariri', 'Mazin Yousif']",2008,"The increased complexity of hardware and software resources and the asynchronous interaction among components (such as servers, end devices, network, services and software) make fault detection and recovery very challenging. In this paper, we present innovative concepts for fault detection, root cause analysis and self-healing architectures analyzing the duration of pattern transition sequences during an execution window. In this approach, all interactions among components of Pervasive Computing Systems (PCS) are monitored and analyzed. We use three-dimensional array of features to capture spatial and temporal variability to be used by an anomaly analysis engine to immediately generate an alert when abnormal behavior pattern is captured indicating some kind of software or hardware failure. The main contributions of this paper include the innovative analysis methodology and feature selection to detect and identify anomalous behavior. Evaluating the effectiveness of this approach to detect faults injected asynchronously shows a detection rate of above 99.9% with no occurrences of false alarms for a wide range of scenarios.",https://doi.org/10.1145/1387269.1387294,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 19}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault detection' OR 'failure detection')"", 'index': 67}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 33}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault detection' OR 'failure detection')"", 'index': 19}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 10}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault detection' OR 'failure detection')"", 'index': 3}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 9}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault detection' OR 'failure detection')"", 'index': 19}, {'database': 'ACM', 'search_string': ""'classification' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 26}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault detection' OR 'failure detection')"", 'index': 10}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 78}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 16}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 3}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 21}]",10.0,,,,['failure-detection'],,Proceedings of the 5th international conference on Pervasive services,True,['failure-management'],,True,,,,10.1145/1387269.1387294,Conference Paper,"['performance objectives', 'faults', 'abnormality detection', 'pattern profiling', 'interaction analysis']",
422,A comparative study of pairwise regression techniques for problem determination,"['Mohammad A. Munawar', 'Paul A. S. Ward']",2007,"Many runtime metrics can be collected from modern software systems. Stable statistical relationships exist among these metrics. Deviation from these stable relationships indicates potential problems, allowing diagnosis of failures. There exist many modeling techniques to represent these relationships. However, which one to use is a question that has yet to be studied. In this paper we compare the use of simple linear regression (SLR) to some of its more complex variants, including autoregressive regression and locally weighted regression. We consider the component coverage, model robustness, accuracy of diagnosis, and computation cost. Our study finds that while more flexible models can improve diagnosis accuracy, they achieve it at the cost of reduced robust-ness. In particular, we found the autoregressive regression model with exogenous input (ARX) to provide the most accurate diagnosis; however, it is the least robust of the techniques considered and the second most expensive. This study also finds that smoothing and other data transformations can noticeably improve results of SLR, thus providing an efficient alternative to ARX.",https://doi.org/10.1145/1321211.1321227,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 26}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault localization' OR 'failure localization')"", 'index': 97}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault localization' OR 'failure localization')"", 'index': 25}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 22}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 60}]",29.0,,,,['failure-detection'],,Proceedings of the 2007 conference of the center for advanced studies on Collaborative research,True,['failure-management'],,,,,,10.1145/1321211.1321227,Conference Paper,,
423,A survey of online failure prediction methods,"['Felix Salfner', 'Maren Lenk', 'Miroslaw Malek']",2010,"With the ever-growing complexity and dynamicity of computer systems, proactive fault management is an effective approach to enhancing availability. Online failure prediction is the key to such techniques. In contrast to classical reliability methods, online failure prediction is based on runtime monitoring and a variety of models and methods that use the current state of a system and, frequently, the past experience as well. This survey describes these methods. To capture the wide spectrum of approaches concerning this area, a taxonomy has been developed, whose different approaches are explained and major concepts are described in detail.",https://doi.org/10.1145/1670679.1670680,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 64}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 3}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 5}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 1}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 2}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 53}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 11}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('remediation' OR 'recovery')"", 'index': 91}, {'database': 'ACM', 'search_string': ""'classification' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 92}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prediction' OR 'failure prediction')"", 'index': 1}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 15}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 42}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 1}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 65}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 1}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 2}]",536.0,,,['survey'],['failure-prediction'],,ACM Computing Surveys (CSUR),True,['failure-management'],True,True,['system-failure-prediction'],,,10.1145/1670679.1670680,Journal Article,"['failure prediction', 'Error', 'fault', 'prediction metrics', 'runtime monitoring']",
424,CoFlux: robustly correlating KPIs by fluctuations for service troubleshooting,"['%A Ya Su', 'Youjian Zhao', 'Wentao Xia', 'Rong Liu', 'Jiahao Bu', 'Jing Zhu', 'Yuanpu Cao', 'Haibin Li', 'Chenhao Niu', 'Yiyin Zhang et al.']",2019,"Internet-based service companies monitor a large number of KPIs (Key Performance Indicators) to ensure their service quality and reliability. Correlating KPIs by fluctuations reveals interactions between KPIs under anomalous situations and can be extremely useful for service troubleshooting. However, such a KPI flux-correlation has been little studied so far in the domain of Internet service operations management. A major challenge is how to automatically and accurately separate fluctuations from normal variations in KPIs with different structural characteristics (such as seasonal, trend and stationary) for a large number of KPIs. In this paper, we propose CoFlux, an unsupervised approach, to automatically (without manual selection of algorithm fitting and parameter tuning) determine whether two KPIs are correlated by fluctuations, in what temporal order they fluctuate, and whether they fluctuate in the same direction. CoFlux's robust feature engineering and robust correlation score computation enable it to work well against the diverse KPI characteristics. Our extensive experiments have demonstrated that CoFlux achieves the best F1-Scores of 0.84 (0.90), 0.92 (0.95), 0.95 (0.99), in answering these three questions, in the two real datasets from a top global Internet company, respectively. Moreover, we showed that CoFlux is effective in assisting service troubleshooting through the applications of alert compression, recommending Top N causes, and constructing fluctuation propagation chains.",https://doi.org/10.1145/3326285.3329048,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 78}, {'database': 'ACM', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 2}, {'database': 'ACM', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 2}, {'database': 'ACM', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 51}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 73}]",1.0,['correlation'],['kpis'],['new-method'],['failure-detection'],,Proceedings of the International Symposium on Quality of Service,True,['failure-management'],,,['anomaly-detection'],,,10.1145/3326285.3329048,Conference Paper,"['service operation and management', 'key performance indicator', 'time series', 'service troubleshooting', 'fluctuation correlation']",
425,Predicting faults in high performance computing systems: an in-depth survey of the state-of-the-practice,"['David Jauk', 'Dai Yang', 'Martin Schulz']",2019,"As we near exascale, resilience remains a major technical hurdle. Any technique with the goal of achieving resilience suffers from having to be reactive, as failures can appear at any time. A wide body of research aims at predicting failures, i.e., forecasting failures so that evasive actions can be taken while the system is still fully functional, which has the benefit of giving insight into the global system state. This research area has grown very diverse with a large number of approaches, yet is currently poorly classified, making it hard to understand the impact and coverage of existing work. In this paper, we perform an extensive survey of existing literature in failure prediction by analyzing and comparing more than 30 different failure prediction approaches. We develop a taxonomy, which aids in categorizing the methods, and we show how this can help us to understand the state-of-the-practice of this field and to identify opportunities, gaps as well as future work.",https://doi.org/10.1145/3295500.3356185,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 92}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 52}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 6}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 83}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 74}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 43}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 38}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 68}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 17}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 83}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prediction' OR 'failure prediction')"", 'index': 19}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 18}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 16}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 61}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 21}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 87}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 25}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 44}]",2.0,,,['survey'],['failure-prediction'],['hpc'],"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis",True,['failure-management'],,,"['hardware-failure-prediction', 'system-failure-prediction']",,,10.1145/3295500.3356185,Conference Paper,"['resillience', 'exascale computing', 'fault prediction', 'high performance computing']",
426,Time-Series Anomaly Detection Service at Microsoft,"['Hansheng Ren', 'Bixiong Xu', 'Yujing Wang', 'Chao Yi', 'Congrui Huang', 'Xiaoyu Kou', 'Tony Xing', 'Mao Yang', 'Jie Tong', 'Qi Zhang']",2019,"Large companies need to monitor various metrics (for example, Page Views and Revenue) of their applications and services in real time. At Microsoft, we develop a time-series anomaly detection service which helps customers to monitor the time-series continuously and alert for potential incidents on time. In this paper, we introduce the pipeline and algorithm of our anomaly detection service, which is designed to be accurate, efficient and general. The pipeline consists of three major modules, including data ingestion, experimentation platform and online compute. To tackle the problem of time-series anomaly detection, we propose a novel algorithm based on Spectral Residual (SR) and Convolutional Neural Network (CNN). Our work is the first attempt to borrow the SR model from visual saliency detection domain to time-series anomaly detection. Moreover, we innovatively combine SR and CNN together to improve the performance of SR model. Our approach achieves superior experimental results compared with state-of-the-art baselines on both public datasets and Microsoft production data.",https://doi.org/10.1145/3292500.3330680,True,"[{'database': 'ACM', 'search_string': ""'AIOps'"", 'index': 5}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 20}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 16}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 46}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 48}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 42}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 43}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 55}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 92}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1373}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 29}]",2.0,['cnn'],['kpis'],['novel-use'],['failure-detection'],,Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining,True,['failure-management'],,,['anomaly-detection'],,,10.1145/3292500.3330680,Conference Paper,"['anomaly detection', 'spectral residual', 'time-series']",
427,Mining metrics to predict component failures,"['Nachiappan Nagappan', 'Thomas Ball', 'Andreas Zeller']",2006,"What is it that makes software fail? In an empirical study of the post-release defect history of five Microsoft software systems, we found that failure-prone software entities are statistically correlated with code complexity measures. However, there is no single set of complexity metrics that could act as a universally best defect predictor. Using principal component analysis on the code metrics, we built regression models that accurately predict the likelihood of post-release defects for new entities. The approach can easily be generalized to arbitrary projects; in particular, predictors obtained from one project can also be significant for new, similar projects.",https://doi.org/10.1145/1134285.1134349,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 25}]",838.0,"['dimensionality-reduction', 'logistic-regression']",['code-metrics'],"['discussion', 'novel-use']",['failure-prevention'],['source-code'],Proceedings of the 28th international conference on Software engineering,True,['failure-management'],True,,['software-defect-prediction'],True,1.0,10.1145/1134285.1134349,Conference Paper,"['empirical study', 'principal component analysis', 'bug database', 'complexity metrics', 'regression model']",
428,Extended Hoeffding Adaptive Tree based-Server Load Prediction in Cloud Computing environment,"['Hajer Toumi', 'Zaki Brahmi', 'Mohammed Mohsen Gammoudi']",2020,"Cloud Computing (CC) enables client-server relationship in order to release users from computational and storage responsibility. As multi-tenant environment, Cloud providers are dealing, in one hand, with multiple concurrent users each of which exhibits a different and variable behavior over time and in the other hand, with a performance interference due to the co-location of multiple virtual machines (VMs) in the same server. Therefore, a real time server load prediction is needed in order to ensure efficient resource provisioning. While classical data mining based techniques suffer from important evaluation time and are enable to react to changes as it arrives, stream mining techniques can provide a real time prediction and changes detection. Thus, in this paper we used a well known stream mining technique, Hoeffding Adaptive Tree (HAT), in order to provide real time server load prediction. The aim of our proposed technique is to detect and react on the fly to different kind of changes that can affect the server load. Therefore, we augmented HAT by ensemble drift detectors in order to produce more accurate prediction. In order to evaluate our proposed technique HAT-ADS, we first compared it with a well known load prediction technique based on Bayesian approach. Then we compared our solution with another HAT based techniques. Overall, The experimentation showed that HAT-ADS proved important flexibility to various types of changes providing high accuracy with quick evaluation time and small memory footprint.",https://doi.org/10.1145/3368474.3368475,True,"[{'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 26}, {'database': 'ACM', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 24}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 26}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 8}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 46}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 4}]",0.0,,,['new-method'],['workload-prediction'],,Proceedings of the International Conference on High Performance Computing in Asia-Pacific Region,True,['resource-provisioning'],,,,,,10.1145/3368474.3368475,Conference Paper,"['Concept Drift', 'Stream Mining', 'Cloud Computing', 'VM interference', 'Real-Time Prediction']",
429,Distributed reinforcement learning for power limited many-core system performance optimization,"['Zhuo Chen', 'Diana Marculescu']",2015,"As power density emerges as the main constraint for many-core systems, controlling power consumption under the Thermal Design Power (TDP) while maximizing the performance becomes increasingly critical. To dynamically save power, Dynamic Voltage Frequency Scaling (DVFS) techniques have proved to be effective and are widely available commercially. In this paper, we present an On-line Distributed Reinforcement Learning (OD-RL) based DVFS control algorithm for many-core system performance improvement under …",https://ieeexplore.ieee.org/document/7092630,True,"[{'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('remediation' OR 'recovery')"", 'index': 72}]",63.0,,,,['power-management'],,"Proceedings of the 2015 Design, Automation & Test in Europe Conference & Exhibition",True,['resource-provisioning'],,,,,,,Conference Paper,,
430,Machine Learning Methods for Reliable Resource Provisioning in Edge-Cloud Computing: A Survey,"['Thang Le Duc', 'Rafael García Leiva', 'Paolo Casari', 'Per-Olov Östberg']",2019,"Large-scale software systems are currently designed as distributed entities and deployed in cloud data centers. To overcome the limitations inherent to this type of deployment, applications are increasingly being supplemented with components instantiated closer to the edges of networks—a paradigm known as edge computing. The problem of how to efficiently orchestrate combined edge-cloud applications is, however, incompletely understood, and a wide range of techniques for resource and application management are currently in use.This article investigates the problem of reliable resource provisioning in joint edge-cloud environments, and surveys technologies, mechanisms, and methods that can be used to improve the reliability of distributed applications in diverse and heterogeneous network environments. Due to the complexity of the problem, special emphasis is placed on solutions to the characterization, management, and control of complex distributed applications using machine learning approaches. The survey is structured around a decomposition of the reliable resource provisioning problem into three categories of techniques: workload characterization and prediction, component placement and system consolidation, and application elasticity and remediation. Survey results are presented along with a problem-oriented discussion of the state-of-the-art. A summary of identified challenges and an outline of future research directions are presented to conclude the article.",https://doi.org/10.1145/3341145,True,"[{'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('remediation' OR 'recovery')"", 'index': 87}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 48}, {'database': 'ACM', 'search_string': ""'regression' AND ('remediation' OR 'recovery')"", 'index': 94}, {'database': 'ACM', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 33}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 22}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 83}, {'database': 'ACM', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 23}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault detection' OR 'failure detection')"", 'index': 75}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('cloud')"", 'index': 85}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('remediation' OR 'recovery')"", 'index': 48}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 54}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 2}, {'database': 'ACM', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 71}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 12}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 5}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 28}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 49}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('remediation' OR 'recovery')"", 'index': 47}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('remediation' OR 'recovery')"", 'index': 50}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 59}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('cloud computing')"", 'index': 5}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 57}, {'database': 'ACM', 'search_string': ""'logistic regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 91}]",3.0,,,['survey'],"['resource-consolidation', 'workload-prediction']",,ACM Computing Surveys (CSUR),True,['resource-provisioning'],,,,,,10.1145/3341145,Journal Article,"['consolidation', 'Reliability', 'remediation', 'machine learning', 'autoscaling', 'optimization', 'edge computing', 'distributed systems', 'placement', 'cloud computing']",
431,Predicting breakdowns in cloud services (with SPIKE),"['Jianfeng Chen', 'Joymallya Chakraborty', 'Philip Clark', 'Kevin Haverlock', 'Snehit Cherian', 'Tim Menzies']",2019,"Maintaining web-services is a mission-critical task where any down-time means loss of revenue and reputation (of being a reliable service provider). In the current competitive web services market, such a loss of reputation causes extensive loss of future revenue. To address this issue, we developed SPIKE, a data mining tool which can predict upcoming service breakdowns, half an hour into the future. Such predictions let an organization alert and assemble the tiger team to address the problem (e.g. by reconfiguring cloud hardware in order to reduce the likelihood of that breakdown). SPIKE utilizes (a) regression tree learning (with CART); (b) synthetic minority over-sampling (to handle how rare spikes are in our data); (c) hyperparameter optimization (to learn best settings for our local data) and (d) a technique we called ""topology sampling"" where training vectors are built from extensive details of an individual node plus summary details on all their neighbors. In the experiments reported here, SPIKE predicted service spikes 30 minutes into future with recalls and precision of 75% and above. Also, SPIKE performed relatively better than other widely-used learning methods (neural nets, random forests, logistic regression).",https://doi.org/10.1145/3338906.3340450,True,"[{'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 49}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 1}, {'database': 'ACM', 'search_string': ""'classification' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 31}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('cloud')"", 'index': 14}, {'database': 'ACM', 'search_string': ""'classification' AND ('cloud')"", 'index': 52}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 32}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 87}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 83}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 60}, {'database': 'ACM', 'search_string': ""'regression' AND ('cloud')"", 'index': 8}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 31}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('cloud computing')"", 'index': 68}, {'database': 'ACM', 'search_string': ""'regression' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 21}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 68}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 18}, {'database': 'arxiv', 'search_string': ""'logistic regression' AND ('cloud')"", 'index': 10}, {'database': 'arxiv', 'search_string': ""'regression' AND ('cloud')"", 'index': 47}, {'database': 'arxiv', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 33}]",0.0,['regression-tree'],,"['comparison', 'novel-use']",['failure-prediction'],,Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering,True,['failure-management'],,,['system-failure-prediction'],,,10.1145/3338906.3340450,Conference Paper,"['parameter tuning', 'Cloud', 'data mining', 'optimization']",
432,A comparison of machine learning algorithms for proactive hard disk drive failure detection,"['Teerat Pitakrat', 'André van Hoorn', 'Lars Grunske']",2013,"Failures or unexpected events are inevitable in critical and complex systems. Proactive failure detection is an approach that aims to detect such events in advance so that preventative or recovery measures can be planned, thus improving system availability. Machine learning techniques have been successfully applied to learn patterns from available datasets and to classify or predict to which class a new instance of data belongs. In this paper, we evaluate and compare the performance of 21 machine learning algorithms by using them for proactive hard disk drive failure detection. For this comparison, we use WEKA as an experimentation platform and benchmark publicly available datasets of hard disk drives that are used to predict imminent failures before the actual failures occur. The results show that different algorithms are suitable for different applications based on the desired prediction quality and the tolerated training and prediction time.",https://doi.org/10.1145/2465470.2465473,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 45}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault detection' OR 'failure detection')"", 'index': 47}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 3}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 45}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prevention' OR 'failure prevention')"", 'index': 38}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 33}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault detection' OR 'failure detection')"", 'index': 12}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 6}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault detection' OR 'failure detection')"", 'index': 2}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prevention' OR 'failure prevention')"", 'index': 2}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault detection' OR 'failure detection')"", 'index': 2}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prevention' OR 'failure prevention')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 5}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prevention' OR 'failure prevention')"", 'index': 20}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 73}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 7}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault detection' OR 'failure detection')"", 'index': 14}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('remediation' OR 'recovery')"", 'index': 59}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('remediation' OR 'recovery')"", 'index': 58}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault prevention' OR 'failure prevention')"", 'index': 3}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault detection' OR 'failure detection')"", 'index': 3}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 3}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prevention' OR 'failure prevention')"", 'index': 9}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prevention' OR 'failure prevention')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 71}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prevention' OR 'failure prevention')"", 'index': 11}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prevention' OR 'failure prevention')"", 'index': 3}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault detection' OR 'failure detection')"", 'index': 23}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prevention' OR 'failure prevention')"", 'index': 62}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prevention' OR 'failure prevention')"", 'index': 12}]",11.0,,,,['failure-prediction'],,Proceedings of the 4th international ACM Sigsoft symposium on Architecting critical systems,True,['failure-management'],,True,,,,10.1145/2465470.2465473,Conference Paper,"['machine learning', 'hard disk drive failures', 'proactive failure detection']",
433,Monitorless: Predicting Performance Degradation in Cloud Applications with Machine Learning,"['Johannes Grohmann', 'Patrick K. Nicholson', 'Jesus Omana Iglesias', 'Samuel Kounev', 'Diego Lugones']",2019,"Today, software operation engineers rely on application key performance indicators (KPIs) for sizing and orchestrating cloud resources dynamically. KPIs are monitored to assess the achievable performance and to configure various cloud-specific parameters such as flavors of instances and autoscaling rules, among others. Usually, keeping KPIs within acceptable levels requires application expertise which is expensive and can slow down the continuous delivery of software. Expertise is required because KPIs are normally based on application-specific quality-of-service metrics, like service response time and processing rate, instead of generic platform metrics, like those typical across various environments (e.g., CPU and memory utilization, I/O rate, etc.) In this paper, we investigate the feasibility of outsourcing the management of application performance from developers to cloud operators. In the same way that the serverless paradigm allows the execution environment to be fully managed by a third party, we discuss a monitorless model to streamline application deployment by delegating performance management. We show that training a machine learning model with platform-level data, collected from the execution of representative containerized services, allows inferring application KPI degradation. This is an opportunity to simplify operations as engineers can rely solely on platform metrics -- while still fulfilling application KPIs -- to configure portable and application agnostic rules and other cloud-specific parameters to automatically trigger actions such as autoscaling, instance migration, network slicing, etc. Results show that monitorless infers KPI degradation with an accuracy of 97% and, notably, it performs similarly to typical autoscaling solutions, even when autoscaling rules are optimally tuned with knowledge of the expected workload.",https://doi.org/10.1145/3361525.3361543,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 73}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 3}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 95}, {'database': 'ACM', 'search_string': ""'classification' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 22}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('cloud')"", 'index': 82}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 20}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 4}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 58}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 11}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('cloud computing')"", 'index': 30}, {'database': 'ACM', 'search_string': ""'regression' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 14}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 13}]",1.0,"['random-forest', 'logistic-regression', 'multilayer-perceptron', 'support-vector-machine', 'decision-tree']","['kpis', 'host-metrics']","['comparison', 'novel-use']",['workload-prediction'],,Proceedings of the 20th International Middleware Conference,True,['resource-provisioning'],,,,,,10.1145/3361525.3361543,Conference Paper,"['DevOps', 'Machine learning', 'Monitoring', 'Cloud computing']",
434,Discriminative Sequential Pattern Mining for Software Failure Detection,"['Hao Du', 'Yongchi Su', 'Chunping Li']",2016,"Software event sequence is a software behavior trace which is produced when software is running. Analyzing the database of software event sequences, we present a novel method to distinguish normal and abnormal behaviors for the purpose of software failure detection. Sequence classification has been a challenge task since sequences have the high-order temporal characteristics and make the number of patterns extremely massive. We select the frequent closed unique iterative patterns as candidate features, mine out the discriminative binary and numerical patterns for sequence classification, and further give an insight into the discriminative power improvement by feature combinations. The experimental results on synthetic and real-life datasets show the validity of our method.",https://doi.org/10.1145/2908446.2908453,True,"[{'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault detection' OR 'failure detection')"", 'index': 11}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault detection' OR 'failure detection')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault detection' OR 'failure detection')"", 'index': 5}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 3}]",0.0,,,,['failure-detection'],,Proceedings of the 10th International Conference on Informatics and Systems,True,['failure-management'],,,['anomaly-detection'],,,10.1145/2908446.2908453,Conference Paper,"['Software Failure Detection', 'Sequential Pattern Mining', 'Data Classification']",
435,Internet service performance failure detection,"['Amy Ward', 'Peter Glynn', 'Kathy Richardson']",1998,"The increasing complexity of computer networks and our increasing dependence on them means enforcing reliability requirements is both more challenging and more critical. The expansion of network services to include both traditional interconnect services and user-oriented services such as the web and email has guaranteed both the increased complexity of networks and the increased importance of their performance. The first step toward increasing reliability is early detection of network performance failures. Here we consider the applicability of statistical model frameworks under the most general assumptions possible. Using measurements from corporate proxy servers, we test the framework against real world failures. The results of these experiments show we can detect failures, but with some tradeoff questions. The pull is in the warning time: either we miss early warning signs or we report some false warnings. Finally, we offer insight into the problem of failure diagnosis.",https://doi.org/10.1145/306225.306237,True,"[{'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault detection' OR 'failure detection')"", 'index': 13}]",29.0,,,,['failure-detection'],,ACM SIGMETRICS Performance Evaluation Review,True,['failure-management'],,,,,,10.1145/306225.306237,Journal Article,,
436,Classification of software behaviors for failure detection: a discriminative pattern mining approach,"['David Lo', 'Hong Cheng', 'Jiawei Han', 'Siau-Cheng Khoo', 'Chengnian Sun']",2009,"Software is a ubiquitous component of our daily life. We often depend on the correct working of software systems. Due to the difficulty and complexity of software systems, bugs and anomalies are prevalent. Bugs have caused billions of dollars loss, in addition to privacy and security threats. In this work, we address software reliability issues by proposing a novel method to classify software behaviors based on past history or runs. With the technique, it is possible to generalize past known errors and mistakes to capture failures and anomalies. Our technique first mines a set of discriminative features capturing repetitive series of events from program execution traces. It then performs feature selection to select the best features for classification. These features are then used to train a classifier to detect failures. Experiments and case studies on traces of several benchmark software systems and a real-life concurrency bug from MySQL server show the utility of the technique in capturing failures and anomalies. On average, our pattern-based classification technique outperforms the baseline approach by 24.68% in accuracy.",https://doi.org/10.1145/1557019.1557083,True,"[{'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault detection' OR 'failure detection')"", 'index': 18}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault detection' OR 'failure detection')"", 'index': 2}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault localization' OR 'failure localization')"", 'index': 62}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault localization' OR 'failure localization')"", 'index': 71}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 2}]",205.0,"['support-vector-machine', 'pattern-matching']",['traces'],['new-method'],['failure-detection'],['software'],Proceedings of the 15th ACM SIGKDD international conference on Knowledge discovery and data mining,True,['failure-management'],True,,['anomaly-detection'],True,57.0,10.1145/1557019.1557083,Conference Paper,"['software behaviors', 'pattern-based classification', 'iterative patterns', 'sequential database', 'failure detection', 'closed unique patterns']",
437,Failure detection and localization in component based systems by online tracking,"['Haifeng Chen', 'Guofei Jiang', 'Cristian Ungureanu', 'Kenji Yoshihira']",2005,"The increasing complexity of today's systems makes fast and accurate failure detection essential for their use in mission-critical applications. Various monitoring methods provide a large amount of data about system's behavior. Analyzing this data with advanced statistical methods holds the promise of not only detecting the errors faster, but also detecting errors which are difficult to catch with current monitoring tools. Two challenges to building such detection tools are: the high dimensionality of observation data, which makes the models expensive to apply, and frequent system changes, which make the models expensive to update. In this paper, we present algorithms to reduce the dimensionality of data in a way that makes it easy to adapt to system changes. We decompose the observation data into signal and noise subspaces. Two statistics, the Hotelling T2 score and squared prediction error (SPE) are calculated to represent the data characteristics in signal and noise subspaces respectively. Instead of tracking the original data, we use a sequentially discounting expectation maximization (SDEM) algorithm to learn the distribution of the two extracted statistics. A failure event can then be detected based on the abnormal change of the distribution. Applying our technique to component interaction data in a simple e-commerce application shows better accuracy than building independent profiles for each component. Additionally, experiments on synthetic data show that the detection accuracy is high even for changing systems.",https://doi.org/10.1145/1081870.1081968,True,"[{'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault detection' OR 'failure detection')"", 'index': 29}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault localization' OR 'failure localization')"", 'index': 21}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 3}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault localization' OR 'failure localization')"", 'index': 63}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 8}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 2}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 11}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault detection' OR 'failure detection')"", 'index': 17}]",17.0,,,,['failure-detection'],,Proceedings of the eleventh ACM SIGKDD international conference on Knowledge discovery in data mining,True,['failure-management'],,True,,,,10.1145/1081870.1081968,Conference Paper,"['distributed computing', 'subspace decomposition', 'online tracking', 'failure detection', 'statistics', 'internet services']",
438,Autonomous Fault Detection in Self-Healing Systems: Comparing Hidden Markov Models and Artificial Neural Networks,"['Chris Schneider', 'Adam Barker', 'Simon Dobson']",2014,"Autonomously detecting and recovering from faults is one approach for reducing the operational complexity and costs associated with managing computing environments. We present a novel methodology for autonomously generating investigation leads that help identify systems faults. Specifically, when historical feature data is present, Hidden Markov Models can be used to heuristically identify the root cause of a fault in an unsupervised manner. This approach improves the state of the art by allowing self-healing systems to detect faults with greater autonomy than existing methodologies, and thus further reduce operational costs.",https://doi.org/10.1145/2553062.2553065,True,"[{'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault detection' OR 'failure detection')"", 'index': 85}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault detection' OR 'failure detection')"", 'index': 22}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 11}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('remediation' OR 'recovery')"", 'index': 20}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 9}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 84}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 14}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 4}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 19}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 83}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 25}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault detection' OR 'failure detection')"", 'index': 40}]",5.0,,,,['failure-detection'],,Proceedings of International Workshop on Adaptive Self-tuning Computing Systems,True,['failure-management'],,True,,,,10.1145/2553062.2553065,Conference Paper,"['fault detection', 'machine learning', 'autonomic computing', 'artificial neural networks', 'hidden markov models', 'self-healing systems']",
439,Infrastructure fault detection and prediction in edge cloud environments,"['Mbarka Soualhia', 'Chunyan Fu', 'Foutse Khomh']",2019,"As an emerging 5G system component, edge cloud becomes one of the key enablers to provide services such us mission critical, IoT and content delivery applications. However, because of limited fail-over mechanisms in edge clouds, faults (e.g., CPU or HDD faults) are highly undesirable. When infrastructure faults occur in edge clouds, they can accumulate and propagate; leading to severe degradation of system and application performance. It is therefore crucial to identify these faults early on and mitigate them. In this paper, we propose a framework to detect and predict several faults at infrastructure-level of edge clouds using supervised machine learning and statistical techniques. The proposed framework is composed of three main components responsible for: (1) data pre-processing, (2) fault detection, and (3) fault prediction. The results show that the framework allows to timely detect and predict several faults online. For instance, using Support Vector Machine (SVM), Random Forest(RF) and Neural Network(NN)models, the framework is able to detect non-fatal CPU and HDD overload faults with an F1 score of more than 95%. For the prediction, the Convolutional Neural Network (CNN) and Long Short Term Memory (LSTM) have comparable accuracy at 96.47% vs. 96.88% for CPU-overload fault and 85.52% vs. 88.73% for network fault.",https://doi.org/10.1145/3318216.3363305,True,"[{'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault detection' OR 'failure detection')"", 'index': 90}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 97}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault detection' OR 'failure detection')"", 'index': 27}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 16}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault detection' OR 'failure detection')"", 'index': 5}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('cloud')"", 'index': 81}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault detection' OR 'failure detection')"", 'index': 3}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 12}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 68}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 46}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 66}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault detection' OR 'failure detection')"", 'index': 29}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 23}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('remediation' OR 'recovery')"", 'index': 33}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 28}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault detection' OR 'failure detection')"", 'index': 15}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 45}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 30}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 46}, {'database': 'ACM', 'search_string': ""'logistic regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 85}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 69}]",0.0,"['random-forest', 'multilayer-perceptron', 'support-vector-machine', 'cnn', 'markov-model', 'rnn', 'bayesian-network', 'autoregression']","['software-metrics', 'host-metrics']","['comparison', 'novel-use']","['failure-prediction', 'failure-detection']","['vm', 'cpu', 'hard-drive']",Proceedings of the 4th ACM/IEEE Symposium on Edge Computing,True,['failure-management'],,True,['hardware-failure-prediction'],,,10.1145/3318216.3363305,Conference Paper,,
440,Hard disk Drive Failure Prediction Challenges in Machine Learning for Multi-variate Time Series,['Jie Yu'],2019,"Hard disk drive failure prediction (HDDFP) is an active area of machine learning applications. While recent work shows very promising results with high failure recall (95%) and precision based on SMART attributes, challenges remain that call for improvement in the machine learning pipeline. This paper starts with an introduction of the topic and a summary of recent work. Some challenges applicable to the existing solutions are then illustrated with an example using Backblaze dataset and its HDDFP rule. A main result of the paper is a rigorous formulation of the HDDFP problem as a MIMO dynamic system problem to tackle the challenges. It is also shown that the general formulation can help the existing classification method by enhancing the prediction lead time requirement. Though presented in the context of the HDDFP problem, the findings and thought process are applicable to other dynamic system failure prediction, and in some degree to the IoT and time series based analytics in general.",https://doi.org/10.1145/3373419.3373437,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 1}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 2}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('IT operations')"", 'index': 12}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 74}, {'database': 'ACM', 'search_string': ""'classification' AND ('IT operations')"", 'index': 5}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 52}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 15}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 97}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 8}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('IT operations')"", 'index': 2}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('IT operations')"", 'index': 2}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 9}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 46}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('IT operations')"", 'index': 61}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 95}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('IT operations')"", 'index': 3}]",0.0,,['host-metrics'],,['failure-prediction'],['hard-drive'],Proceedings of the 2019 3rd International Conference on Advances in Image Processing,True,['failure-management'],,,['hardware-failure-prediction'],,,10.1145/3373419.3373437,Conference Paper,"['dynamic system', 'hard disk', 'machine learning', 'IoT', 'multi-variate', 'Failure prediction', 'SMART', 'time series', 'big data']",
441,Memory failure prediction using online learning,"['Xiaoming Du', 'Cong Li']",2018,"Occurring frequently in datacenters, dynamic random access memory (DRAM) errors are the leading cause of the failures among various hardware components. DRAM failure analysis is one of the most important topics in hardware reliability, availability, and serviceability. Though with comprehensive studies of DRAM failure modes in prior work, a mechanism of predicting future failures on DRAM components is not available today. In this paper we address the problem of predicting the failures on micro-level DRAM components including cells, rows, and columns. A DRAM failure is the combined effect of the wear level of a DRAM fault and the implicit runtime context. Correctly predicting DRAM failures quantifies the impact to DRAM reliability and enables advanced error-prevention mechanisms such as efficient page retirement or dynamic substitution with spare DRAM components. We propose an online learning method, repeatedly taking the historical memory failure data of an individual server as the input to predict its failure occurrences in the near future. The learning algorithm embeds a kernel function to evaluate how well the current error observation follows certain previous observations in its failure history. It then performs the prediction by discovering the implicit patterns online. The algorithm strengthens failure confidence with the history of repeated errors from hard faults while washing out occasional errors from soft faults. It also adapts to the unobservable time-varying context by penalizing the change in failure observations across cycles. We further augment the algorithm with a mechanism which propagates failure confidence scores to nearby cells including the ones without error history. Empirical evaluation demonstrates that the failure prediction approach consistently outperforms the baseline methods based on historical error statistics.",https://doi.org/10.1145/3240302.3240309,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 2}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 3}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 1}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 1}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 1}]",1.0,,,['new-method'],['failure-prediction'],['memory'],Proceedings of the International Symposium on Memory Systems,True,['failure-management'],,,['hardware-failure-prediction'],,,10.1145/3240302.3240309,Conference Paper,"['kernel method', 'DRAM reliability', 'online learning', 'failure propagation', 'DRAM failure prediction']",
442,Misclassification cost-sensitive fault prediction models,"['Yue Jiang', 'Bojan Cukic']",2009,"Traditionally, software fault prediction models are built by assuming a uniform misclassification cost. In other words, cost implications of misclassifying a faulty module as fault free are assumed to be the same as the cost implications of misclassifying a fault free module as faulty. In reality, these two types of misclassification costs are rarely equal. They are project-specific, reflecting the characteristics of the domain in which the program operates. In this paper, using project information from a public repository, we analyze the benefits of techniques which incorporate misclassification costs in the development of software fault prediction models. We find that cost-sensitive learning does not provide operational points which outperform cost-insensitive classifiers. However, an advantage of cost-sensitive modeling is the explicit choice of the operational threshold appropriate for the cost differential.",https://doi.org/10.1145/1540438.1540466,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 4}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault detection' OR 'failure detection')"", 'index': 58}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 7}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 3}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 6}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 5}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 4}]",17.0,,,,['failure-prediction'],,Proceedings of the 5th International Conference on Predictor Models in Software Engineering,True,['failure-management'],,True,,,,10.1145/1540438.1540466,Conference Paper,"['misclassification cost', 'fault prediction', 'machine learning', 'cost-sensitive']",
443,Mutation-aware fault prediction,"['David Bowes', 'Tracy Hall', 'Mark Harman', 'Yue Jia', 'Federica Sarro', 'Fan Wu']",2016,"We introduce mutation-aware fault prediction, which leverages additional guidance from metrics constructed in terms of mutants and the test cases that cover and detect them. We report the results of 12 sets of experiments, applying 4 different predictive modelling techniques to 3 large real-world systems (both open and closed source). The results show that our proposal can significantly (p ≤ 0.05) improve fault prediction performance. Moreover, mutation-based metrics lie in the top 5% most frequently relied upon fault predictors in 10 of the 12 sets of experiments, and provide the majority of the top ten fault predictors in 9 of the 12 sets of experiments.",https://doi.org/10.1145/2931037.2931039,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 5}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 6}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 2}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 3}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 3}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prediction' OR 'failure prediction')"", 'index': 2}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault localization' OR 'failure localization')"", 'index': 22}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 3}]",9.0,,,,['failure-prediction'],,Proceedings of the 25th International Symposium on Software Testing and Analysis,True,['failure-management'],,True,,,,10.1145/2931037.2931039,Conference Paper,"['Software Defect Prediction', 'Mutation Testing', 'Empirical Study', 'Software Metrics', 'Software Fault Prediction']",
444,Disk Failure Prediction in Data Centers via Online Learning,"['Jiang Xiao', 'Zhuang Xiong', 'Song Wu', 'Yusheng Yi', 'Hai Jin', 'Kan Hu']",2018,"Disk failure has become a major concern with the rapid expansion of storage systems in data centers. Based on SMART (Self-Monitoring, Analysis and Reporting Technology) attributes, many researchers derive disk failure prediction models using machine learning techniques. Despite the significant developments, the majority of works rely on offline training and thereby hinder their adaption to the continuous update of forthcoming data, suffering from the 'model aging' problem. We are therefore motivated to uncover the root cause -- the dynamic SMART distribution for 'model aging', aiming to resolve the performance degradation as to pave a comprehensive study in practice. In this paper, we introduce a novel disk failure prediction model using Online Random Forests (ORFs). Our ORF-based model can automatically evolve with sequential arrival of data on-the-fly and thus is highly adaptive to the variance of SMART distribution over time. Moreover, it has favourable advantage against the offline counterparts in terms of superior prediction performance. Experiments on real-world datasets show that our ORF model converges rapidly to the offline random forests and achieves stable failure detection rates of 93-99% with low false alarm rates. Furthermore, we demonstrate the ability of our approach on maintaining stable prediction performance for the long-term usage in data centers.",https://doi.org/10.1145/3225058.3225106,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 8}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault detection' OR 'failure detection')"", 'index': 19}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault detection' OR 'failure detection')"", 'index': 34}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 11}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 45}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 4}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 7}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 7}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 70}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 4}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 5}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 7}]",20.0,['random-forest'],['host-metrics'],['novel-use'],['failure-prediction'],['hard-drive'],Proceedings of the 47th International Conference on Parallel Processing,True,['failure-management'],True,True,['hardware-failure-prediction'],True,36.0,10.1145/3225058.3225106,Conference Paper,"['hard disk drive', 'storage system reliability', 'Failure prediction', 'online learning', 'SMART']",['online']
445,An iterative semi-supervised approach to software fault prediction,"['Huihua Lu', 'Bojan Cukic', 'Mark Culp']",2011,"Background: Many statistical and machine learning techniques have been implemented to build predictive fault models. Traditional methods are based on supervised learning. Software metrics for a module and corresponding fault information, available from previous projects, are used to train a fault prediction model. This approach calls for a large size of training data set and enables the development of effective fault prediction models. In practice, data collection costs, the lack of data from earlier projects or product versions may make large fault prediction training data set unattainable. Small size of the training set that may be available from the current project is known to deteriorate the performance of the fault predictive model. In semi-supervised learning approaches, software modules with known or unknown fault content can be used for training. Aims: To implement and evaluate a semi-supervised learning approach in software fault prediction. Methods: We investigate an iterative semi-supervised approach to software quality prediction in which a base supervised learner is used within a semi-supervised application. Results: We varied the size of labeled software modules from 2% to 50% of all the modules in the project. After tracking the performance of each iteration in the semi-supervised algorithm, we observe that semi-supervised learning improves fault prediction if the number of initially labeled software modules exceeds 5%. Conclusion: The semi-supervised approach outperforms the corresponding supervised learning approach when both use random forest as base classification algorithm.",https://doi.org/10.1145/2020390.2020405,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 11}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 18}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 9}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 8}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 11}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 11}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 7}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 12}]",8.0,,,,['failure-prediction'],,Proceedings of the 7th International Conference on Predictive Models in Software Engineering,True,['failure-management'],,True,,,,10.1145/2020390.2020405,Conference Paper,"['fault prediction', 'semi-supervised learning']",
446,Investigating fault prediction capabilities of five prediction models for software quality,"['Deepak Banthia', 'Atul Gupta']",2012,"Predicting faults in software modules can lead to a high quality and more effective software development process to follow. However, the results of a fault prediction model have to be properly interpreted before incorporating them into any decision making. Most of the earlier studies have used the prediction accuracy as the main criteria to compare amongst competing fault prediction models. However, we show that besides accuracy, other criteria like number of false positives and false negatives can equally be important to choose a candidate model for fault prediction. We have used five NASA software data sets in our experiment. Our results suggest that the performance of Simple Logistic is better than the others on raw data sets whereas the performance of Neural Network was found to be better when we applied dimensionality reduction method on raw data sets. When we used data pre-processing techniques, the prediction accuracy of Random Forest was found to be better in both cases i.e. with and without dimensionality reduction but reliability of Simple Logistic was better than Random Forest because it had less number of fault negatives.",https://doi.org/10.1145/2245276.2231975,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 16}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 29}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 15}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 12}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 18}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 10}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 10}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 21}]",4.0,"['multilayer-perceptron', 'random-forest', 'dimensionality-reduction', 'logistic-regression']",,['comparison'],['failure-prevention'],['software'],Proceedings of the 27th Annual ACM Symposium on Applied Computing,True,['failure-management'],,,['software-defect-prediction'],,,10.1145/2245276.2231975,Conference Paper,"['fault prediction', 'quality assurance', 'effort estimation', 'attribute selection', 'fault prediction models']",
447,PreFix: Switch Failure Prediction in Datacenter Networks,"['%A Shenglin Zhang', 'Ying Liu', 'Weibin Meng', 'Zhiling Luo', 'Jiahao Bu', 'Sen Yang', 'Peixian Liang', 'Dan Pei', 'Jun Xu', 'Yuzhi Zhang et al.']",2018,"In modern datacenter networks (DCNs), failures of network devices are the norm rather than the exception, and many research efforts have focused on dealing with failures after they happen. In this paper, we take a different approach by predicting failures, thus the operators can intervene and ""fix"" the potential failures before they happen. Specifically, in our proposed system, named PreFix, we aim to determine during runtime whether a switch failure will happen in the near future. The prediction is based on the measurements of the current switch system status and historical switch hardware failure cases that have been carefully labelled by network operators. Our key observation is that failures of the same switch model share some common syslog patterns before failures occur, and we can apply machine learning methods to extract the common patterns for predicting switch failures. Our novel set of features (message template sequence, frequency, seasonality and surge) for machine learning can efficiently deal with the challenges of noises, sample imbalance, and computation overhead. We evaluated PreFix on a data set collected from 9397 switches (3 different switch models) deployed in more than 20 datacenters owned by a top global search engine in a 2-year period. PreFix achieved an average of 61.81% recall and 1.84x10 -5 false positive ratio, outperforming the other failure prediction methods for computers and ISP devices.",https://doi.org/10.1145/3219617.3219643,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 18}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 20}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 3}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 4}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 28}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 18}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault detection' OR 'failure detection')"", 'index': 69}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault detection' OR 'failure detection')"", 'index': 81}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 91}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 51}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 53}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 31}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 32}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 23}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 55}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 27}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 28}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 58}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 54}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 79}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 30}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 15}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 87}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 15}, {'database': 'ACM', 'search_string': ""'regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 91}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 27}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 83}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 19}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 15}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 16}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 17}, {'database': 'ACM', 'search_string': ""'logistic regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 10}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 32}]",6.0,,,,['failure-prediction'],['switch'],Abstracts of the 2018 ACM International Conference on Measurement and Modeling of Computer Systems,True,['failure-management'],,,['hardware-failure-prediction'],,,10.1145/3219617.3219643,Conference Paper,"['operations', 'machine learning', 'failure prediction', 'datacenter']",
448,Software fault prediction for object oriented systems: a literature review,"['Ruchika Malhotra', 'Ankita Jain']",2011,"There always has been a demand to produce efficient and high quality software. There are various object oriented metrics that measure various properties of the software like coupling, cohesion, inheritance etc. which affect the software to a large extent. These metrics can be used in predicting important quality attributes such as fault proneness, maintainability, effort, productivity and reliability. Early prediction of fault proneness will help us to focus on testing resources and use them only on the classes which are predicted to be fault-prone. Thus, this will help in early phases of software development to give a measurement of quality assessment.This paper provides the review of the previous studies which are related to software metrics and the fault proneness. In other words, it reviews several journals and conference papers on software fault prediction. There is large number of software metrics proposed in the literature. Each study uses a different subset of these metrics and performs the analysis using different datasets. Also, the researchers have used different approaches such as Support vector machines, naive bayes network, random forest, artificial neural network, decision tree, logistic regression etc. Thus, this study focuses on the metrics used, dataset used and the evaluation or analysis method used by various authors. This review will be beneficial for the future studies as various researchers and practitioners can use it for comparative analysis.",https://doi.org/10.1145/2020976.2020991,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 32}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault detection' OR 'failure detection')"", 'index': 15}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault detection' OR 'failure detection')"", 'index': 64}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 46}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 26}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 14}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 12}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 15}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 83}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 27}]",2.0,,,,['failure-prediction'],,ACM SIGSOFT Software Engineering Notes,True,['failure-management'],,,,,,10.1145/2020976.2020991,Journal Article,"['empirical validation', 'metrics', 'software quality', 'fault proneness', 'object oriented']",
449,Transfer Learning based Failure Prediction for Minority Disks in Large Data Centers of Heterogeneous Disk Systems,"['Ji Zhang', 'Ke Zhou', 'Ping Huang', 'Xubin He', 'Zhili Xiao', 'Bin Cheng', 'Yongguang Ji', 'Yinhu Wang']",2019,"The storage system in large scale data centers is typically built upon thousands or even millions of disks, where disk failures constantly happen. A disk failure could lead to serious data loss and thus system unavailability or even catastrophic consequences if the lost data cannot be recovered. While replication and erasure coding techniques have been widely deployed to guarantee storage availability and reliability, disk failure prediction is gaining popularity as it has the potential to prevent disk failures from occurring in the first place. Recent trends have turned toward applying machine learning approaches based on disk SMART attributes for disk failure predictions. However, traditional machine learning (ML) approaches require a large set of training data in order to deliver good predictive performance. In large-scale storage systems, new disks enter gradually to augment the storage capacity or to replace failed disks, leading storage systems to consist of small amounts of new disks from different vendors and/or different models from the same vendor as time goes on. We refer to this relatively small amount of disks as minority disks. Due to the lack of sufficient training data, traditional ML approaches fail to deliver satisfactory predictive performance in evolving storage systems which consist of heterogeneous minority disks. To address this challenge and improve the predictive performance for minority disks in large data centers, we propose a minority disk failure prediction model named TLDFP based on a transfer learning approach. Our evaluation results on two realistic datasets have demonstrated that TLDFP can deliver much more precise results, compared to four popular prediction models based on traditional ML algorithms and two state-of-the-art transfer learning methods.",https://doi.org/10.1145/3337821.3337881,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 35}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 5}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 48}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault detection' OR 'failure detection')"", 'index': 65}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 59}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 98}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 37}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 33}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 40}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 66}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 17}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 19}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 22}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 37}]",0.0,"['random-forest', 'rnn', 'multilayer-perceptron', 'regression-tree']",['host-metrics'],"['new-method', 'comparison']",['failure-prediction'],"['storage', 'hard-drive']",Proceedings of the 48th International Conference on Parallel Processing,True,['failure-management'],,,['hardware-failure-prediction'],,,10.1145/3337821.3337881,Conference Paper,,
450,Machine Learning Methods for Predicting Failures in Hard Drives: A Multiple-Instance Application,"['Joseph F. Murray', 'Gordon F. Hughes', 'Kenneth Kreutz-Delgado']",2005,"We compare machine learning methods applied to a difficult real-world problem: predicting computer hard-drive failure using attributes monitored internally by individual drives. The problem is one of detecting rare events in a time series of noisy and nonparametrically-distributed data. We develop a new algorithm based on the multiple-instance learning framework and the naive Bayesian classifier (mi-NB) which is specifically designed for the low false-alarm case, and is shown to have promising performance. Other methods compared are support vector machines (SVMs), unsupervised clustering, and non-parametric statistical tests (rank-sum and reverse arrangements). The failure-prediction performance of the SVM, rank-sum and mi-NB algorithm is considerably better than the threshold method currently implemented in drives, while maintaining low false alarm rates. Our results suggest that nonparametric statistical tests should be considered for learning problems involving detecting rare events in time series data. An appendix details the calculation of rank-sum significance probabilities in the case of discrete, tied observations, and we give new recommendations about when the exact calculation should be used instead of the commonly-used normal approximation. These normal approximations may be particularly inaccurate for rare event problems like hard drive failures.",https://dl.acm.org/doi/10.5555/1046920.1088699,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 36}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 36}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 51}]",263.0,"['naive-bayes', 'support-vector-machine', 'clustering', 'gaussian-mixture-model']",,,['failure-prediction'],['hard-drive'],The Journal of Machine Learning Research,True,['failure-management'],True,,['hardware-failure-prediction'],True,28.0,10.5555/1046920.1088699,Journal Article,,
451,An Empirical Study on Fault Prediction using Token-Based Approach,"['Ishleen Kaur', 'Neha Bajpai']",2016,"Since exhaustive testing is not possible, prediction of fault prone modules can be used for prioritizing the components of a software system. Various approaches have been proposed for the prediction of fault prone modules. Most of them uses module metrics as quality estimators. In this study, we proposed a tokenbased approach and combine the metric evaluated from our approach with the module metrics to further improve the prediction results. We conducted the experiment on an open source project for evaluating the approach. The proposed approach is further compared with the existing fault prone filtering technique. The results show that the proposed approach is an improvement over fault prone filtering technique.",https://doi.org/10.1145/2979779.2979811,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 43}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 60}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 17}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 32}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prediction' OR 'failure prediction')"", 'index': 15}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 16}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 18}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 20}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 57}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 36}]",0.0,,,,['failure-prevention'],,Proceedings of the International Conference on Advances in Information Communication Technology & Computing,True,['failure-management'],,,['software-defect-prediction'],,,10.1145/2979779.2979811,Conference Paper,"['fault', 'software testing', 'logistic regression', 'Classification', 'fault prone modules', 'software metrics']",
452,An Attention-augmented Deep Architecture for Hard Drive Status Monitoring in Large-scale Storage Systems,"['Ji Wang', 'Weidong Bao', 'Lei Zheng', 'Xiaomin Zhu', 'Philip S. Yu']",2019,"Data centers equipped with large-scale storage systems are critical infrastructures in the era of big data. The enormous amount of hard drives in storage systems magnify the failure probability, which may cause tremendous loss for both data service users and providers. Despite a set of reactive fault-tolerant measures such as RAID, it is still a tough issue to enhance the reliability of large-scale storage systems. Proactive prediction is an effective method to avoid possible hard-drive failures in advance. A series of models based on the SMART statistics have been proposed to predict impending hard-drive failures. Nonetheless, there remain some serious yet unsolved challenges like the lack of explainability of prediction results. To address these issues, we carefully analyze a dataset collected from a real-world large-scale storage system and then design an attention-augmented deep architecture for hard-drive health status assessment and failure prediction. The deep architecture, composed of a feature integration layer, a temporal dependency extraction layer, an attention layer, and a classification layer, cannot only monitor the status of hard drives but also assist in failure cause diagnoses. The experiments based on real-world datasets show that the proposed deep architecture is able to assess the hard-drive status and predict the impending failures accurately. In addition, the experimental results demonstrate that the attention-augmented deep architecture can reveal the degradation progression of hard drives automatically and assist administrators in tracing the cause of hard drive failures.",https://doi.org/10.1145/3340290,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 56}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 7}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault detection' OR 'failure detection')"", 'index': 34}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault detection' OR 'failure detection')"", 'index': 53}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 78}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 46}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 31}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 41}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 44}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 20}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 22}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 27}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 48}]",0.0,['rnn'],,,['failure-prediction'],['hardware'],ACM Transactions on Storage (TOS),True,['failure-management'],,,['hardware-failure-prediction'],,,10.1145/3340290,Journal Article,"['SMART', 'attention mechanism', 'Hard drive failure', 'recurrent neural network', 'deep neural network']",
453,"SSD failures in the field: symptoms, causes, and prediction models","['Jacob Alter', 'Ji Xue', 'Alma Dimnaku', 'Evgenia Smirni']",2019,"In recent years, solid state drives (SSDs) have become a staple of high-performance data centers for their speed and energy efficiency. In this work, we study the failure characteristics of 30,000 drives from a Google data center spanning six years. We characterize the workload conditions that lead to failures and illustrate that their root causes differ from common expectation but remain difficult to discern. Particularly, we study failure incidents that result in manual intervention from the repair process. We observe high levels of infant mortality and characterize the differences between infant and non-infant failures. We develop several machine learning failure prediction models that are shown to be surprisingly accurate, achieving high recall and low false positive rates. These models are used beyond simple prediction as they aid us to untangle the complex interaction of workload characteristics that lead to failures and identify failure root causes from monitored symptoms.",https://doi.org/10.1145/3295500.3356172,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 57}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 86}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 49}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 34}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 63}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 45}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 62}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 14}, {'database': 'ACM', 'search_string': ""'regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 84}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 26}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 31}, {'database': 'ACM', 'search_string': ""'logistic regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 7}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 51}]",0.0,"['random-forest', 'logistic-regression', 'similarity-matching', 'multilayer-perceptron', 'support-vector-machine', 'decision-tree']",['host-metrics'],"['comparison', 'novel-use']",['failure-prediction'],['ssd'],"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis",True,['failure-management'],True,True,['hardware-failure-prediction'],False,,10.1145/3295500.3356172,Conference Paper,,
454,Desh: deep learning for system health prediction of lead times to failure in HPC,"['Anwesha Das', 'Frank Mueller', 'Charles Siegel', 'Abhinav Vishnu']",2018,"Today's large-scale supercomputers encounter faults on a daily basis. Exascale systems are likely to experience even higher fault rates due to increased component count and density. Triggering resilience-mitigating techniques remains a challenge due to the absence of well defined failure indicators. System logs consist of unstructured text that obscures essential system health information contained within. In this context, efficient failure prediction via log mining can enable proactive recovery mechanisms to increase reliability. This work aims to predict node failures that occur in supercomputing systems via long short-term memory (LSTM) networks that exploit recurrent neural networks (RNNs). Our framework, Desh1 (Deep Learning for System Health), diagnoses and predicts failures with short lead times. Desh identifies failure indicators with enhanced training and classification for generic applicability to logs from operating systems and software components without the need to modify any of them. Desh uses a novel three-phase deep learning approach to (1) train to recognize chains of log events leading to a failure, (2) re-train chain recognition of events augmented with expected lead times to failure, and (3) predict lead times during testing/inference deployment to predict which specific node fails in how many minutes. Desh obtains as high as 3 minutes average lead time with no less than 85% recall and 83% accuracy to take proactive actions on the failing nodes, which could be used to migrate computation to healthy nodes.",https://doi.org/10.1145/3208040.3208051,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 62}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 9}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 96}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 74}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 68}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 50}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 35}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 44}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prediction' OR 'failure prediction')"", 'index': 24}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 27}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('remediation' OR 'recovery')"", 'index': 34}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 30}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 50}]",7.0,,,,['failure-prediction'],,Proceedings of the 27th International Symposium on High-Performance Parallel and Distributed Computing,True,['failure-management'],,True,,,,10.1145/3208040.3208051,Conference Paper,"['LSTM', 'log mining', 'deep learning', 'lead times', 'HPC', 'anomaly detection', 'node failures', 'failure prediction']",
455,A Principled Approach to HPC Event Monitoring,"['Alireza Goudarzi', 'Dorian Arnold', 'Darko Stefanovic', 'Kurt B. Ferreira', 'Guy Feldman']",2015,"As high-performance computing (HPC) systems become larger and more complex, fault tolerance becomes a greater concern. At the same time, the data volume collected to help in understanding and mitigating hardware and software faults and failures also becomes prohibitively large. We argue that the HPC community must adopt more systematic approaches to system event logging as opposed to the current, ad hoc, strategies based on practitioner intuition and experience. Specifically, we show that event correlation and prediction can increase our understanding of fault behavior and can become critical components of effective fault tolerance strategies. While event correlation and prediction have been used in HPC contexts, we offer new insights about their potential capabilities. Using event logs from the computer failure data repository (cfdr) (1) we use cross and partial correlations to observe conditional correlations in HPC event data; (2) we use information theory to understand the fundamental predictive power of HPC failure data; (3) we study neural networks for failure prediction; and (4) finally, we use principal component analysis to understand to what extent dimensionality reduction can apply to HPC event data. This work results in the following insights that can inform HPC event monitoring: ad hoc correlations or ones based on direct correlations can be deficient or even misleading; highly accurate failure prediction may only require small windows of failure event history; and principal component analysis can significantly reduce HPC event data without loss of relevant information.",https://doi.org/10.1145/2751504.2751506,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 64}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 81}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 38}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 28}]",0.0,,,,['failure-prediction'],,Proceedings of the 5th Workshop on Fault Tolerance for HPC at eXtreme Scale,True,['failure-management'],,,,,,10.1145/2751504.2751506,Conference Paper,"['failure prediction', 'event analysis', 'fault-tolerance', 'resource monitoring']",
456,Predicting HDD failures from compound SMART attributes,"['Shiri Gaber', 'Oshry Ben-Harush', 'Amihai Savir']",2017,"Hard Disk Drives (HDD s) sometimes fail with no apparent reason; some SMART (Self-Monitoring, Analysis and Reporting Technology) attributes present strong correlations with drive failures[3], yet a drive may also fail without (supposedly) any previous indication. Most of the host systems today utilize alert-methods which are reactive by nature-a drive is indicated to fail when some SMART attribute exceeds its vendor defined threshold for valid operation[4]. This approach does not take the cross correlation between different attributes into account and the fact that thresholds vary across different vendors. Unlike more conventional studies that focused on reliability statistics such as the annualized failure rate (AFR) at the population level[5], [6], we use machine learning (ML) algorithms that attempt to predict the failure of individual drives. Former ML approaches applied to the drive prediction failure domain include methods for dealing with sequential data, such as sliding windows and hidden Markov models [1], or anomaly detection algorithms, adhering to the often-low proportion of failed drives in the population [4] and [2]. We present a mechanism that performs a sophisticated samples aggregation inside a distributed database, allowing for the efficient extraction of compound features representing the behavior of the drive during a continuous time window prior its failure. We use here an open source dataset from BACKBLAZE, comprising extensive SMART information collected daily from a large drive population at its data center Thus we can use not only the last cumulative SMART counts but also include new features extracted over different time windows of drives last operational days. These capture the dynamics of the drives attributes, for example their growth rate which is highly indicative of the drive s failure probability [7]. Our results suggest that the use of compound features reduce the amount of false positives, which is primary performance measure for algorithms in the drive reliability domain[3], by as much as 60%. In a consecutive evaluation scenario a final decision is made based on a collection of the most recent data samples. This way we are able to capture more soon-to-fail drives at low cost to the false-positive rate. On average, a drive is predicted to fail long enough in advance (30 days) to allow for the modification of business strategy in fields such as drive replacement logistics. This approach, when implemented in a production environment may have a direct effect on business savings related to drive logistics.",https://doi.org/10.1145/3078468.3081875,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 65}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 88}]",2.0,,,,['failure-prediction'],,Proceedings of the 10th ACM International Systems and Storage Conference,True,['failure-management'],,,,,,10.1145/3078468.3081875,Conference Paper,,
457,Predicting Node failure in cloud service systems,"['%A Qingwei Lin', 'Ken Hsieh', 'Yingnong Dang', 'Hongyu Zhang', 'Kaixin Sui', 'Yong Xu', 'Jian-Guang Lou', 'Chenggang Li', 'Youjiang Wu', 'Randolph Yao et al.']",2018,"In recent years, many traditional software systems have migrated to cloud computing platforms and are provided as online services. The service quality matters because system failures could seriously affect business and user experience. A cloud service system typically contains a large number of computing nodes. In reality, nodes may fail and affect service availability. In this paper, we propose a failure prediction technique, which can predict the failure-proneness of a node in a cloud service system based on historical data, before node failure actually happens. The ability to predict faulty nodes enables the allocation and migration of virtual machines to the healthy nodes, therefore improving service availability. Predicting node failure in cloud service systems is challenging, because a node failure could be caused by a variety of reasons and reflected by many temporal and spatial signals. Furthermore, the failure data is highly imbalanced. To tackle these challenges, we propose MING, a novel technique that combines: 1) a LSTM model to incorporate the temporal data, 2) a Random Forest model to incorporate spatial data; 3) a ranking model that embeds the intermediate results of the two models as feature inputs and ranks the nodes by their failure-proneness, 4) a cost-sensitive function to identify the optimal threshold for selecting the faulty nodes. We evaluate our approach using real-world data collected from a cloud service system. The results confirm the effectiveness of the proposed approach. We have also successfully applied the proposed approach in real industrial practice.",https://doi.org/10.1145/3236024.3236060,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 66}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('cloud')"", 'index': 58}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 80}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 48}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 29}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 42}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 76}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('remediation' OR 'recovery')"", 'index': 95}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prediction' OR 'failure prediction')"", 'index': 23}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('cloud computing')"", 'index': 86}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 23}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 29}, {'database': 'ACM', 'search_string': ""'logistic regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 26}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 49}]",3.0,,,,['failure-prediction'],,Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering,True,['failure-management'],,,,,,10.1145/3236024.3236060,Conference Paper,"['cloud service systems', 'service availability', 'Failure prediction', 'node failure', 'maintenance']",
458,Unsupervised Learning for Log Data Analysis Based on Behavior and Attribute Features,"['Xiaojuan Wang', 'Defu Wang', 'Yong Zhang', 'Lei Jin', 'Mei Song']",2019,"In some special application environments, network fault can lead to loss of important information or even mission failures, resulting in unpredictable losses. Therefore, it has certain research significance and practical value to evaluate the network status and predict the possible faults before performing the key tasks. Based on the logs collected by the router board in the real network, this paper analyses the behavior type, attribute information and the corresponding status value, and detects the hidden fault or network attack, so as to provide early warning information for operators. We propose a deep neural network model utilizing Long Short-Term Memory (LSTM) to predict the current number of level-1 logs. By comparing the predicted number of level-1 logs, it can detect abnormal behavior such as a surge in the number of logs. What's more, we perform semantic analysis on attribute information to construct attribute syntax forest, which assists maintenance staff to monitor the system through key fingerprint information in the log. In addition, we adopt attribute information and status value to train the unsupervised learning algorithm models such as Isolation Forest, OneClassSVM and LocalOutlierFactor. What's more, this paper analyses the results to find out the causes of log surge, and to assist operators in subsequent maintenance of the system.",https://doi.org/10.1145/3349341.3349460,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 87}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 59}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 91}, {'database': 'ACM', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 45}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 79}, {'database': 'ACM', 'search_string': ""'regression' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 16}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault prediction' OR 'failure prediction')"", 'index': 28}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 43}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 54}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 75}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 11}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 54}, {'database': 'ACM', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 41}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 70}]",0.0,,['logs'],,['failure-detection'],['router'],Proceedings of the 2019 International Conference on Artificial Intelligence and Computer Science,True,['failure-management'],,,['anomaly-detection'],,,10.1145/3349341.3349460,Conference Paper,"['Log Analysis', 'Network Fault', 'LSTM', 'Unsupervised Machine Learning']",
459,Predictive Analysis in Network Function Virtualization,"['Zhijing Li', 'Zihui Ge', 'Ajay Mahimkar', 'Jia Wang', 'Ben Y. Zhao', 'Haitao Zheng', 'Joanne Emmons', 'Laura Ogden']",2018,"Recent deployments of Network Function Virtualization (NFV) architectures have gained tremendous traction. While virtualization introduces benefits such as lower costs and easier deployment of network functions, it adds additional layers that reduce transparency into faults at lower layers. To improve fault analysis and prediction for virtualized network functions (VNF), we envision a runtime predictive analysis system that runs in parallel with existing reactive monitoring systems to provide network operators timely warnings against faulty conditions. In this paper, we propose a deep learning based approach to reliably identify anomaly events from NFV system logs, and perform an empirical study using 18 consecutive months in 2016--2018 of real-world deployment data on virtualized provider edge routers. Our deep learning models, combined with customization and adaptation mechanisms, can successfully identify anomalous conditions that correlate with network trouble tickets. Analyzing these anomalies can help operators to optimize trouble ticket generation and processing rules in order to enable fast, or even proactive actions against faulty conditions.",https://doi.org/10.1145/3278532.3278547,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 93}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 17}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 47}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 89}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 45}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 21}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prediction' OR 'failure prediction')"", 'index': 60}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 41}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 25}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 19}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 56}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 95}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 51}]",1.0,['multilayer-perceptron'],,,['failure-detection'],"['network', 'router']",Proceedings of the Internet Measurement Conference 2018,True,['failure-management'],,,,,,10.1145/3278532.3278547,Conference Paper,"['Network Function Virtualization', 'Machine Learning']",
460,Finding Anomalies in Network System Logs with Latent Variables,"['Kazuki Otomo', 'Satoru Kobayashi', 'Kensuke Fukuda', 'Hiroshi Esaki']",2018,"System logs are useful to understand the status of and detect faults in large scale networks. However, due to their diversity and volume of these logs, log analysis requires much time and effort. In this paper, we propose a log event anomaly detection method for large-scale networks without pre-processing and feature extraction. The key idea is to embed a large amount of diverse data into hidden states by using latent variables. We evaluate our method with 15 months of system logs obtained from a nation-wide academic network in Japan. Through comparisons with Kleinberg's univariate burst detection and a traditional multivariate analysis (i.e., PCA), we demonstrate that our proposed method detects anomalies and ease troubleshooting of network system faults.",https://doi.org/10.1145/3229607.3229608,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 94}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 42}, {'database': 'ACM', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 28}, {'database': 'ACM', 'search_string': ""'regression' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 10}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 50}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 21}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 5}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 94}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 17}, {'database': 'ACM', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 28}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 62}]",2.0,,,,['failure-detection'],,Proceedings of the 2018 Workshop on Big Data Analytics and Machine Learning for Data Communication Networks,True,['failure-management'],,True,,,,10.1145/3229607.3229608,Conference Paper,"['Variational autoencoder', 'Network log analysis', 'Latent variable analysis']",
461,Application of support vector machine to predict fault prone classes,"['Yogesh Singh', 'Arvinder Kaur', 'Ruchika Malhotra']",2009,"Empirical validation of software metrics to predict quality using machine learning methods is important to ensure their practical relevance in the software organizations. It would also be interesting to know the relationship between object-oriented metrics and fault proneness. In this paper, we build a Support Vector Machine (SVM) model to find the relation-ship between object-oriented metrics given by Chidamber and Kemerer and fault proneness. The proposed model is empirically evaluated using open source software. The performance of the SVM method was evaluated by Receiver Operating Characteristic (ROC) analysis. Based on these results, it is reasonable to claim that such models could help for planning and performing testing by focusing resources on fault- prone parts of the design and code. Thus, the study shows that SVM method may also be used in constructing software quality models. However, similar types of studies are required to be carried out in order to establish the acceptability of the model.",https://doi.org/10.1145/1457516.1457529,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 100}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault detection' OR 'failure detection')"", 'index': 27}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault detection' OR 'failure detection')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 7}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 11}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 72}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 2}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 88}]",10.0,,,,['failure-prediction'],,ACM SIGSOFT Software Engineering Notes,True,['failure-management'],,True,,,,10.1145/1457516.1457529,Journal Article,,
462,Automated known problem diagnosis with event traces,"['Chun Yuan', 'Ni Lao', 'Ji-Rong Wen', 'Jiwei Li', 'Zheng Zhang', 'Yi-Min Wang', 'Wei-Ying Ma']",2006,"Computer problem diagnosis remains a serious challenge to users and support professionals. Traditional troubleshooting methods relying heavily on human intervention make the process inefficient and the results inaccurate even for solved problems, which contribute significantly to user's dissatisfaction. We propose to use system behavior information such as system event traces to build correlations with solved problems, instead of using only vague text descriptions as in existing practices. The goal is to enable automatic identification of the root cause of a problem if it is a known one, which would further lead to its resolution. By applying statistical learning techniques to classifying system call sequences, we show our approach can achieve considerable accuracy of root cause recognition by studying four case examples.",https://doi.org/10.1145/1217935.1217972,True,"[{'database': 'ACM', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 89}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault localization' OR 'failure localization')"", 'index': 8}, {'database': 'ACM', 'search_string': ""'classification' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 93}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 80}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 12}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault localization' OR 'failure localization')"", 'index': 90}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault localization' OR 'failure localization')"", 'index': 78}, {'database': 'ACM', 'search_string': ""'classification' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 14}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 45}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 94}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 29}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 9}]",100.0,,['traces'],,['root-cause-analysis'],,Proceedings of the 1st ACM SIGOPS/EuroSys European Conference on Computer Systems 2006,True,['failure-management'],,,['root-cause-diagnosis'],,,10.1145/1217935.1217972,Conference Paper,"['root cause analysis', 'support vector machine', 'system call sequences']",
463,Program structure aware fault localization,"['Heng Li', 'Yuzhen Liu', 'Zhenyu Zhang', 'Jian Liu']",2014,"Software testing is always an effective method to show the presence of bugs in programs, while debugging is never an easy task to remove a bug from a program. To facilitate the debugging task, statistical fault localization estimates the location of faults in programs automatically by analyzing the program executions to narrow down the suspicious region. We observe that program structure has strong impacts on the assessed suspiciousness of program elements. However, existing techniques inadequately pay attention to this problem. In this paper, we emphasize the biases caused by program structure in fault localization, and propose a method to address them. Our method is dedicated to boost a fault localization technique by adapting it to various program structures, in a software development process. It collects the suspiciousness of program elements when locating historical faults, statistically captures the biases caused by program structure, and removes such an impact factor from a fault localization result. An empirical study using the Siemens test suite shows that our method can greatly improve the effectiveness of the most representative fault localization Tarantula.",https://doi.org/10.1145/2666581.2666593,True,"[{'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault localization' OR 'failure localization')"", 'index': 1}]",,,,,['root-cause-analysis'],,Proceedings of the International Workshop on Innovative Software Development Methodologies and Practices,True,['failure-management'],,,['fault-localization'],,,10.1145/2666581.2666593,Conference Paper,"['Software testing', 'program structure', 'fault localization']",
464,Causal inference for statistical fault localization,"['George K. Baah', 'Andy Podgurski', 'Mary Jean Harrold']",2010,"This paper investigates the application of causal inference methodology for observational studies to software fault localization based on test outcomes and profiles. This methodology combines statistical techniques for counterfactual inference with causal graphical models to obtain causal-effect estimates that are not subject to severe confounding bias. The methodology applies Pearl's Back-Door Criterion to program dependence graphs to justify a linear model for estimating the causal effect of covering a given statement on the occurrence of failures. The paper also presents the analysis of several proposed-fault localization metrics and their relationships to our causal estimator. Finally, the paper presents empirical results demonstrating that our model significantly improves the effectiveness of fault localization.",https://doi.org/10.1145/1831708.1831717,True,"[{'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 3}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault localization' OR 'failure localization')"", 'index': 4}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault localization' OR 'failure localization')"", 'index': 4}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 5}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 2}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault localization' OR 'failure localization')"", 'index': 11}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault localization' OR 'failure localization')"", 'index': 1}]",57.0,,,,['root-cause-analysis'],,Proceedings of the 19th international symposium on Software testing and analysis,True,['failure-management'],,,,,,10.1145/1831708.1831717,Conference Paper,"['program analysis', 'debugging', 'potential outcome model', 'causal inference', 'fault localization']",
465,An automated model-based debugging approach,"['Cemal Yilmaz', 'Clay Williams']",2007,"Program debugging is a difficult and time-consuming task. Our ultimate goal in this work is to help developers reduce the space of potential root causes for failures, which can, in turn, improve the turn around time for bug fixes. We propose a novel and very different approach. Rather then focusing on how a program behaves by analyzing its source code and/or execution traces, we concentrate on how it should behave with respect to a given behavioral model. We identify and verify slices of the behavior model, that, once implemented wrong in the program, can potentially lead to failures. Not only do we identify functional differences between the programand its model, but we also provide a ranked list of diagnoses which might explain (or be associated with) these differences. Our experiments suggest that the proposed approach can be quite effective in reducing the search space for potential root causes for failures",https://doi.org/10.1145/1321631.1321659,True,"[{'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 49}]",12.0,,,,['root-cause-analysis'],,Proceedings of the twenty-second IEEE/ACM international conference on Automated software engineering,True,['failure-management'],,,,,,10.1145/1321631.1321659,Conference Paper,"['model-based problem determination', 'fault localization', 'automated debugging']",
466,Identifying impactful service system problems via log analysis,"['Shilin He', 'Qingwei Lin', 'Jian-Guang Lou', 'Hongyu Zhang', 'Michael R. Lyu', 'Dongmei Zhang']",2018,"Logs are often used for troubleshooting in large-scale software systems. For a cloud-based online system that provides 24/7 service, a huge number of logs could be generated every day. However, these logs are highly imbalanced in general, because most logs indicate normal system operations, and only a small percentage of logs reveal impactful problems. Problems that lead to the decline of system KPIs (Key Performance Indicators) are impactful and should be fixed by engineers with a high priority. Furthermore, there are various types of system problems, which are hard to be distinguished manually. In this paper, we propose Log3C, a novel clustering-based approach to promptly and precisely identify impactful system problems, by utilizing both log sequences (a sequence of log events) and system KPIs. More specifically, we design a novel cascading clustering algorithm, which can greatly save the clustering time while keeping high accuracy by iteratively sampling, clustering, and matching log sequences. We then identify the impactful problems by correlating the clusters of log sequences with system KPIs. Log3C is evaluated on real-world log data collected from an online service system at Microsoft, and the results confirm its effectiveness and efficiency. Furthermore, our approach has been successfully applied in industrial practice.",https://doi.org/10.1145/3236024.3236083,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 6}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 21}, {'database': 'ACM', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 3}, {'database': 'ACM', 'search_string': ""'regression' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 3}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 21}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 3}, {'database': 'ACM', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 4}]",7.0,,,,['root-cause-analysis'],,Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering,True,['failure-management'],,,,,,10.1145/3236024.3236083,Conference Paper,"['Problem Identification', 'Service Systems', 'Log Analysis', 'Clustering']",
467,Modeling virtualized applications using machine learning techniques,"['Sajib Kundu', 'Raju Rangaswami', 'Ajay Gulati', 'Ming Zhao', 'Kaushik Dutta']",2012,"With the growing adoption of virtualized datacenters and cloud hosting services, the allocation and sizing of resources such as CPU, memory, and I/O bandwidth for virtual machines (VMs) is becoming increasingly important. Accurate performance modeling of an application would help users in better VM sizing, thus reducing costs. It can also benefit cloud service providers who can offer a new charging model based on the VMs' performance instead of their configured sizes. In this paper, we present techniques to model the performance of a VM-hosted application as a function of the resources allocated to the VM and the resource contention it experiences. To address this multi-dimensional modeling problem, we propose and refine the use of two machine learning techniques: artificial neural network (ANN) and support vector machine (SVM). We evaluate these modeling techniques using five virtualized applications from the RUBiS and Filebench suite of benchmarks and demonstrate that their median and 90th percentile prediction errors are within 4.36% and 29.17% respectively. These results are substantially better than regression based approaches as well as direct applications of machine learning techniques without our refinements. We also present a simple and effective approach to VM sizing and empirically demonstrate that it can deliver optimal results for 65% of the sizing problems that we studied and produces close-to-optimal sizes for the remaining 35%.",https://doi.org/10.1145/2151024.2151028,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 78}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 35}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 77}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 24}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 47}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 86}, {'database': 'ACM', 'search_string': ""'logistic regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 77}]",78.0,,,,['workload-prediction'],,Proceedings of the 8th ACM SIGPLAN/SIGOPS conference on Virtual Execution Environments,True,['resource-provisioning'],,True,,,,10.1145/2151024.2151028,Conference Paper,"['VM sizing', 'cloud data centers', 'performance modeling', 'machine learning', 'virtualization']",
468,Obtaining dynamic scheduling policies with simulation and machine learning,"['Danilo Carastan-Santos', 'Raphael Y. de Camargo']",2017,"Dynamic scheduling of tasks in large-scale HPC platforms is normally accomplished using ad-hoc heuristics, based on task characteristics, combined with some backfilling strategy. Defining heuristics that work efficiently in different scenarios is a difficult task, specially when considering the large variety of task types and platform architectures. In this work, we present a methodology based on simulation and machine learning to obtain dynamic scheduling policies. Using simulations and a workload generation model, we can determine the characteristics of tasks that lead to a reduction in the mean slowdown of tasks in an execution queue. Modeling these characteristics using a nonlinear function and applying this function to select the next task to execute in a queue improved the mean task slowdown in synthetic workloads. When applied to real workload traces from highly different machines, these functions still resulted in performance improvements, attesting the generalization capability of the obtained heuristics.",https://doi.org/10.1145/3126908.3126955,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 98}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 46}]",5.0,,,,,,"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis",True,['resource-provisioning'],,,,,,10.1145/3126908.3126955,Conference Paper,"['high performance computing', 'machine learning', 'scheduling', 'simulation']",
469,A machine learning approach to TCP throughput prediction,"['Mariyam Mirza', 'Joel Sommers', 'Paul Barford', 'Xiaojin Zhu']",2010,"TCP throughput prediction is an important capability for networks where multiple paths exist between data senders and receivers. In this paper, we describe a new lightweight method for TCP throughput prediction. Our predictor uses Support Vector Regression (SVR); prediction is based on both prior file transfer history and measurements of simple path properties. We evaluate our predictor in a laboratory setting where ground truth can be measured with perfect accuracy. We report the performance of our predictor for oracular and practical measurements of path properties over a wide range of traffic conditions and transfer sizes. For bulk transfers in heavy traffic using oracular measurements, TCP throughput is predicted within 10% of the actual value 87% of the time, representing nearly a threefold improvement in accuracy over prior history-based methods. For practical measurements of path properties, predictions can be made within 10% of the actual value nearly 50% of the time, approximately a 60% improvement over history-based methods, and with much lower measurement traffic overhead. We implement our predictor in a tool called Path-Perf, test it in the wide area, and show that PathPerf predicts TCP throughput accurately over diverse wide area paths.",https://doi.org/10.1109/TNET.2009.2037812,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 100}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 49}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 90}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 91}]",67.0,"['support-vector-machine', 'linear-regression']",,,['workload-prediction'],['network'],Proceedings of the 2007 ACM SIGMETRICS international conference on Measurement and modeling of computer systems,True,['resource-provisioning'],,,,,,10.1109/TNET.2009.2037812,Journal Article,"['active measurements', 'support vector regression', 'TCP throughput prediction', 'machine learning']",['low-overhead']
470,Empowering automatic data-center management with machine learning,"['Josep Ll. Berral', 'Ricard Gavaldà', 'Jordi Torres']",2013,"The Cloud as computing paradigm has become nowadays crucial for most Internet business models. Managing and optimizing its performance on a moment-by-moment basis is not easy given as the amount and diversity of elements involved (hardware, applications, workloads, customer needs...). Here we show how a combination of scheduling algorithms and data mining techniques helps improving the performance and profitability of a data-center running virtualized web-services. We model the data-center's main resources (CPU, memory, IO), quality of service (viewed as response time), and workloads (incoming streams of requests) from past executions. We show how these models to help scheduling algorithms make better decisions about job and resource allocation, aiming for a balance between throughput, quality of service, and power consumption.",https://doi.org/10.1145/2480362.2480397,True,"[{'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 7}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 14}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 37}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 72}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 71}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 27}, {'database': 'ACM', 'search_string': ""'regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 5}]",6.0,,,,,,Proceedings of the 28th Annual ACM Symposium on Applied Computing,True,['resource-provisioning'],,True,,,,10.1145/2480362.2480397,Conference Paper,"['web-services', 'machine learning', 'cloud computing', 'modeling']",
471,Minimizing data center cooling and server power costs,"['Ehsan Pakbaznia', 'Massoud Pedram']",2009,"This paper focuses on power minimization in a data center accounting for both the information technology equipment and the air conditioning power usage. In particular we address the server consolidation (on/off state assignment) concurrently with the task assignment. We formulate the resulting optimization problem as an Integer Linear Programming problem and present a heuristic algorithm that solves it in polynomial time. Experimental results show an average of 13% power saving for different data center utilization rates compared to a baseline task assignment technique, which does not perform server consolidation.",https://doi.org/10.1145/1594233.1594268,True,"[{'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 9}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 70}]",107.0,['integer-programming'],,,['power-management'],['power'],Proceedings of the 2009 ACM/IEEE international symposium on Low power electronics and design,True,['resource-provisioning'],,,,,,10.1145/1594233.1594268,Conference Paper,['datacenter'],
472,Autonomic multi-agent management of power and performance in data centers,"['Rajarshi Das', 'Jeffrey O. Kephart', 'Charles Lefurgy', 'Gerald Tesauro', 'David W. Levine', 'Hoi Chan']",2008,"The rapidly rising cost and environmental impact of energy consumption in data centers has become a multi-billion dollar concern globally. In response, the IT Industry is actively engaged in a first-to-market race to develop energy-conserving hardware and software solutions that do not sacrifice performance objectives. In this work we demonstrate a prototype of an integrated data center power management solution that employs server management tools, appropriate sensors and monitors, and an agent-based approach to achieve specified power and performance objectives. By intelligently turning off servers under low-load conditions, we can achieve over 25% power savings over the unmanaged case without incurring SLA penalties for typical daily and weekly periodic demands seen in webserver farms.",https://dl.acm.org/doi/10.5555/1402795.1402816,True,"[{'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 46}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 71}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 11}]",153.0,,,,['power-management'],,Proceedings of the 7th international joint conference on Autonomous agents and multiagent systems: industrial track,True,['resource-provisioning'],,False,,,,10.5555/1402795.1402816,Conference Paper,"['power measurement', 'multicriteria utility functions', 'energy savings', 'green data center', 'power management', 'policy-based management', 'data center']",
473,Renewable and cooling aware workload management for sustainable data centers,"['Zhenhua Liu', 'Yuan Chen', 'Cullen Bash', 'Adam Wierman', 'Daniel Gmach', 'Zhikui Wang', 'Manish Marwah', 'Chris Hyser']",2012,"Recently, the demand for data center computing has surged, increasing the total energy footprint of data centers worldwide. Data centers typically comprise three subsystems: IT equipment provides services to customers; power infrastructure supports the IT and cooling equipment; and the cooling infrastructure removes heat generated by these subsystems. This work presents a novel approach to model the energy flows in a data center and optimize its operation. Traditionally, supply-side constraints such as energy or cooling availability were treated independently from IT workload management. This work reduces electricity cost and environmental impact using a holistic approach that integrates renewable supply, dynamic pricing, and cooling supply including chiller and outside air cooling, with IT workload planning to improve the overall sustainability of data center operations. Specifically, we first predict renewable energy as well as IT demand. Then we use these predictions to generate an IT workload management plan that schedules IT workload and allocates IT resources within a data center according to time varying power supply and cooling efficiency. We have implemented and evaluated our approach using traces from real data centers and production systems. The results demonstrate that our approach can reduce both the recurring power costs and the use of non-renewable energy by as much as 60% compared to existing techniques, while still meeting the Service Level Agreements.",https://doi.org/10.1145/2254756.2254779,True,"[{'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 50}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 5}]",244.0,,,['new-method'],"['resource-consolidation', 'power-management', 'workload-prediction']",,Proceedings of the 12th ACM SIGMETRICS/PERFORMANCE joint international conference on Measurement and Modeling of Computer Systems,True,['resource-provisioning'],,True,,,,10.1145/2254756.2254779,Conference Paper,"['sustainable data center', 'demand shaping', 'renewable energy', 'cooling optimization', 'scheduling']",
474,Towards energy-aware scheduling in data centers using machine learning,"['Josep Ll. Berral', 'Íñigo Goiri', 'Ramón Nou', 'Ferran Julià', 'Jordi Guitart', 'Ricard Gavaldà', 'Jordi Torres']",2010,"As energy-related costs have become a major economical factor for IT infrastructures and data-centers, companies and the research community are being challenged to find better and more efficient power-aware resource management strategies. There is a growing interest in ""Green"" IT and there is still a big gap in this area to be covered. In order to obtain an energy-efficient data center, we propose a framework that provides an intelligent consolidation methodology using different techniques such as turning on/off machines, power-aware consolidation algorithms, and machine learning techniques to deal with uncertain information while maximizing performance. For the machine learning approach, we use models learned from previous system behaviors in order to predict power consumption levels, CPU loads, and SLA timings, and improve scheduling decisions. Our framework is vertical, because it considers from watt consumption to workload features, and cross-disciplinary, as it uses a wide variety of techniques. We evaluate these techniques with a framework that covers the whole control cycle of a real scenario, using a simulation with representative heterogeneous workloads, and we measure the quality of the results according to a set of metrics focused toward our goals, besides traditional policies. The results obtained indicate that our approach is close to the optimal placement and behaves better when the level of uncertainty increases.",https://doi.org/10.1145/1791314.1791349,True,"[{'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 60}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 86}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 26}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 13}, {'database': 'ACM', 'search_string': ""'regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 26}]",260.0,"['linear-regression', 'decision-tree']","['requests', 'sla', 'host-metrics']",['novel-use'],"['scheduling', 'power-management']",,Proceedings of the 1st International Conference on Energy-Efficient Computing and Networking,True,['resource-provisioning'],True,True,,,,10.1145/1791314.1791349,Conference Paper,"['power efficiency', 'simulation', 'machine learning', 'data center', 'scheduling']",
475,PeerWatch: a fault detection and diagnosis tool for virtualized consolidation systems,"['Hui Kang', 'Haifeng Chen', 'Guofei Jiang']",2010,"Server virtualization is now becoming an effective means to consolidate numerous applications into a small number of machines. While such a strategy can lead to significant savings in power and hardware cost, it may complicate the fault management task due to the increasing scalability and complexity in the virtualized environment. In this paper, we propose PeerWatch, a fault detection and diagnosis tool specially designed for virtualized consolidation systems. Based on the observation that each application usually reveals itself in multiple instances in the virtualized data center, PeerWatch introduces a statistical technique, canonical correlation analysis (CCA), to extract the correlated characteristics between multiple application instances. The extracted correlations are utilized to examine the status of each application instance. If some correlations drop significantly during the operation, PeerWatch regards that the system is in faulty situation and produces alarms. PeerWatch is robust to system dynamics, compared to traditional fault detection techniques and thus can avoid a lot of false alarms. Once the fault has been detected, PeerWatch proposes a diagnosis process that also takes advantage of the multiple instances feature in the virtualized systems. The diagnosis combines the spatial and temporal analysis on the measurement data across multiple instances before and after the failure. As a result, PeerWatch can obtain much accurate clues about the fault root cause. Experimental results in our virtualized testbed system have demonstrated the effectiveness of the proposed detection and diagnosis tool.",https://doi.org/10.1145/1809049.1809070,True,"[{'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 65}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 92}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault localization' OR 'failure localization')"", 'index': 57}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 24}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 86}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 95}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 16}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault localization' OR 'failure localization')"", 'index': 34}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 14}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 22}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 53}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 33}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault detection' OR 'failure detection')"", 'index': 85}]",24.0,,,,['failure-detection'],,Proceedings of the 7th international conference on Autonomic computing,True,['failure-management'],,,,,,10.1145/1809049.1809070,Conference Paper,"['fault diagnosis', 'virtualization', 'canonical correlation analysis', 'data center', 'fault detection']",
476,Automated control of multiple virtualized resources,"['Pradeep Padala', 'Kai-Yuan Hou', 'Kang G. Shin', 'Xiaoyun Zhu', 'Mustafa Uysal', 'Zhikui Wang', 'Sharad Singhal', 'Arif Merchant']",2009,"Virtualized data centers enable sharing of resources among hosted applications. However, it is difficult to satisfy service-level objectives(SLOs) of applications on shared infrastructure, as application workloads and resource consumption patterns change over time. In this paper, we present AutoControl, a resource control system that automatically adapts to dynamic workload changes to achieve application SLOs. AutoControl is a combination of an online model estimator and a novel multi-input, multi-output (MIMO) resource controller. The model estimator captures the complex relationship between application performance and resource allocations, while the MIMO controller allocates the right amount of multiple virtualized resources to achieve application SLOs. Our experimental evaluation with RUBiS and TPC-W benchmarks along with production-trace-driven workloads indicates that AutoControl can detect and mitigate CPU and disk I/O bottlenecks that occur over time and across multiple nodes by allocating each resource accordingly. We also show that AutoControl can be used to provide service differentiation according to the application priorities during resource contention.",https://doi.org/10.1145/1519065.1519068,True,"[{'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 84}]",296.0,['autoregression'],,,"['resource-consolidation', 'workload-prediction']",,Proceedings of the 4th ACM European conference on Computer systems,True,['resource-provisioning'],True,,,,,10.1145/1519065.1519068,Conference Paper,"['data center', 'control theory', 'application qos', 'automated control', 'virtualization', 'resource management', 'server consolidation']",
477,A regression-based approach to scalability prediction,"['Bradley J. Barnes', 'Barry Rountree', 'David K. Lowenthal', 'Jaxk Reeves', 'Bronis de Supinski', 'Martin Schulz']",2008,"Many applied scientific domains are increasingly relying on large-scale parallel computation. Consequently, many large clusters now have thousands of processors. However, the ideal number of processors to use for these scientific applications varies with both the input variables and the machine under consideration, and predicting this processor count is rarely straightforward. Accurate prediction mechanisms would provide many benefits, including improving cluster efficiency and identifying system configuration or hardware issues that impede performance. We explore novel regression-based approaches to predict parallel program scalability. We use several program executions on a small subset of the processors to predict execution time on larger numbers of processors. We compare three different regression-based techniques: one based on execution time only; another that uses per-processor information only; and a third one based on the global critical path. These techniques provide accurate scaling predictions, with median prediction errors between 6.2% and 17.3% for seven applications.",https://doi.org/10.1145/1375527.1375580,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 33}]",172.0,['linear-regression'],"['runs', 'host-metrics', 'tasks']",['novel-use'],['workload-prediction'],['cluster'],Proceedings of the 22nd annual international conference on Supercomputing,True,['resource-provisioning'],,,,,,10.1145/1375527.1375580,Conference Paper,"['regression', 'modeling', 'prediction', 'scalability', 'MPI']",
478,Autonomic mix-aware provisioning for non-stationary data center workloads,"['Rahul Singh', 'Upendra Sharma', 'Emmanuel Cecchet', 'Prashant Shenoy']",2010,"Online Internet applications see dynamic workloads that fluctuate over multiple time scales. This paper argues that the non-stationarity in Internet application workloads, which causes the request mix to change over time, can have a significant impact on the overall processing demands imposed on data center servers. We propose a novel mix-aware dynamic provisioning technique that handles both the non-stationarity in the workload as well as changes in request volumes when allocating server capacity in Internet data centers. Our technique employs the k-means clustering algorithm to automatically determine the workload mix and a queuing model to predict the server capacity for a given workload mix. We implement a prototype provisioning system that incorporates our technique and experimentally evaluate its efficacy on a laboratory Linux data center running the TPC-W web benchmark. Our results show that our k-means clustering technique accurately captures workload mix changes in Internet applications. We also demonstrate that mix-aware dynamic provisioning eliminates SLA violations due to under-provisioning with non-stationary web workloads, and that it offers a better resource usage by reducing over-provisioning when compared to a baseline provisioning approach that only reacts to workload volume changes. We also present a case study of our provisioning approach on Amazon's EC2 cloud platform.",https://doi.org/10.1145/1809049.1809053,True,"[{'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 26}, {'database': 'ACM', 'search_string': ""'clustering' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 9}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 14}, {'database': 'ACM', 'search_string': ""'classification' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 13}]",89.0,,,,['workload-prediction'],,Proceedings of the 7th international conference on Autonomic computing,True,['resource-provisioning'],,True,,,,10.1145/1809049.1809053,Conference Paper,"['modeling', 'availability', 'performance']",
479,Clustering event logs using iterative partitioning,"['Adetokunbo A.O. Makanju', 'A. Nur Zincir-Heywood', 'Evangelos E. Milios']",2009,"The importance of event logs, as a source of information in systems and network management cannot be overemphasized. With the ever increasing size and complexity of today's event logs, the task of analyzing event logs has become cumbersome to carry out manually. For this reason recent research has focused on the automatic analysis of these log files. In this paper we present IPLoM (Iterative Partitioning Log Mining), a novel algorithm for the mining of clusters from event logs. Through a 3-Step hierarchical partitioning process IPLoM partitions log data into its respective clusters. In its 4th and final stage IPLoM produces cluster descriptions or line formats for each of the clusters produced. Unlike other similar algorithms IPLoM is not based on the Apriori algorithm and it is able to find clusters in data whether or not its instances appear frequently. Evaluations show that IPLoM outperforms the other algorithms statistically significantly, and it is also able to achieve an average F-Measure performance 78% when the closest other algorithm achieves an F-Measure performance of 10%.",https://doi.org/10.1145/1557019.1557154,True,"[{'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 34}, {'database': 'ACM', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 80}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 35}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 50}]",101.0,['clustering'],"['events', 'logs']",,['failure-detection'],,Proceedings of the 15th ACM SIGKDD international conference on Knowledge discovery and data mining,True,['failure-management'],,,['anomaly-detection'],,,10.1145/1557019.1557154,Conference Paper,"['fault management', 'event log mining', 'telecommunications']",
480,Symptom-based problem determination using log data abstraction,"['Liang Huang', 'Xiaodi Ke', 'Kenny Wong', 'Serge Mankovskii']",2010,"System failures in industry are expensive, and the increasingly stringent requirements on performance and reliability of enterprise systems have made the detection and diagnosis of system failures crucial and challenging. Log files generated at the system runtime are considered to contain the representations of failure symptoms, and thus become one of the most important sources used for system monitoring and failure diagnosis. A number of studies suggest that data mining and machine learning can help in dealing with the vast amount of log data for a complex enterprise system. Log data abstraction techniques have been proposed, but have not been well studied for failure detection and problem determination. In this research, we investigate the effects of using an unsupervised log data abstraction method to aid the supervised learning processes of problem determination. Additionally, we compare the efficiency of associative classification methods for failure diagnosis against Bayesian Learning technique and C4.5 that have been proved good both in documentation classification and failure diagnosis. Our experimental results show that two associative classification methods outperform Naive Bayes and C4.5 when applied on non-abstracted logs, and unsupervised log abstraction helps to improve the performance of log-based problem determination significantly in terms of the precision, F-measure, and efficiency.",https://doi.org/10.1145/1923947.1923979,True,"[{'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault prediction' OR 'failure prediction')"", 'index': 35}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 76}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault detection' OR 'failure detection')"", 'index': 33}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault detection' OR 'failure detection')"", 'index': 65}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 73}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 42}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 25}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 25}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 20}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 94}]",9.0,['decision-tree'],['logs'],"['comparison', 'novel-use']",['root-cause-analysis'],,Proceedings of the 2010 Conference of the Center for Advanced Studies on Collaborative Research,True,['failure-management'],,True,,,,10.1145/1923947.1923979,Conference Paper,,
481,Probabilistic performance modeling of virtualized resource allocation,"['Brian J. Watson', 'Manish Marwah', 'Daniel Gmach', 'Yuan Chen', 'Martin Arlitt', 'Zhikui Wang']",2010,"Virtualization technologies enable organizations to dynamically flex their IT resources based on workload fluctuations and changing business needs. However, only through a formal understanding of the relationship between application performance and virtualized resource allocation can over-provisioning or over-loading of physical IT resources be avoided. In this paper, we examine the probabilistic relationships between virtualized CPU allocation, CPU contention, and application response time, to enable autonomic controllers to satisfy service level objectives (SLOs) while more effectively utilizing IT resources. We show that with only minimal knowledge of application and system behaviors, our methodology can model the probability distribution of response time with a mean absolute error of less than 6% when compared with the measured response time distribution. We then demonstrate the usefulness of a probabilistic approach with case studies. We apply basic laws of probability to our model to investigate whether and how CPU allocation and contention affect application response time, correcting for their effects on CPU utilization. We find mean absolute differences of 8-10% between the modeled response time distributions of certain allocation states, and a similar difference when we add CPU contention. This methodology is general, and should also be applicable to non-CPU virtualized resources and other performance modeling problems.",https://doi.org/10.1145/1809049.1809067,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('IT operations')"", 'index': 3}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('IT operations')"", 'index': 10}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('IT operations')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('IT operations')"", 'index': 55}]",64.0,,,,['workload-prediction'],,Proceedings of the 7th international conference on Autonomic computing,True,['resource-provisioning'],,,,,,10.1145/1809049.1809067,Conference Paper,"['probability theory', 'quantile regression', 'performance modeling']",
482,CogNETive: insights and visualization for operations@scale,"['Dean Lorenz', 'Eran Raichstein', 'Katherine Barabash', 'Hillel Kolodner', 'Liran Schour', 'Shelly Garion']",2017,"Operating a cloud-scale service is a huge challenge. There are millions of users worldwide and millions of requests per seconds. For example, Amazon's Simple Storage Service (S3) in 2013 contained two trillion objects and its logs contained 1.1 million log lines per second, which are approximately 10 PB of log records per year (see [1]). Cloud scale implies thousands of servers and network elements, and hundreds of services from multiple cross-regional data centers. Cloud service operation data is scattered over various types of semi-structured and unstructured logs (e.g., application, error, debug), telemetry and network data, as well as customer service records. It is therefore extremely difficult for the multiple owners and administrators in such systems, coming from different units of the organization, to follow the possible paths and system alternatives in order to detect problems, solve issues and understand the service operation.",https://doi.org/10.1145/3078468.3078495,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('IT operations')"", 'index': 6}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('IT operations')"", 'index': 8}, {'database': 'ACM', 'search_string': ""'clustering' AND ('IT operations')"", 'index': 7}, {'database': 'ACM', 'search_string': ""'clustering' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 33}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 74}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('IT operations')"", 'index': 65}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 61}, {'database': 'ACM', 'search_string': ""'regression' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 65}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('IT operations')"", 'index': 7}]",0.0,,,['discussion'],,,Proceedings of the 10th ACM International Systems and Storage Conference,True,['aiops-general'],,,,,,10.1145/3078468.3078495,Conference Paper,,
483,Data-driven resource flexing for network functions visualization,"['Lianjie Cao', 'Sonia Fahmy', 'Puneet Sharma', 'Shandian Zhe']",2018,"Resource flexing is the notion of allocating resources on-demand as workload changes. This is a key advantage of Virtualized Network Functions (VNFs) over their non-virtualized counterparts. However, it is difficult to balance the timeliness and resource efficiency when making resource flexing decisions due to unpredictable workloads and complex VNF processing logic. In this work, we propose an Elastic resource flexing system for Network functions VIrtualization (ENVI) that leverages a combination of VNF-level features and infrastructure-level features to construct a neural-network-based scaling decision engine for generating timely scaling decisions. To adapt to dynamic workloads, we design a window-based rewinding mechanism to update the neural network with emerging workload patterns and make accurate decisions in real time. Our experimental results for real VNFs (IDS Suricata and caching proxy Squid) using workloads generated based on real-world traces, show that ENVI provisions significantly fewer (up to 26%) resources without violating service level objectives, compared to commonly used rule-based scaling policies.",https://doi.org/10.1145/3230718.3230725,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('IT operations')"", 'index': 18}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('IT operations')"", 'index': 26}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('IT operations')"", 'index': 2}, {'database': 'ACM', 'search_string': ""'classification' AND ('IT operations')"", 'index': 38}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('IT operations')"", 'index': 23}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('IT operations')"", 'index': 3}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('IT operations')"", 'index': 33}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('IT operations')"", 'index': 15}]",1.0,,,,['resource-consolidation'],,Proceedings of the 2018 Symposium on Architectures for Networking and Communications Systems,True,['resource-provisioning'],,,,,,10.1145/3230718.3230725,Conference Paper,,
484,Software fault localization using feature selection,"['Shounak Roychowdhury', 'Sarfraz Khurshid']",2011,"Manually locating and fixing faults can be tedious and hard. Recent years have seen much progress in automated techniques for fault localization. A particularly promising approach is to analyze passing and failing runs to compute how likely each statement is to be faulty. Techniques based on this approach have so far largely focused on either using statistical analysis or similarity based algorithms, which have a natural application in evaluating such runs. We present a novel approach to fault localization using feature selection techniques from machine learning. Our insight is that each additional failing or passing run can provide significantly diverse amount of information, which can help localize faults in code -- the statements with maximum feature diversity information can point to most suspicious lines of code. Experimental results show that our approach outperforms state-of-the-art approaches for localizing faults in most subject programs of the Siemens suite, which have previously been used to evaluate several fault localization techniques.",https://doi.org/10.1145/2070821.2070823,True,"[{'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault localization' OR 'failure localization')"", 'index': 1}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault localization' OR 'failure localization')"", 'index': 5}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault localization' OR 'failure localization')"", 'index': 5}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 1}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault localization' OR 'failure localization')"", 'index': 3}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault localization' OR 'failure localization')"", 'index': 1}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault localization' OR 'failure localization')"", 'index': 3}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault localization' OR 'failure localization')"", 'index': 15}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault localization' OR 'failure localization')"", 'index': 2}]",8.0,,['test-cases'],,['root-cause-analysis'],,Proceedings of the International Workshop on Machine Learning Technologies in Software Engineering,True,['failure-management'],,True,['fault-localization'],,,10.1145/2070821.2070823,Conference Paper,"['automated debugging', 'fault localization', 'feature selection', 'machine learning', 'statistical debugging', 'RELIEF']",
485,A new hybrid algorithm for software fault localization,"['Jeongho Kim', 'Jonghee Park', 'Eunseok Lee']",2015,"We previously presented a spectrum-based fault localization (SFL) technique, which we named Hybrid, that localizes a bug by using the program hit spectra and test results. We also proposed a distinct mechanism for test data that enables the SFL algorithms to localize fault in a more precise manner than what would be possible with the original test data. However, there was a limitation of the Hybrid algorithm. In that the technique only showed better performance when using distinct test data. Therefore, in the current work, we improve over Hybrid by analyzing more than 30 types of existing algorithms. After choosing the appropriate algorithms, we adopted their specific strengths through experimentation. Finally, we developed the novel Combination Algorithms (CAL). In our experimental study, we used the Siemens test program to confirm that our technique was more precise than the state-of-the-art SFL algorithm D Star and Heuristic III for both original and distinct test data. In particular, our technique localizes a fault under 2 percent on average as well as decreases the coverage of the reading code by a third of the source code.",https://doi.org/10.1145/2701126.2701207,True,"[{'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault localization' OR 'failure localization')"", 'index': 2}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault localization' OR 'failure localization')"", 'index': 8}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 13}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault localization' OR 'failure localization')"", 'index': 7}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault localization' OR 'failure localization')"", 'index': 38}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 24}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 69}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 93}]",4.0,,,,['root-cause-analysis'],,Proceedings of the 9th International Conference on Ubiquitous Information Management and Communication,True,['failure-management'],,True,,,,10.1145/2701126.2701207,Conference Paper,"['spectrum based fault localization', 'execution trace', 'suspicious code', 'fault localization', 'program debugging']",
486,Taking the Blame Game out of Data Centers Operations with NetPoirot,"['Behnaz Arzani', 'Selim Ciraci', 'Boon Thau Loo', 'Assaf Schuster', 'Geoff Outhred']",2016,"Today, root cause analysis of failures in data centers is mostly done through manual inspection. More often than not, cus- tomers blame the network as the culprit. However, other components of the system might have caused these failures. To troubleshoot, huge volumes of data are collected over the entire data center. Correlating such large volumes of diverse data collected from different vantage points is a daunting task even for the most skilled technicians. In this paper, we revisit the question: how much can you infer about a failure in the data center using TCP statistics collected at one of the endpoints? Using an agent that cap- tures TCP statistics we devised a classification algorithm that identifies the root cause of failure using this information at a single endpoint. Using insights derived from this classi- fication algorithm we identify dominant TCP metrics that indicate where/why problems occur in the network. We val- idate and test these methods using data that we collect over a period of six months in a production data center.",https://doi.org/10.1145/2934872.2934884,True,"[{'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault localization' OR 'failure localization')"", 'index': 16}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 38}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 41}, {'database': 'ACM', 'search_string': ""'classification' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 78}, {'database': 'ACM', 'search_string': ""'classification' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 78}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault localization' OR 'failure localization')"", 'index': 48}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 35}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 41}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 13}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 13}, {'database': 'ACM', 'search_string': ""'classification' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 33}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 42}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 54}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 5}]",21.0,,,,['root-cause-analysis'],,Proceedings of the 2016 ACM SIGCOMM Conference,True,['failure-management'],,,,,,10.1145/2934872.2934884,Conference Paper,"['Network performance analysis', 'Network monitoring', 'Network transport protocols', 'Network reliability', 'Network manageability']",
487,Automatically learning semantic features for defect prediction,"['Song Wang', 'Taiyue Liu', 'Lin Tan']",2016,"Software defect prediction, which predicts defective code regions, can help developers find bugs and prioritize their testing efforts. To build accurate prediction models, previous studies focus on manually designing features that encode the characteristics of programs and exploring different machine learning algorithms. Existing traditional features often fail to capture the semantic differences of programs, and such a capability is needed for building accurate prediction models. To bridge the gap between programs' semantics and defect prediction features, this paper proposes to leverage a powerful representation-learning algorithm, deep learning, to learn semantic representation of programs automatically from source code. Specifically, we leverage Deep Belief Network (DBN) to automatically learn semantic features from token vectors extracted from programs' Abstract Syntax Trees (ASTs). Our evaluation on ten open source projects shows that our automatically learned semantic features significantly improve both within-project defect prediction (WPDP) and cross-project defect prediction (CPDP) compared to traditional features. Our semantic features improve WPDP on average by 14.7% in precision, 11.5% in recall, and 14.2% in F1. For CPDP, our semantic features based approach outperforms the state-of-the-art technique TCA+ with traditional features by 8.9% in F1.",https://doi.org/10.1145/2884781.2884804,True,"[{'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault localization' OR 'failure localization')"", 'index': 21}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault localization' OR 'failure localization')"", 'index': 98}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault localization' OR 'failure localization')"", 'index': 21}]",239.0,"['multilayer-perceptron', 'bayesian-network']",['source-code'],"['new-method', 'comparison']",['failure-prevention'],['source-code'],Proceedings of the 38th International Conference on Software Engineering,True,['failure-management'],True,True,['software-defect-prediction'],True,13.0,10.1145/2884781.2884804,Conference Paper,,
488,Performance Anomaly Detection and Bottleneck Identification,"['Olumuyiwa Ibidunmoye', 'Francisco Hernández-Rodriguez', 'Erik Elmroth']",2015,"In order to meet stringent performance requirements, system administrators must effectively detect undesirable performance behaviours, identify potential root causes, and take adequate corrective measures. The problem of uncovering and understanding performance anomalies and their causes (bottlenecks) in different system and application domains is well studied. In order to assess progress, research trends, and identify open challenges, we have reviewed major contributions in the area and present our findings in this survey. Our approach provides an overview of anomaly detection and bottleneck identification research as it relates to the performance of computing systems. By identifying fundamental elements of the problem, we are able to categorize existing solutions based on multiple factors such as the detection goals, nature of applications and systems, system observability, and detection methods.",https://doi.org/10.1145/2791120,True,"[{'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault localization' OR 'failure localization')"", 'index': 34}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('anomaly detection' OR 'outlier detection')"", 'index': 31}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 93}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault detection' OR 'failure detection')"", 'index': 54}, {'database': 'ACM', 'search_string': ""'regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 21}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault localization' OR 'failure localization')"", 'index': 31}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 11}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 44}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 36}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault localization' OR 'failure localization')"", 'index': 80}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 91}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 29}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('anomaly detection' OR 'outlier detection')"", 'index': 27}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 32}, {'database': 'ACM', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 72}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 93}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 87}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 36}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 89}, {'database': 'ACM', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 63}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 18}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault localization' OR 'failure localization')"", 'index': 31}, {'database': 'ACM', 'search_string': ""'logistic regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 74}]",59.0,,"['kpis', 'host-metrics']",['survey'],['failure-detection'],[],ACM Computing Surveys (CSUR),True,['failure-management'],True,True,['anomaly-detection'],,,10.1145/2791120,Journal Article,"['performance anomaly detection', 'Systems performance', 'performance problem identification', 'bottleneck detection']",
489,GRANO: interactive graph-based root cause analysis for cloud-native distributed data platform,"['Hanzhang Wang', 'Phuong Nguyen', 'Jun Li', 'Selcuk Kopru', 'Gene Zhang', 'Sanjeev Katariya', 'Sami Ben-Romdhane']",2019,"We demonstrate Grano1, an end-to-end anomaly detection and root cause analysis (or RCA for short) system for cloud-native distributed data platform by providing a holistic view of the system component topology, alarms and application events. Grano provides: a Detection Layer to process large amount of time-series monitoring data to detect anomalies at logical and physical system components; an Anomaly Graph Layer with novel graph modeling and algorithms for leveraging system topology data and detection results to identify the root cause relevance at the system component level; and an Application Layer that automatically notifies on-call personnel and presents real-time and on-demand RCA support through an interactive graph interface. The system is deployed and evaluated using eBay's production data to help on-call personnel to shorten the identification of root cause from hours to minutes.",https://doi.org/10.14778/3352063.3352105,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 8}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 23}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 11}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 5}]",2.0,['graph-mining'],,,"['root-cause-analysis', 'failure-detection']",,Proceedings of the VLDB Endowment,True,['failure-management'],,,['rca-others'],,,10.14778/3352063.3352105,Journal Article,,
490,Black-box problem diagnosis in parallel file systems,"['Michael P. Kasick', 'Jiaqi Tan', 'Rajeev Gandhi', 'Priya Narasimhan']",2010,"We focus on automatically diagnosing different performance problems in parallel file systems by identifying, gathering and analyzing OS-level, black-box performance metrics on every node in the cluster. Our peer-comparison diagnosis approach compares the statistical attributes of these metrics across I/O servers, to identify the faulty node. We develop a root-cause analysis procedure that further analyzes the affected metrics to pinpoint the faulty resource (storage or network), and demonstrate that this approach works commonly across stripe-based parallel file systems. We demonstrate our approach for realistic storage and network problems injected into three different file-system benchmarks (dd, IOzone, and Post-Mark), in both PVFS and Lustre clusters.",https://dl.acm.org/doi/10.5555/1855511.1855515,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 22}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 66}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 19}]",77.0,,"['software-metrics', 'host-metrics']",,['root-cause-analysis'],"['cluster', 'filesystem']",Proceedings of the 8th USENIX conference on File and storage technologies,True,['failure-management'],,,['root-cause-diagnosis'],,,10.5555/1855511.1855515,Conference Paper,,
491,Hound: Causal Learning for Datacenter-scale Straggler Diagnosis,"['Pengfei Zheng', 'Benjamin C. Lee']",2018,"Stragglers are exceptionally slow tasks within a job that delay its completion. Stragglers, which are uncommon within a single job, are pervasive in datacenters with many jobs. A large body of research has focused on mitigating datacenter stragglers, but relatively little research has focused on systematically and rigorously identifying their root causes. We present Hound, a statistical machine learning framework that infers the causes of stragglers from traces of datacenter-scale jobs. Hound is designed to achieve several objectives: datacenter-scale diagnosis, interpretable models, unbiased inference, and computational efficiency. We demonstrate Hound's capabilities for a production trace from Google's warehouse-scale datacenters and two Spark traces from Amazon EC2 clusters.",https://doi.org/10.1145/3179420,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 40}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 63}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 67}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 81}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 37}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 25}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 98}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 35}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 90}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 15}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('cloud computing')"", 'index': 100}, {'database': 'ACM', 'search_string': ""'logistic regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 11}]",5.0,,,,['root-cause-analysis'],,Proceedings of the ACM on Measurement and Analysis of Computing Systems,True,['failure-management'],,True,,,,10.1145/3179420,Journal Article,"['causal reasoning', 'topic modeling', 'distributed system', 'performance modeling', 'performance diagnosis', 'datacenter', 'machine learning']",
492,Diagnosis of recurrent faults using log files,"['Thomas Reidemeister', 'Mohammad Ahmad Munawar', 'Miao Jiang', 'Paul A. S. Ward']",2009,"Enterprise software systems (ESS) are becoming larger and increasingly complex. Failure in business-critical systems is expensive, leading to consequences such as loss of critical data, loss of sales, customer dissatisfaction, even law suits. Therefore, detecting failures and diagnosing their root-cause in a timely manner is essential. Many studies suggest that a large fraction of failures encountered in practice are recurrent (i.e., they have been seen before). Fast and accurate detection of these failures can accelerate problem determination, and thereby improve system reliability. To this effect, we explore machine learning techniques, including the Naïve Bayes classifier, partially-supervised learning, and decision trees (using C4.5), to automatically recognize symptoms of recurrent faults and to derive detection rules from samples of log data. This work focuses on log files, since they are readily available and they do not put any additional computational burden on the component generating the data. The methods explored in this work can aid the development of tools to assist support personnel in problem determination tasks. Instead of requiring the operators to manually define patterns for identifying recurrent problems, such tools can be trained using prior, solved and unsolved cases from existing support databases.",https://doi.org/10.1145/1723028.1723031,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 55}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 64}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 32}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 6}, {'database': 'ACM', 'search_string': ""'classification' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 59}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 100}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 31}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 47}]",3.0,"['naive-bayes', 'decision-tree']",['logs'],,['root-cause-analysis'],,Proceedings of the 2009 Conference of the Center for Advanced Studies on Collaborative Research,True,['failure-management'],,,,,,10.1145/1723028.1723031,Conference Paper,,
493,System performance anomaly detection using tracing data analysis,"['Iman Kohyarnejadfard', 'Mahsa Shakeri', 'Daniel Aloise']",2019,"In recent years, distributed systems have become increasingly complex as they grow in both scale and functionality. Such complexity makes these systems prone to performance anomalies. Efficient anomaly detection frameworks enable rapid recovery mechanisms to increase the system's reliability. In this paper, we present an anomaly detection approach for practical monitoring of processes running on a system to detect anomalous vectors of system calls. Our proposed methodology employs a Linux tracing toolkit (LTTng) to monitor the processes running on a system and extracts the streams of system calls. The system calls streams are split into short sequences using a sliding window strategy. Unlike previous studies, our proposed approach computes the execution time of system calls in addition to the frequency of each individual call in a window. Finally, a multi-class support vector machine approach is applied to evaluate the performance of the system and detect the anomalous sequences. A comprehensive experimental study on a real dataset collected using LTTng demonstrates that our proposed method is able to distinguish normal sequences from anomalous ones with CPU or memory related problems.",https://doi.org/10.1145/3323933.3324085,True,"[{'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('anomaly detection' OR 'outlier detection')"", 'index': 15}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 30}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 27}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('remediation' OR 'recovery')"", 'index': 23}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 43}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 80}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 91}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 63}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 100}, {'database': 'ACM', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 47}]",0.0,['support-vector-machine'],"['software-metrics', 'host-metrics']",,['failure-detection'],,Proceedings of the 2019 5th International Conference on Computer and Technology Applications,True,['failure-management'],,,,,,10.1145/3323933.3324085,Conference Paper,"['Anomaly detection', 'Time series', 'Linux Tracing', 'Data mining', 'Performance evaluation', 'Machine learning']",
494,Unsupervised Anomaly Detection via Variational Auto-Encoder for Seasonal KPIs in Web Applications,"['%A Haowen Xu', 'Wenxiao Chen', 'Nengwen Zhao', 'Zeyan Li', 'Jiahao Bu', 'Zhihan Li', 'Ying Liu', 'Youjian Zhao', 'Dan Pei', 'Yang Feng et al.']",2018,"To ensure undisrupted business, large Internet companies need to closely monitor various KPIs (e.g., Page Views, number of online users, and number of orders) of its Web applications, to accurately detect anomalies and trigger timely troubleshooting/mitigation. However, anomaly detection for these seasonal KPIs with various patterns and data quality has been a great challenge, especially without labels. In this paper, we proposed Donut, an unsupervised anomaly detection algorithm based on VAE. Thanks to a few of our key techniques, Donut greatly outperforms a state-of-arts supervised ensemble approach and a baseline VAE approach, and its best F-scores range from 0.75 to 0.9 for the studied KPIs from a top global Internet company. We come up with a novel KDE interpretation of reconstruction for Donut, making it the first VAE-based anomaly detection algorithm with solid theoretical explanation.",https://doi.org/10.1145/3178876.3185996,True,"[{'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('anomaly detection' OR 'outlier detection')"", 'index': 51}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 79}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 64}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 57}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 91}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 243}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 102}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 54}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 11}]",89.0,['autoencoder'],['kpis'],['novel-use'],['failure-detection'],['web-server'],Proceedings of the 2018 World Wide Web Conference,True,['failure-management'],True,,['anomaly-detection'],True,63.0,10.1145/3178876.3185996,Conference Paper,"['seasonal KPI', 'variational auto-encoder', 'anomaly detection']",
495,Robust log-based anomaly detection on unstable log data,"['%A Xu Zhang', 'Yong Xu', 'Qingwei Lin', 'Bo Qiao', 'Hongyu Zhang', 'Yingnong Dang', 'Chunyu Xie', 'Xinsheng Yang', 'Qian Cheng', 'Ze Li et al.']",2019,"Logs are widely used by large and complex software-intensive systems for troubleshooting. There have been a lot of studies on log-based anomaly detection. To detect the anomalies, the existing methods mainly construct a detection model using log event data extracted from historical logs. However, we find that the existing methods do not work well in practice. These methods have the close-world assumption, which assumes that the log data is stable over time and the set of distinct log events is known. However, our empirical study shows that in practice, log data often contains previously unseen log events or log sequences. The instability of log data comes from two sources: 1) the evolution of logging statements, and 2) the processing noise in log data. In this paper, we propose a new log-based anomaly detection approach, called LogRobust. LogRobust extracts semantic information of log events and represents them as semantic vectors. It then detects anomalies by utilizing an attention-based Bi-LSTM model, which has the ability to capture the contextual information in the log sequences and automatically learn the importance of different log events. In this way, LogRobust is able to identify and handle unstable log events and sequences. We have evaluated LogRobust using logs collected from the Hadoop system and an actual online service system of Microsoft. The experimental results show that the proposed approach can well address the problem of log instability and achieve accurate and robust results on real-world, ever-changing log data.",https://doi.org/10.1145/3338906.3338931,True,"[{'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('anomaly detection' OR 'outlier detection')"", 'index': 58}, {'database': 'ACM', 'search_string': ""'regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 45}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 16}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 50}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 88}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 77}, {'database': 'ACM', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 31}, {'database': 'ACM', 'search_string': ""'regression' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 12}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 67}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 89}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 71}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 6}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 19}, {'database': 'ACM', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 31}]",7.0,['rnn'],['logs'],,['failure-detection'],['hadoop'],Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering,True,['failure-management'],True,,['anomaly-detection'],,,10.1145/3338906.3338931,Conference Paper,"['Anomaly Detection', 'Data Quality', 'Deep Learning', 'Log Analysis', 'Log Instability']",['robustness']
496,Shrink: a tool for failure diagnosis in IP networks,"['Srikanth Kandula', 'Dina Katabi', 'Jean-Philippe Vasseur']",2005,"Faults in an IP network have various causes such as the failure of one or more routers at the IP layer, fiber-cuts, failure of physical elements at the optical layer, or extraneous causes like power outages. These faults are usually detected as failures of a set of dependent logical entities--the IP links affected by the failed components. We present Shrink, a tool for root cause analysis of network faults which, given a set of failed IP links, identifies the underlying cause of the faulty state. Shrink models the diagnosis problem as a Bayesian network. It has two main contributions. First, it effectively accounts for noisy measurement and inaccurate mapping between the IP and optical layers. Second, it has an efficient inference algorithm that finds the most likely failure causes in polynomial time and with bounded errors. We compare Shrink with two prior approaches and show that it substantially improves the performance.",https://doi.org/10.1145/1080173.1080178,True,"[{'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 40}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault localization' OR 'failure localization')"", 'index': 42}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 90}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 44}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 12}]",230.0,['bayesian-network'],"['configuration', 'events']",,['root-cause-analysis'],"['network', 'link']",Proceedings of the 2005 ACM SIGCOMM workshop on Mining network data,True,['failure-management'],True,,['fault-localization'],True,79.0,10.1145/1080173.1080178,Conference Paper,"['fault diagnosis', 'Bayesian', 'shrink', 'IP networks', 'SRLG', 'optical']",
497,"Capturing, indexing, clustering, and retrieving system history","['Ira Cohen', 'Steve Zhang', 'Moises Goldszmidt', 'Julie Symons', 'Terence Kelly', 'Armando Fox']",2005,"We present a method for automatically extracting from a running system an indexable signature that distills the essential characteristic from a system state and that can be subjected to automated clustering and similarity-based retrieval to identify when an observed system state is similar to a previously-observed state. This allows operators to identify and quantify the frequency of recurrent problems, to leverage previous diagnostic efforts, and to establish whether problems seen at different installations of the same site are similar or distinct. We show that the naive approach to constructing these signatures based on simply recording the actual ``raw'' values of collected measurements is ineffective, leading us to a more sophisticated approach based on statistical modeling and inference. Our method requires only that the system's metric of merit (such as average transaction response time) as well as a collection of lower-level operational metrics be collected, as is done by existing commercial monitoring tools. Even if the traces have no annotations of prior diagnoses of observed incidents (as is typical), our technique successfully clusters system states corresponding to similar problems, allowing diagnosticians to identify recurring problems and to characterize the ``syndrome'' of a group of problems. We validate our approach on both synthetic traces and several weeks of production traces from a customer-facing geoplexed 24 x 7 system; in the latter case, our approach identified a recurring problem that had required extensive manual diagnosis, and also aided the operators in correcting a previous misdiagnosis of a different problem.",https://doi.org/10.1145/1095810.1095821,True,"[{'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 98}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault localization' OR 'failure localization')"", 'index': 95}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault localization' OR 'failure localization')"", 'index': 54}, {'database': 'ACM', 'search_string': ""'classification' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 82}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 28}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 33}]",338.0,"['similarity-matching', 'clustering']","['kpis', 'traces']",['new-method'],['root-cause-analysis'],,Proceedings of the twentieth ACM symposium on Operating systems principles,True,['failure-management'],True,,['rca-others'],True,84.0,10.1145/1095810.1095821,Conference Paper,"['performance objectives', 'information retrieval', 'clustering', 'bayesian networks', 'signatures']",
498,Energy Efficiency Techniques in Cloud Computing: A Survey and Taxonomy,"['Tarandeep Kaur', 'Inderveer Chana']",2015,"The increase in energy consumption is the most critical problem worldwide. The growth and development of complex data-intensive applications have promulgated the creation of huge data centers that have heightened the energy demand. In this article, the need for energy efficiency is emphasized by discussing the dual role of cloud computing as a major contributor to increasing energy consumption and as a method to reduce energy wastage. This article comprehensively and comparatively studies existing energy efficiency techniques in cloud computing and provides the taxonomies for the classification and evaluation of the existing studies. The article concludes with a summary providing valuable suggestions for future enhancements.",https://doi.org/10.1145/2742488,True,"[{'database': 'ACM', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 9}, {'database': 'ACM', 'search_string': ""'clustering' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 56}, {'database': 'ACM', 'search_string': ""'classification' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 70}, {'database': 'ACM', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 28}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 26}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 51}]",115.0,,,['survey'],['power-management'],,ACM Computing Surveys (CSUR),True,['resource-provisioning'],True,,,,,10.1145/2742488,Journal Article,"['multicores', 'Resource Management System (RMS)', 'data center', 'virtualization', 'Cloud computing', 'Information and Communication Technology (ICT)', 'consolidation', 'resource scheduling', 'Virtual Machines (VMs)', 'energy efficiency']",
499,Survey on Machine Learning based scheduling in Cloud Computing,"['Naveen Kumar Gondhi', 'Ayushi Gupta']",2017,"In the modern era, cloud computing gains a lot of attention due to its various features such as it is simple to use, minimum cost, and mostly low power consumption. Many algorithms and techniques have been proposed for scheduling of virtual machines to provide dynamic load balancing, dynamic scalability and reallocation of resources. Intelligent algorithms are used for the optimization of results and minimizing the makespan scheduling while utilizing the resources efficiently based on dynamic environment. This paper reviews various intelligent scheduling algorithms such as Genetic Algorithm (GA), Simulated annealing (SA), Tabu Search (TS), Ant Colony Optimization (ACO), Particle Swarm Optimization (PSO), Artificial Immune System (AIS), Bacterial Foraging Algorithm (BF), Fish Swarm Optimization Algorithm (FS), Cat Swarm Optimization Algorithm (CS), Firefly Algorithm (FF), Cuckoo Search Algorithm (CS), Artificial Bee Colony (ABC), Bat Algorithm (BA).",https://doi.org/10.1145/3059336.3059352,True,"[{'database': 'ACM', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 15}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 1}, {'database': 'ACM', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 39}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 87}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 46}]",7.0,,,['survey'],['scheduling'],,"Proceedings of the 2017 International Conference on Intelligent Systems, Metaheuristics & Swarm Intelligence",True,['resource-provisioning'],True,,,,,10.1145/3059336.3059352,Conference Paper,"['Machine learning', 'Cloud Computing', 'Scheduling Algorithm']",
500,The cloud computing load forecasting algorithm based on wavelet support vector machine,"['Wei Zhong', 'Yi Zhuang', 'Jian Sun', 'Jingjing Gu']",2017,"In this paper, we propose a model based on the wavelet support vector machine(WSVM), which combines the wavelet transform's advantage of analyzing the cycle and frequency of the input signal with the support vector machine's characteristic of nonlinear regression analysis, to model the task load in the cloud computing center. Then we propose a cloud computing load forecasting algorithm based on WSVM. Finally, we verify the forecasting results using the data set of Google cloud computing center. The results prove that the algorithm we proposed performs better comparing with the similar forecasting algorithms in forecasting effect and accuracy.",https://doi.org/10.1145/3014812.3014852,True,"[{'database': 'ACM', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 16}, {'database': 'ACM', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 13}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 42}, {'database': 'ACM', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 40}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 1}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 13}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 3}, {'database': 'ACM', 'search_string': ""'regression' AND ('cloud')"", 'index': 100}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 62}]",5.0,,,,,,Proceedings of the Australasian Computer Science Week Multiconference,True,['resource-provisioning'],,True,,,,10.1145/3014812.3014852,Conference Paper,"['wavelet transform', 'support vector machine', 'load forecasting', 'cloud computing']",
501,QoS-Aware Autonomic Resource Management in Cloud Computing: A Systematic Review,"['Sukhpal Singh', 'Inderveer Chana']",2015,"As computing infrastructure expands, resource management in a large, heterogeneous, and distributed environment becomes a challenging task. In a cloud environment, with uncertainty and dispersion of resources, one encounters problems of allocation of resources, which is caused by things such as heterogeneity, dynamism, and failures. Unfortunately, existing resource management techniques, frameworks, and mechanisms are insufficient to handle these environments, applications, and resource behaviors. To provide efficient performance of workloads and applications, the aforementioned characteristics should be addressed effectively. This research depicts a broad methodical literature analysis of autonomic resource management in the area of the cloud in general and QoS (Quality of Service)-aware autonomic resource management specifically. The current status of autonomic resource management in cloud computing is distributed into various categories. Methodical analysis of autonomic resource management in cloud computing and its techniques are described as developed by various industry and academic groups. Further, taxonomy of autonomic resource management in the cloud has been presented. This research work will help researchers find the important characteristics of autonomic resource management and will also help to select the most suitable technique for autonomic resource management in a specific application along with significant future research directions.",https://doi.org/10.1145/2843889,True,"[{'database': 'ACM', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 22}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 16}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 70}, {'database': 'ACM', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 48}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 15}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault prediction' OR 'failure prediction')"", 'index': 94}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault prediction' OR 'failure prediction')"", 'index': 67}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 46}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 69}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 76}]",62.0,[],,['survey'],['resource-consolidation'],,ACM Computing Surveys (CSUR),True,['resource-provisioning'],True,True,,,,10.1145/2843889,Journal Article,"['self-healing', 'autonomic management', 'cloud computing', 'Resource provisioning', 'resource scheduling', 'self-protecting', 'self-configuring', 'service-level agreement', 'autonomic computing', 'quality of service', 'self-management', 'self-optimizing', 'resource management', 'grid computing', 'autonomic cloud computing']",
502,Issues and Challenges of Load Balancing Techniques in Cloud Computing: A Survey,"['Pawan Kumar', 'Rakesh Kumar']",2019,"With the growth in computing technologies, cloud computing has added a new paradigm to user services that allows accessing Information Technology services on the basis of pay-per-use at any time and any location. Owing to flexibility in cloud services, numerous organizations are shifting their business to the cloud and service providers are establishing more data centers to provide services to users. However, it is essential to provide cost-effective execution of tasks and proper utilization of resources. Several techniques have been reported in the literature to improve performance and resource use based on load balancing, task scheduling, resource management, quality of service, and workload management. Load balancing in the cloud allows data centers to avoid overloading/underloading in virtual machines, which itself is a challenge in the field of cloud computing. Therefore, it becomes a necessity for developers and researchers to design and implement a suitable load balancer for parallel and distributed cloud environments. This survey presents a state-of-the-art review of issues and challenges associated with existing load-balancing techniques for researchers to develop more effective algorithms.",https://doi.org/10.1145/3281010,True,"[{'database': 'ACM', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 26}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 86}, {'database': 'ACM', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 58}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 45}]",9.0,,,['survey'],['scheduling'],,ACM Computing Surveys (CSUR),True,['resource-provisioning'],True,,,,,10.1145/3281010,Journal Article,"['cloud computing', 'workload management', 'virtual machine', 'resource allocation', 'task scheduling', 'optimization', 'Load balancing']",
503,PerfScope: Practical Online Server Performance Bug Inference in Production Cloud Computing Infrastructures,"['Daniel J. Dean', 'Hiep Nguyen', 'Xiaohui Gu', 'Hui Zhang', 'Junghwan Rhee', 'Nipun Arora', 'Geoff Jiang']",2014,"Performance bugs which manifest in a production cloud computing infrastructure are notoriously difficult to diagnose because of both the difficulty of reproducing those bugs and the lack of debugging information. In this paper, we present PerfScope, a practical online performance bug inference tool to help the developer understand how a performance bug happened during the production run. PerfScope achieves online bug inference to obviate the need for offline bug reproduction. PerfScope does not require application source code or any runtime instrumentation to the production system. PerfScope is application-agnostic, which can support both interpreted and compiled programs running inside a cloud infrastructure. We have implemented PerfScope and tested it using real performance bugs on seven popular open source server systems (Hadoop, HDFS, Cassandra, Tomcat, Apache, Lighttpd, MySQL). The results show that PerfScope can narrow down the search scope of the bug-related functions to a small percentage (0.03-2.3%) and rank the real bug-related functions within top five candidates in the majority of cases. PerfScope only imposes on average 1.8% runtime overhead to the tested server applications.",https://doi.org/10.1145/2670979.2670987,True,"[{'database': 'ACM', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 48}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 40}, {'database': 'ACM', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 36}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault localization' OR 'failure localization')"", 'index': 73}]",25.0,,,,['root-cause-analysis'],"['hadoop', 'cassandra', 'tomcat', 'apache', 'lighttpd', 'mysql', 'hdfs']",Proceedings of the ACM Symposium on Cloud Computing,True,['failure-management'],,,,,,10.1145/2670979.2670987,Conference Paper,"['Reliability', 'Performance', 'Testing and Debugging']",
504,Mapping Virtual Machines onto Physical Machines in Cloud Computing: A Survey,"['Ilia Pietri', 'Rizos Sakellariou']",2016,"Cloud computing enables users to provision resources on demand and execute applications in a way that meets their requirements by choosing virtual resources that fit their application resource needs. Then, it becomes the task of cloud resource providers to accommodate these virtual resources onto physical resources. This problem is a fundamental challenge in cloud computing as resource providers need to map virtual resources onto physical resources in a way that takes into account the providers’ optimization objectives. This article surveys the relevant body of literature that deals with this mapping problem and how it can be addressed in different scenarios and through different objectives and optimization techniques. The evaluation aspects of different solutions are also considered. The article aims at both identifying and classifying research done in the area adopting a categorization that can enhance understanding of the problem.",https://doi.org/10.1145/2983575,True,"[{'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 7}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 29}, {'database': 'ACM', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 24}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 10}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 41}]",68.0,,,,['scheduling'],,ACM Computing Surveys (CSUR),True,['resource-provisioning'],True,True,,,,10.1145/2983575,Journal Article,"['cloud computing', 'VM placement', 'VM scheduling', 'VM configuration']",
505,Brownout Approach for Adaptive Management of Resources and Applications in Cloud Computing Systems: A Taxonomy and Future Directions,"['Minxian Xu', 'Rajkumar Buyya']",2019,"Cloud computing has been regarded as an emerging approach to provisioning resources and managing applications. It provides attractive features, such as an on-demand model, scalability enhancement, and management cost reduction. However, cloud computing systems continue to face problems such as hardware failures, overloads caused by unexpected workloads, or the waste of energy due to inefficient resource utilization, which all result in resource shortages and application issues such as delays or saturation. A paradigm, the brownout, has been applied to handle these issues by adaptively activating or deactivating optional parts of applications or services to manage resource usage in cloud computing system. Brownout has successfully shown that it can avoid overloads due to changes in workload and achieve better load balancing and energy saving effects. This article proposes a taxonomy of the brownout approach for managing resources and applications adaptively in cloud computing systems and carries out a comprehensive survey. It identifies open challenges and offers future research directions.",https://doi.org/10.1145/3234151,True,"[{'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 20}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 87}, {'database': 'ACM', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 59}]",10.0,,,['survey'],['resource-consolidation'],,ACM Computing Surveys (CSUR),True,['resource-provisioning'],True,True,,,,10.1145/3234151,Journal Article,"['optional services', 'brownout', 'adaptive management', 'Cloud computing', 'quality of service']",
506,Self-adaptive provisioning of virtualized resources in cloud computing,"['Jia Rao', 'Xiangping Bu', 'Kun Wang', 'Cheng-Zhong Xu']",2011,"In this paper, we propose a distributed learning mechanism that facilitates self-adaptive virtual machines resource provisioning. We treat cloud resource allocation as a distributed learning task, in which each VM being a highly autonomous agent submits resource requests according to its own benefit. The mechanism evaluates the requests and replies with feedback. We develop a reinforcement learning algorithm with a highly efficient representation of experiences as the heart of the VM side learning engine. We prototype the mechanism and the distributed learning algorithm in an iBalloon system. Experiment results on a Xen-based cloud testbed demonstrate the effectiveness of iBalloon.",https://doi.org/10.1145/2007116.2007157,True,"[{'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 31}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 32}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 40}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 42}]",20.0,,,,,,Proceedings of the ACM SIGMETRICS joint international conference on Measurement and modeling of computer systems,True,['resource-provisioning'],,True,,,,10.1145/2007116.2007157,Journal Article,"['cloud management', 'autonomic computing', 'reinforcement learning']",
507,Auto-Scaling Web Applications in Clouds: A Taxonomy and Survey,"['Chenhao Qu', 'Rodrigo N. Calheiros', 'Rajkumar Buyya']",2018,"Web application providers have been migrating their applications to cloud data centers, attracted by the emerging cloud computing paradigm. One of the appealing features of the cloud is elasticity. It allows cloud users to acquire or release computing resources on-demand, which enables web application providers to automatically scale the resources provisioned to their applications without human intervention under a dynamic workload to minimize resource cost while satisfying Quality of Service (QoS) requirements. In this paper, we comprehensively analyze the challenges that remain in auto-scaling web applications in clouds and review the developments in this field. We present a taxonomy of auto-scalers according to the identified challenges and key properties. We analyze the surveyed works and map them to the taxonomy to identify the weaknesses in this field. Moreover, based on the analysis, we propose new future directions that can be explored in this area.",https://doi.org/10.1145/3148149,True,"[{'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 66}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 52}, {'database': 'ACM', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 62}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 11}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 60}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 526}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud')"", 'index': 591}]",86.0,,,['survey'],['resource-consolidation'],,ACM Computing Surveys (CSUR),True,['resource-provisioning'],True,True,,,,10.1145/3148149,Journal Article,"['Auto-scaling', 'cloud computing', 'web application']",
508,UBL: unsupervised behavior learning for predicting performance anomalies in virtualized cloud systems,"['Daniel Joseph Dean', 'Hiep Nguyen', 'Xiaohui Gu']",2012,"Infrastructure-as-a-Service (IaaS) clouds are prone to performance anomalies due to their complex nature. Although previous work has shown the effectiveness of using statistical learning to detect performance anomalies, existing schemes often assume labelled training data, which requires significant human effort and can only handle previously known anomalies. We present an Unsupervised Behavior Learning (UBL) system for IaaS cloud computing infrastructures. UBL leverages Self-Organizing Maps to capture emergent system behaviors and predict unknown anomalies. For scalability, UBL uses residual resources in the cloud infrastructure for behavior learning and anomaly prediction with little add-on cost. We have implemented a prototype of the UBL system on top of the Xen platform and conducted extensive experiments using a range of distributed systems. Our results show that UBL can predict performance anomalies with high accuracy and achieve sufficient lead time for automatic anomaly prevention. UBL supports large-scale infrastructure-wide behavior learning with negligible overhead.",https://doi.org/10.1145/2371536.2371572,True,"[{'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 76}]",55.0,,,,['failure-detection'],,Proceedings of the 9th international conference on Autonomic computing,True,['failure-management'],,,['anomaly-detection'],,,10.1145/2371536.2371572,Conference Paper,"['anomaly prediction', 'unsupervised system behavior learning', 'cloud computing']",
509,SDN Flow Entry Management Using Reinforcement Learning,"['Ting-Yu Mu', 'Ala Al-Fuqaha', 'Khaled Shuaib', 'Farag M. Sallabi', 'Junaid Qadir']",2018,"Modern information technology services largely depend on cloud infrastructures to provide their services. These cloud infrastructures are built on top of datacenter networks (DCNs) constructed with high-speed links, fast switching gear, and redundancy to offer better flexibility and resiliency. In this environment, network traffic includes long-lived (elephant) and short-lived (mice) flows with partitioned and aggregated traffic patterns. Although SDN-based approaches can efficiently allocate networking resources for such flows, the overhead due to network reconfiguration can be significant. With limited capacity of Ternary Content-Addressable Memory (TCAM) deployed in an OpenFlow enabled switch, it is crucial to determine which forwarding rules should remain in the flow table, and which rules should be processed by the SDN controller in case of a table-miss on the SDN switch. This is needed in order to obtain the flow entries that satisfy the goal of reducing the long-term control plane overhead introduced between the controller and the switches. To achieve this goal, we propose a machine learning technique that utilizes two variations of reinforcement learning (RL) algorithms-the first of which is traditional reinforcement learning algorithm based while the other is deep reinforcement learning based. Emulation results using the RL algorithm show around 60% improvement in reducing the long-term control plane overhead, and around 14% improvement in the table-hit ratio compared to the Multiple Bloom Filters (MBF) method given a fixed size flow table of 4KB.",https://doi.org/10.1145/3281032,True,"[{'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 29}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 34}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 75}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 308}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 3}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 65}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 1}]",3.0,,,,,,ACM Transactions on Autonomous and Adaptive Systems (TAAS),True,['resource-provisioning'],,True,,,,10.1145/3281032,Journal Article,"['Openflow', 'big data', 'ternary content addressable memory', 'elephant and mice flows', 'Flow entry', 'reinforcement learning', 'machine learning', 'MiniNet', 'software defined networking (SDN)']",
510,Resource and performance prediction at high utilization for N-Tier cloud-based service systems,"['Wenbin Zhang', 'Yuliang Shi', 'Yongqing Zheng', 'Lei Liu', 'Lizhen Cui']",2017,"One of the key objectives of cloud computing systems is to meet the service level agreements (SLAs) under conditions of high resource utilization. Cloud service providers often need to design policies for resource sharing and performance optimization. As a result, being able to predict the performance and resource utilizations prior to implementing these policies is important to the dynamic provisioning of services by cloud providers. It is a significant and difficult challenge due to the fact that requests for resources often interact with each other in complex ways. Moreover, the dynamics of the cloud environment bring more problems to predicting the performance of a running query or workload. Hence, an accurate situation-aware model which can capture the complex interactions among resource requests is useful for addressing this challenge. To this end, we propose an efficient and highly accurate resource and performance prediction framework which takes into account the interactions among concurrently running resource requests for n-tier service systems. The proposed framework extends the Gaussian process and kernel canonical correlation analysis techniques and is able to dynamically adapt to variations in workload and physical resource usage. The proposed framework has been trained and evaluated extensively with a realistic multi-tier cloud application benchmark - the RUBiS benchmark system. The results demonstrate that the framework yields highly accurate performance and resource usage predictions, especially under high resource utilization conditions.",https://doi.org/10.1145/3014812.3014857,True,"[{'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 57}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 95}, {'database': 'ACM', 'search_string': ""'regression' AND ('cloud')"", 'index': 72}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 28}]",1.0,,,,,,Proceedings of the Australasian Computer Science Week Multiconference,True,['resource-provisioning'],,True,,,,10.1145/3014812.3014857,Conference Paper,"['cloud', 'performance', 'prediction', 'resource', 'benchmark']",
511,Tracking adaptive performance models using dynamic clustering of user classes,"['Hamoun Ghanbari', 'Cornel Barna', 'Marin Litoiu', 'Murray Woodside', 'Tao Zheng', 'Johnny Wong', 'Gabriel Iszlai']",2011,"Estimation techniques have been largely applied to track hidden performance parameters (e.g. service demands) of web based software systems. In this paper we investigate dynamic multiclass modeling of such systems, with variable classes of service, aiming at finding a low complexity model yet with enough accuracy. We propose a combination of clustering algorithm and tracking filter for effective grouping of classes of services. The tracking estimator is based on a layered queuing model with parameters for CPU demands and the user load intensity of each class of service. Clustering uses the K-means algorithm. The target application is autonomic control of web clusters, where changes occur at different rates and amplitudes and at random time instants. Experiments show that the tracking is effective, and reveal good filter settings for different variations.",https://doi.org/10.1145/1958746.1958774,True,"[{'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 70}]",12.0,,,,,,Proceedings of the 2nd ACM/SPEC International Conference on Performance engineering,True,['resource-provisioning'],,,,,,10.1145/1958746.1958774,Conference Paper,,
512,DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep Learning,"['Min Du', 'Feifei Li', 'Guineng Zheng', 'Vivek Srikumar']",2017,"Anomaly detection is a critical step towards building a secure and trustworthy system. The primary purpose of a system log is to record system states and significant events at various critical points to help debug system failures and perform root cause analysis. Such log data is universally available in nearly all computer systems. Log data is an important and valuable resource for understanding system status and performance issues; therefore, the various system logs are naturally excellent source of information for online monitoring and anomaly detection. We propose DeepLog, a deep neural network model utilizing Long Short-Term Memory (LSTM), to model a system log as a natural language sequence. This allows DeepLog to automatically learn log patterns from normal execution, and detect anomalies when log patterns deviate from the model trained from log data under normal execution. In addition, we demonstrate how to incrementally update the DeepLog model in an online fashion so that it can adapt to new log patterns over time. Furthermore, DeepLog constructs workflows from the underlying system log so that once an anomaly is detected, users can diagnose the detected anomaly and perform root cause analysis effectively. Extensive experimental evaluations over large log data have shown that DeepLog has outperformed other existing log-based anomaly detection methods based on traditional data mining methodologies.",https://doi.org/10.1145/3133956.3134015,True,"[{'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 94}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 73}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 6}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 67}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 28}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 34}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 13}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 18}]",223.0,['rnn'],['logs'],,['failure-detection'],,Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security,True,['failure-management'],True,,['anomaly-detection'],True,61.0,10.1145/3133956.3134015,Conference Paper,"['log data analysis', 'deep learning', 'anomaly detection']",['online']
513,Towards highly reliable enterprise network services via inference of multi-level dependencies,"['Paramvir Bahl', 'Ranveer Chandra', 'Albert Greenberg', 'Srikanth Kandula', 'David A. Maltz', 'Ming Zhang']",2007,"Localizing the sources of performance problems in large enterprise networks is extremely challenging. Dependencies are numerous, complex and inherently multi-level, spanning hardware and software components across the network and the computing infrastructure. To exploit these dependencies for fast, accurate problem localization, we introduce an Inference Graph model, which is well-adapted to user-perceptible problems rooted in conditions giving rise to both partial service degradation and hard faults. Further, we introduce the Sherlock system to discover Inference Graphs in the operational enterprise, infer critical attributes, and then leverage the result to automatically detect and localize problems. To illuminate strengths and limitations of the approach, we provide results from a prototype deployment in a large enterprise network, as well as from testbed emulations and simulations. In particular, we find that taking into account multi-level structure leads to a 30% improvement in fault localization, as compared to two-level approaches.",https://doi.org/10.1145/1282380.1282383,True,"[{'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault localization' OR 'failure localization')"", 'index': 23}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 45}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 91}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 85}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 65}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 24}]",448.0,['graph-mining'],"['traces', 'packets']",['new-method'],['root-cause-analysis'],"['server', 'network', 'link', 'router', 'service']","Proceedings of the 2007 conference on Applications, technologies, architectures, and protocols for computer communications",True,['failure-management'],True,True,['fault-localization'],True,71.0,10.1145/1282380.1282383,Conference Paper,"['dependencies', 'probabilistic inference', 'network and service management', 'fault localization']",
514,Dynamic syslog mining for network failure monitoring,"['Kenji Yamanishi', 'Yuko Maruyama']",2005,"Syslog monitoring technologies have recently received vast attentions in the areas of network management and network monitoring. They are used to address a wide range of important issues including network failure symptom detection and event correlation discovery. Syslogs are intrinsically dynamic in the sense that they form a time series and that their behavior may change over time. This paper proposes a new methodology of dynamic syslog mining in order to detect failure symptoms with higher confidence and to discover sequential alarm patterns among computer devices. The key ideas of dynamic syslog mining are 1) to represent syslog behavior using a mixture of Hidden Markov Models, 2) to adaptively learn the model using an on-line discounting learning algorithm in combination with dynamic selection of the optimal number of mixture components, and 3) to give anomaly scores using universal test statistics with a dynamically optimized threshold. Using real syslog data we demonstrate the validity of our methodology in the scenarios of failure symptom detection, emerging pattern identification, and correlation discovery.",https://doi.org/10.1145/1081870.1081927,True,"[{'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault localization' OR 'failure localization')"", 'index': 51}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 32}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault localization' OR 'failure localization')"", 'index': 16}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 25}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 12}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 61}]",211.0,['markov-model'],['logs'],['novel-use'],['failure-detection'],,Proceedings of the eleventh ACM SIGKDD international conference on Knowledge discovery in data mining,True,['failure-management'],,,['anomaly-detection'],,,10.1145/1081870.1081927,Conference Paper,"['model selection', 'correlation analysis', 'syslog mining', 'probabilistic modeling', 'failure detection']",['online']
515,Fault prediction under the microscope: a closer look into HPC systems,"['Ana Gainaru', 'Franck Cappello', 'Marc Snir', 'William Kramer']",2012,"A large percentage of computing capacity in today's large high-performance computing systems is wasted because of failures. Consequently current research is focusing on providing fault tolerance strategies that aim to minimize fault's effects on applications. By far the most popular technique is the checkpointrestart strategy. A complement to this classical approach is failure avoidance, by which the occurrence of a fault is predicted and preventive measures are taken. This requires a reliable prediction system to anticipate failures and their locations. Thus far, research in this field has used ideal predictors that were not implemented in real HPC systems. In this paper, we merge signal analysis concepts with data mining techniques to extend the ELSA (Event Log Signal Analyzer) toolkit and offer an adaptive and more efficient prediction module. Our goal is to provide models that characterize the normal behavior of a system and the way faults affect it. Being able to detect deviations from normality quickly is the foundation of accurate fault prediction. However, this is challenging because component failure dynamics are heterogeneous in space and time. To this end, a large part of the paper is focused on a detailed analysis of the prediction method, by applying it to two large-scale systems and by investigating the characteristics and bottlenecks of each step of the prediction process. Furthermore, we analyze the prediction's precision and recall impact on current checkpointing strategies and highlight future improvements and directions for research in this field.",https://ieeexplore.ieee.org/document/6468487,True,"[{'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 56}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 36}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 22}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault detection' OR 'failure detection')"", 'index': 95}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault prediction' OR 'failure prediction')"", 'index': 14}]",125.0,"['correlation', 'rule-mining']",['logs'],,"['root-cause-analysis', 'failure-detection', 'failure-prevention']",['hpc'],"Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis",True,['failure-management'],False,,"['anomaly-detection', 'fault-localization', 'system-failure-prediction']",,,,Conference Paper,"['signal analysis', 'fault tolerance', 'large-scale HPC systems', 'fault detection']",['online']
516,ActiveSLA: a profit-oriented admission control framework for database-as-a-service providers,"['Pengcheng Xiong', 'Yun Chi', 'Shenghuo Zhu', 'Junichi Tatemura', 'Calton Pu', 'Hakan HacigümüŞ']",2011,"The system overload is a common problem in a Database-as-a-Serice (DaaS) environment because of unpredictable and bursty workloads from various clients. Due to the service delivery nature of DaaS, such system overload usually has direct economic impact on the service provider, who has to pay penalties if the system performance does not meet clients' service level agreements (SLAs). In this paper, we investigate techniques that prevent system overload by using admission control. We propose a profit-oriented admission control framework, called ActiveSLA, for DaaS providers. ActiveSLA is an end-to-end framework that consists of two components. First, a prediction module estimates the probability for a new query to finish the execution before its deadline. Second, based on the predicted probability, a decision module determines whether or not to admit the given query into the database system. The decision is made with the profit optimization objective, where the expected profit is derived from the service level agreements between a service provider and its clients. We present extensive real system experiments with standard database benchmarks, under different traffic patterns, DBMS settings, and SLAs. The results demonstrate that ActiveSLA is able to make admission control decisions that are both more accurate and more profit-effective than several state-of-the-art methods.",https://doi.org/10.1145/2038916.2038931,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 76}]",41.0,,,,['resource-consolidation'],,Proceedings of the 2nd ACM Symposium on Cloud Computing,True,['resource-provisioning'],,,,,,10.1145/2038916.2038931,Conference Paper,"['cloud computing', 'machine learning', 'admission control', 'database-as-a-service']",
517,A unified model for holistic power usage in cloud datacenter servers,"['Peter Garraghan', 'Yaser Al-Anii', 'Jon Summers', 'Harvey Thompson', 'Nik Kapur', 'Karim Djemame']",2016,"Cloud datacenters are compute facilities formed by hundreds and thousands of heterogeneous servers requiring significant power requirements to operate effectively. Servers are composed by multiple interacting sub-systems including applications, microelectronic processors, and cooling which reflect their respective power profiles via different parameters. What is presently unknown is how to accurately model the holistic power usage of the entire server when including all these sub-systems together. This becomes increasingly challenging when considering diverse utilization patterns, server hardware characteristics, air and liquid cooling techniques, and importantly quantifying the non-electrical energy cost imposed by cooling operation. Such a challenge arises due to the need for multi-disciplinary expertise required to study server operation holistically. This work provides a unified model for capturing holistic power usage within Cloud datacenter servers. Constructed through controlled laboratory experiments, the model captures the relationship of server power usage between software, hardware, and cooling agnostic of architecture and cooling type (air and liquid). An exciting prospect is the ability to quantify the amount of non-electrical power consumed through cooling, allowing for more realistic and accurate server power profiles. This work represents the first empirically supported analysis and modeling of holistic power usage for Cloud datacenter servers, and bridges a significant gap between computer science and mechanical engineering research. Model validation through experiments demonstrates an average standard error of 3% for server power usage within both air and liquid cooled environments.",https://doi.org/10.1145/2996890.2996896,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 88}, {'database': 'ACM', 'search_string': ""'regression' AND ('cloud')"", 'index': 43}]",5.0,,,,,,Proceedings of the 9th International Conference on Utility and Cloud Computing,True,['resource-provisioning'],,,,,,10.1145/2996890.2996896,Conference Paper,"['cloud datacenters', 'holistic energy', 'server power modeling']",
518,A Survey and Taxonomy of Self-Aware and Self-Adaptive Cloud Autoscaling Systems,"['Tao Chen', 'Rami Bahsoon', 'Xin Yao']",2016,"Autoscaling system can reconfigure cloud-based services and applications, through various configurations of cloud software and provisions of hardware resources, to adapt to the changing environment at runtime. Such a behavior offers the foundation for achieving elasticity in modern cloud computing paradigm. Given the dynamic and uncertain nature of the shared cloud infrastructure, cloud autoscaling system has been engineered as one of the most complex, sophisticated and intelligent artifacts created by human, aiming to achieve self-aware, self-adaptive and dependable runtime scaling. Yet, existing Self-aware and Self-adaptive Cloud Autoscaling System (SSCAS) is not mature to a state that it can be reliably exploited in the cloud. In this article, we survey the state-of-the-art research studies on SSCAS and provide a comprehensive taxonomy for this field. We present detailed analysis of the results and provide insights on open challenges, as well as the promising directions that are worth investigated in the future work of this area of research. Our survey and taxonomy contribute to the fundamentals of engineering more intelligent autoscaling systems in the cloud.",https://doi.org/10.1145/3190507,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('cloud computing')"", 'index': 99}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 44}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 97}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 672}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud')"", 'index': 325}]",29.0,,,['survey'],['resource-consolidation'],,ACM Computing Surveys (CSUR),True,['resource-provisioning'],True,True,,,,10.1145/3190507,Journal Article,"['resources provisioning', 'distributed systems', 'Cloud computing', 'self-aware systems', 'auto-scaling', 'self-adaptive systems']",
519,Cross-project defect prediction: a large scale experiment on data vs. domain vs. process,"['Thomas Zimmermann', 'Nachiappan Nagappan', 'Harald Gall', 'Emanuel Giger', 'Brendan Murphy']",2009,"Prediction of software defects works well within projects as long as there is a sufficient amount of data available to train any models. However, this is rarely the case for new software projects and for many companies. So far, only a few have studies focused on transferring prediction models from one project to another. In this paper, we study cross-project defect prediction models on a large scale. For 12 real-world applications, we ran 622 cross-project predictions. Our results indicate that cross-project prediction is a serious challenge, i.e., simply using models from projects in the same domain or with the same process does not lead to accurate predictions. To help software engineers choose models wisely, we identified factors that do influence the success of cross-project predictions. We also derived decision trees that can provide early estimates for precision, recall, and accuracy before a prediction is attempted.",https://doi.org/10.1145/1595696.1595713,True,"[{'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault detection' OR 'failure detection')"", 'index': 12}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 37}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 64}]",261.0,['decision-tree'],,"['new-method', 'novel-use']",['failure-prevention'],['software'],Proceedings of the 7th joint meeting of the European software engineering conference and the ACM SIGSOFT symposium on The foundations of software engineering,True,['failure-management'],,True,['software-defect-prediction'],,,10.1145/1595696.1595713,Conference Paper,"['prediction quality', 'cross-project', 'defect prediction', 'logistic regression', 'churn', 'decision trees']",
520,Exploring program phases for statistical bug localization,"['Varun Modi', 'Subhajit Roy', 'Sanjeev K. Aggarwal']",2013,"Statistical bug isolation techniques attempt to capture a correlation of various program features (like predicates and profiled paths) for debugging. These techniques collect profile data for multiple executions, both with successful and faulty runs, and propose using various statistical tests to capture this correlation. In this paper, we explore the utility of program phases, a concept which is primarily used by computer architects to speed up architectural simulations, for statistical bug isolation. Program phases represent sets of execution intervals in a program's execution where the rates of architectural statistics like branch mispredictions, CPU/Memory usage and cache misses remain almost the same. We found multiple scenarios where coupling program phases with predicates achieves higher accuracy to bug localization than when predicates are used alone. We demonstrate the use of program phases for bug isolation by presenting experimental results and concrete case studies on medium-size programs, showing an improved ranking of the program points that are critical to debugging over when program phases are not used.",https://doi.org/10.1145/2462029.2462034,True,"[{'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault detection' OR 'failure detection')"", 'index': 52}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault localization' OR 'failure localization')"", 'index': 77}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault localization' OR 'failure localization')"", 'index': 12}]",4.0,,,,['root-cause-analysis'],,Proceedings of the 11th ACM SIGPLAN-SIGSOFT Workshop on Program Analysis for Software Tools and Engineering,True,['failure-management'],,,,,,10.1145/2462029.2462034,Conference Paper,"['program phases', 'automated fault isolation', 'statistical debugging']",
521,Applying classification techniques to remotely-collected program execution data,"['Murali Haran', 'Alan Karr', 'Alessandro Orso', 'Adam Porter', 'Ashish Sanil']",2005,"There is an increasing interest in techniques that support measurement and analysis of fielded software systems. One of the main goals of these techniques is to better understand how software actually behaves in the field. In particular, many of these techniques require a way to distinguish, in the field, failing from passing executions. So far, researchers and practitioners have only partially addressed this problem: they have simply assumed that program failure status is either obvious (i.e., the program crashes) or provided by an external source (e.g., the users). In this paper, we propose a technique for automatically classifying execution data, collected in the field, as coming from either passing or failing program runs. (Failing program runs are executions that terminate with a failure, such as a wrong outcome.) We use statistical learning algorithms to build the classification models. Our approach builds the models by analyzing executions performed in a controlled environment (e.g., test cases run in-house) and then uses the models to predict whether execution data produced by a fielded instance were generated by a passing or failing program execution. We also present results from an initial feasibility study, based on multiple versions of a software subject, in which we investigate several issues vital to the applicability of the technique. Finally, we present some lessons learned regarding the interplay between the reliability of classification models and the amount and type of data collected.",https://doi.org/10.1145/1081706.1081732,True,"[{'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault detection' OR 'failure detection')"", 'index': 67}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault localization' OR 'failure localization')"", 'index': 26}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 85}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault localization' OR 'failure localization')"", 'index': 64}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault localization' OR 'failure localization')"", 'index': 59}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault localization' OR 'failure localization')"", 'index': 29}]",37.0,,,,['failure-detection'],,Proceedings of the 10th European software engineering conference held jointly with 13th ACM SIGSOFT international symposium on Foundations of software engineering,True,['failure-management'],,,,,,10.1145/1081706.1081732,Conference Paper,"['software behavior', 'classification', 'machine learning']",
522,Active learning for automatic classification of software behavior,"['James F. Bowring', 'James M. Rehg', 'Mary Jean Harrold']",2004,"A program's behavior is ultimately the collection of all its executions. This collection is diverse, unpredictable, and generally unbounded. Thus it is especially suited to statistical analysis and machine learning techniques. The primary focus of this paper is on the automatic classification of program behavior using execution data. Prior work on classifiers for software engineering adopts a classical batch-learning approach. In contrast, we explore an active-learning paradigm for behavior classification. In active learning, the classifier is trained incrementally on a series of labeled data elements. Secondly, we explore the thesis that certain features of program behavior are stochastic processes that exhibit the Markov property, and that the resultant Markov models of individual program executions can be automatically clustered into effective predictors of program behavior. We present a technique that models program executions as Markov models, and a clustering method for Markov models that aggregates multiple program executions into effective behavior classifiers. We evaluate an application of active learning to the efficient refinement of our classifiers by conducting three empirical studies that explore a scenario illustrating automated test plan augmentation.",https://doi.org/10.1145/1007512.1007539,True,"[{'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault detection' OR 'failure detection')"", 'index': 72}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault localization' OR 'failure localization')"", 'index': 18}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 70}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault localization' OR 'failure localization')"", 'index': 70}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 49}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault localization' OR 'failure localization')"", 'index': 30}]",127.0,"['clustering', 'markov-model']",['runs'],['new-method'],['failure-detection'],['software'],Proceedings of the 2004 ACM SIGSOFT international symposium on Software testing and analysis,True,['failure-management'],,,['anomaly-detection'],,,10.1145/1007512.1007539,Conference Paper,"['software behavior', 'software testing', 'Markov models', 'machine learning']",['online']
523,Generic and Scalable Framework for Automated Time-series Anomaly Detection,"['Nikolay Laptev', 'Saeed Amizadeh', 'Ian Flint']",2015,"This paper introduces a generic and scalable framework for automated anomaly detection on large scale time-series data. Early detection of anomalies plays a key role in maintaining consistency of person's data and protects corporations against malicious attackers. Current state of the art anomaly detection approaches suffer from scalability, use-case restrictions, difficulty of use and a large number of false positives. Our system at Yahoo, EGADS, uses a collection of anomaly detection and forecasting models with an anomaly filtering layer for accurate and scalable anomaly detection on time-series. We compare our approach against other anomaly detection systems on real and synthetic data with varying time-series characteristics. We found that our framework allows for 50-60% improvement in precision and recall for a variety of use-cases. Both the data and the framework are being open-sourced. The open-sourcing of the data, in particular, represents the first of its kind effort to establish the standard benchmark for anomaly detection.",https://doi.org/10.1145/2783258.2788611,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 24}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 48}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('anomaly detection' OR 'outlier detection')"", 'index': 29}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 46}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 35}, {'database': 'ACM', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 76}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault detection' OR 'failure detection')"", 'index': 59}]",150.0,"['clustering', 'similarity-matching']",,['novel-use'],['failure-detection'],[],Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,True,['failure-management'],False,,['anomaly-detection'],,,10.1145/2783258.2788611,Conference Paper,"['time-series', 'scalable anomaly detection', 'anomaly detection']",
524,On-line anomaly detection of deployed software: a statistical machine learning approach,"['George K. Baah', 'Alexander Gray', 'Mary Jean Harrold']",2006,"This paper presents a new machine-learning technique that performs anomaly detection as software is executing in the field. The technique uses a fully observable Markov model where each state in the model emits a number of distinct observations according to a probability distribution, and estimates the model parameters using the Baum-Welch algorithm. The trained model is then deployed with the software to perform anomaly detection. By performing the anomaly detection as the software is executing, faults associated with anomalies can be located and fixed before they cause critical failures in the system, and developers time to debug deployed software can be reduced. This paper also presents a prototype implementation of our technique, along with a case study that shows, for the subjects we studied, the effectiveness of the technique.",https://doi.org/10.1145/1188895.1188911,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 33}, {'database': 'ACM', 'search_string': ""'regression' AND ('fault localization' OR 'failure localization')"", 'index': 59}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('fault localization' OR 'failure localization')"", 'index': 3}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 14}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 36}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault localization' OR 'failure localization')"", 'index': 43}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault localization' OR 'failure localization')"", 'index': 39}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('anomaly detection' OR 'outlier detection')"", 'index': 33}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 10}, {'database': 'ACM', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 96}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault localization' OR 'failure localization')"", 'index': 6}]",24.0,,,,['failure-detection'],,Proceedings of the 3rd international workshop on Software quality assurance,True,['failure-management'],,True,,,,10.1145/1188895.1188911,Conference Paper,"['anomaly diagnosis', 'fault localization', 'machine learning', 'anomaly detection', 'Markov models']",
525,Automated anomaly detection and performance modeling of enterprise applications,"['Ludmila Cherkasova', 'Kivanc Ozonat', 'Ningfang Mi', 'Julie Symons', 'Evgenia Smirni']",2009,"Automated tools for understanding application behavior and its changes during the application lifecycle are essential for many performance analysis and debugging tasks. Application performance issues have an immediate impact on customer experience and satisfaction. A sudden slowdown of enterprise-wide application can effect a large population of customers, lead to delayed projects, and ultimately can result in company financial loss. Significantly shortened time between new software releases further exacerbates the problem of thoroughly evaluating the performance of an updated application. Our thesis is that online performance modeling should be a part of routine application monitoring. Early, informative warnings on significant changes in application performance should help service providers to timely identify and prevent performance problems and their negative impact on the service. We propose a novel framework for automated anomaly detection and application change analysis. It is based on integration of two complementary techniques: (i) a regression-based transaction model that reflects a resource consumption model of the application, and (ii) an application performance signature that provides a compact model of runtime behavior of the application. The proposed integrated framework provides a simple and powerful solution for anomaly detection and analysis of essential performance changes in application behavior. An additional benefit of the proposed approach is its simplicity: It is not intrusive and is based on monitoring data that is typically available in enterprise production environments. The introduced solution further enables the automation of capacity planning and resource provisioning tasks of multitier applications in rapidly evolving IT environments.",https://doi.org/10.1145/1629087.1629089,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 44}]",55.0,,,,['failure-detection'],,ACM Transactions on Computer Systems (TOCS),True,['failure-management'],,,,,,10.1145/1629087.1629089,Journal Article,"['performance modeling', 'capacity planning', 'online algorithms', 'Anomaly detection', 'multitier applications']",
526,Temporal sequence learning and data reduction for anomaly detection,"['Terran Lane', 'Carla E. Brodley']",1999,"The anomaly-detection problem can be formulated as one of learning to characterize the behaviors of an individual, system, or network in terms of temporal sequences of discrete data. We present an approach on the basis of instance-based learning (IBL) techniques. To cast the anomaly-detection task in an IBL framework, we employ an approach that transforms temporal sequences of discrete, unordered observations into a metric space via a similarity measure that encodes intra-attribute dependencies. Classification boundaries are selected from an a posteriori characterization of valid user behaviors, coupled with a domain heuristic. An empirical evaluation of the approach on user command data demonstrates that we can accurately differentiate the profiled user from alternative users when the available features encode sufficient information. Furthermore, we demonstrate that the system detects anomalous conditions quickly — an important quality for reducing potential damage by a malicious user. We present several techniques for reducing data storage requirements of the user profile, including instance-selection methods and clustering. As empirical evaluation shows that a new greedy clustering algorithm reduces the size of the user model by 70%, with only a small loss in accuracy.",https://doi.org/10.1145/322510.322526,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 46}, {'database': 'ACM', 'search_string': ""'regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 83}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('anomaly detection' OR 'outlier detection')"", 'index': 92}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 79}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('anomaly detection' OR 'outlier detection')"", 'index': 47}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('anomaly detection' OR 'outlier detection')"", 'index': 79}]",234.0,"['similarity-matching', 'clustering']",,,['failure-detection'],,Proceedings of the 5th ACM conference on Computer and communications security,True,['failure-management'],,,['anomaly-detection'],,,10.1145/322510.322526,Journal Article,"['clustering', 'data reduction', 'machine learning', 'empirical evaluation', 'anomaly detection', 'instance based learning', 'user profiling']",
527,Opprentice: Towards Practical and Automatic Anomaly Detection Through Machine Learning,"['Dapeng Liu', 'Youjian Zhao', 'Haowen Xu', 'Yongqian Sun', 'Dan Pei', 'Jiao Luo', 'Xiaowei Jing', 'Mei Feng']",2015,"Closely monitoring service performance and detecting anomalies are critical for Internet-based services. However, even though dozens of anomaly detectors have been proposed over the years, deploying them to a given service remains a great challenge, requiring manually and iteratively tuning detector parameters and thresholds. This paper tackles this challenge through a novel approach based on supervised machine learning. With our proposed system, Opprentice (Operators' apprentice), operators' only manual work is to periodically label the anomalies in the performance data with a convenient tool. Multiple existing detectors are applied to the performance data in parallel to extract anomaly features. Then the features and the labels are used to train a random forest classifier to automatically select the appropriate detector-parameter combinations and the thresholds. For three different service KPIs in a top global search engine, Opprentice can automatically satisfy or approximate a reasonable accuracy preference (recall >= 0.66 and precision>= 0.66). More importantly, Opprentice allows operators to label data in only tens of minutes, while operators traditionally have to spend more than ten days selecting and tuning detectors, which may still turn out not to work in the end.",https://doi.org/10.1145/2815675.2815679,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 96}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 27}]",68.0,"['random-forest', 'autoregression']",['kpis'],['novel-use'],['failure-detection'],,Proceedings of the 2015 Internet Measurement Conference,True,['failure-management'],True,True,['anomaly-detection'],True,58.0,10.1145/2815675.2815679,Conference Paper,"['tuning detectors', 'machine learning', 'anomaly detection']",
528,Fault-localization techniques for software systems: a literature review,"['Pragya Agarwal', 'Arun Prakash Agrawal']",2014,"Software is a major component of any computer system. To maintain the quality of software, early fault localization is necessary. Many different fault-localization methods have been used by researchers. Ideally, methods for fault-localization are used in such a way that one is able to detect as many faults as possible using the least resources. But, in general, it is hard to predict a test suite's fault-localization capability. This paper gives a review of the previous studies that are related to software fault-localization methods. It reviews various journal and conference papers on localization of faults and various methods for localization of faults proposed in the literature.",https://doi.org/10.1145/2659118.2659125,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('fault localization' OR 'failure localization')"", 'index': 6}, {'database': 'ACM', 'search_string': ""'clustering' AND ('fault localization' OR 'failure localization')"", 'index': 3}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 9}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault localization' OR 'failure localization')"", 'index': 4}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault localization' OR 'failure localization')"", 'index': 17}]",10.0,,,,['root-cause-analysis'],,ACM SIGSOFT Software Engineering Notes,True,['failure-management'],,True,,,,10.1145/2659118.2659125,Journal Article,"['fault localization', 'software systems', 'test-case prioritization methods', 'debugging']",
529,Mitigating the confounding effects of program dependences for effective fault localization,"['George K. Baah', 'Andy Podgurski', 'Mary Jean Harrold']",2011,"Dynamic program dependences are recognized as important factors in software debugging because they contribute to triggering the effects of faults and propagating the effects to a program's output. The effects of dynamic dependences also produce significant confounding bias when statistically estimating the causal effect of a statement on the occurrence of program failures, which leads to poor fault localization results. This paper presents a novel causal-inference technique for fault localization that accounts for the effects of dynamic data and control dependences and thus, significantly reduces confounding bias during fault localization. The technique employs a new dependence-based causal model together with matching of test executions based on their dynamic dependences. The paper also presents empirical results indicating that the new technique performs significantly better than existing statistical fault-localization techniques as well as our previous fault localization technique based on causal-inference methodology.",https://doi.org/10.1145/2025113.2025136,True,"[{'database': 'ACM', 'search_string': ""'regression' AND ('fault localization' OR 'failure localization')"", 'index': 20}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 35}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault localization' OR 'failure localization')"", 'index': 71}]",41.0,,,,['root-cause-analysis'],,Proceedings of the 19th ACM SIGSOFT symposium and the 13th European conference on Foundations of software engineering,True,['failure-management'],,,,,,10.1145/2025113.2025136,Conference Paper,"['matching', 'causal inference', 'debugging', 'fault localization', 'potential outcome model', 'program analysis']",
530,Feedback-Guided Anomaly Discovery via Online Optimization,"['Md Amran Siddiqui', 'Alan Fern', 'Thomas G. Dietterich', 'Ryan Wright', 'Alec Theriault', 'David W. Archer']",2018,"Anomaly detectors are often used to produce a ranked list of statistical anomalies, which are examined by human analysts in order to extract the actual anomalies of interest. This can be exceedingly difficult and time consuming when most high-ranking anomalies are false positives and not interesting from an application perspective. In this paper, we study how to reduce the analyst's effort by incorporating their feedback about whether the anomalies they investigate are of interest or not. In particular, the feedback will be used to adjust the anomaly ranking after every analyst interaction, ideally moving anomalies of interest closer to the top. Our main contribution is to formulate this problem within the framework of online convex optimization, which yields an efficient and extremely simple approach to incorporating feedback compared to the prior state-of-the-art. We instantiate this approach for the powerful class of tree-based anomaly detectors and conduct experiments on a range of benchmark datasets. The results demonstrate the utility of incorporating feedback and advantages of our approach over the state-of-the-art. In addition, we present results on a significant cybersecurity application where the goal is to detect red-team attacks in real system audit data. We show that our approach for incorporating feedback is able to significantly reduce the time required to identify malicious system entities across multiple attacks on multiple operating systems.",https://doi.org/10.1145/3219819.3220083,True,"[{'database': 'ACM', 'search_string': ""'logistic regression' AND ('anomaly detection' OR 'outlier detection')"", 'index': 31}]",4.0,,,,['failure-detection'],,Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining,True,['failure-management'],,,,,,10.1145/3219819.3220083,Conference Paper,"['online convex optimization', 'anomaly detection', 'anomaly detection feedback', 'feedback in linear model', 'anomaly detection on security']",
531,Robust Anomaly Detection for Multivariate Time Series through Stochastic Recurrent Neural Network,"['Ya Su', 'Youjian Zhao', 'Chenhao Niu', 'Rong Liu', 'Wei Sun', 'Dan Pei']",2019,"Industry devices (i.e., entities) such as server machines, spacecrafts, engines, etc., are typically monitored with multivariate time series, whose anomaly detection is critical for an entity's service quality management. However, due to the complex temporal dependence and stochasticity of multivariate time series, their anomaly detection remains a big challenge. This paper proposes OmniAnomaly, a stochastic recurrent neural network for multivariate time series anomaly detection that works well robustly for various devices. Its core idea is to capture the normal patterns of multivariate time series by learning their robust representations with key techniques such as stochastic variable connection and planar normalizing flow, reconstruct input data by the representations, and use the reconstruction probabilities to determine anomalies. Moreover, for a detected entity anomaly, OmniAnomaly can provide interpretations based on the reconstruction probabilities of its constituent univariate time series. The evaluation experiments are conducted on two public datasets from aerospace and a new server machine dataset (collected and released by us) from an Internet company. OmniAnomaly achieves an overall F1-Score of 0.86 in three real-world datasets, signicantly outperforming the best performing baseline method by 0.09. The interpretation accuracy for OmniAnomaly is up to 0.89.",https://doi.org/10.1145/3292500.3330672,True,"[{'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 65}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 34}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('anomaly detection' OR 'outlier detection')"", 'index': 40}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 62}]",4.0,['rnn'],"['host-metrics', 'network-metrics']","['comparison', 'novel-use']",['failure-detection'],['server'],Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining,True,['failure-management'],True,,['anomaly-detection'],,,10.1145/3292500.3330672,Conference Paper,"['multivariate time series', 'stochastic model', 'anomaly detection', 'recurrent neural network']",
532,AuTO: scaling deep reinforcement learning for datacenter-scale automatic traffic optimization,"['Li Chen', 'Justinas Lingys', 'Kai Chen', 'Feng Liu']",2018,"Traffic optimizations (TO, e.g. flow scheduling, load balancing) in datacenters are difficult online decision-making problems. Previously, they are done with heuristics relying on operators' understanding of the workload and environment. Designing and implementing proper TO algorithms thus take at least weeks. Encouraged by recent successes in applying deep reinforcement learning (DRL) techniques to solve complex online control problems, we study if DRL can be used for automatic TO without human-intervention. However, our experiments show that the latency of current DRL systems cannot handle flow-level TO at the scale of current datacenters, because short flows (which constitute the majority of traffic) are usually gone before decisions can be made. Leveraging the long-tail distribution of datacenter traffic, we develop a two-level DRL system, AuTO, mimicking the Peripheral & Central Nervous Systems in animals, to solve the scalability problem. Peripheral Systems (PS) reside on end-hosts, collect flow information, and make TO decisions locally with minimal delay for short flows. PS's decisions are informed by a Central System (CS), where global traffic information is aggregated and processed. CS further makes individual TO decisions for long flows. With CS&PS, AuTO is an end-to-end automatic TO system that can collect network information, learn from past decisions, and perform actions to achieve operator-defined goals. We implement AuTO with popular machine learning frameworks and commodity servers, and deploy it on a 32-server testbed. Compared to existing approaches, AuTO reduces the TO turn-around time from weeks to ~100 milliseconds while achieving superior performance. For example, it demonstrates up to 48.14% reduction in average flow completion time (FCT) over existing solutions.",https://doi.org/10.1145/3230543.3230551,True,"[{'database': 'ACM', 'search_string': ""'clustering' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 41}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 96}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 17}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 28}]",47.0,"['multilayer-perceptron', 'reinforcement-learning']",,,['resource-consolidation'],['network'],Proceedings of the 2018 Conference of the ACM Special Interest Group on Data Communication,True,['resource-provisioning'],,,,,,10.1145/3230543.3230551,Conference Paper,"['reinforcement learning', 'datacenter networks', 'traffic optimization']",
533,Multi-Tiered On-Demand Resource Scheduling for VM-Based Data Center,"['Ying Song', 'Hui Wang', 'Yaqiong Li', 'Binquan Feng', 'Yuzhong Sun']",2009,"The trend of using virtualization for server consolidation is more and more popular in enterprise data center. However, on-demand resource allocation among the concurrent hosted services in such a virtualized environment is still a challenge. In order to optimize resource allocation among services in data center, this paper proposes a multi-tiered resource scheduling scheme which automatically provides on-demand capacities to the hosted services via resources flowing among VMs. We model the resource flowing using optimiza-tion theory. Based on this model, we present a global re-source flowing algorithm in the multi-tiered resource scheduling scheme. This algorithm preferentially ensures performance of some critical services by degrading of others to some extent when resource competition arises. Using our RAINBOW prototype, we evaluate the multi-tiered resource scheduling scheme with the performance improvements for the most critical services up to 9%~16%, which are 75% of the maximum improvement margin, while performance degradation of others is up to 2%, and leads to 1%~5% im-provements in resource utilization than RAINBOW without resource flowing. Compared with the existent scheme, our work leads to 9% less improvements for critical services, while introduces 39% less degradation to low priority ser-vices.",https://doi.org/10.1109/CCGRID.2009.11,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 11}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 73}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 3}]",80.0,,,,['scheduling'],,Proceedings of the 2009 9th IEEE/ACM International Symposium on Cluster Computing and the Grid,True,['resource-provisioning'],,True,,,,10.1109/CCGRID.2009.11,Conference Paper,"['on-demand', 'Data Center', 'virtualization', 'resource scheduling']",
534,HOLMES: an event-driven solution to monitor data centers through continuous queries and machine learning,"['Pedro Henriques dos Santos Teixeira', 'Ricardo Gomes Clemente', 'Ronald Andreu Kaiser', 'Denis Almeida Vieira']",2010,"Supervisory processes are fundamental when running data center operations striving for fault resilience: any downtime can directly affect the business's income and definitely its reputation. Current monitoring tools rely on experts to configure constant thresholds on single streams, which is not appropriated for dynamic systems and insufficient to capture complex patterns. We present HOLMES, built to support data center experts to anticipate failures with a solution that combines Event Driven Architecture, Complex Event Processing and an unsupervised machine learning algorithm. Based on rules created by the users, the system continuously checks for known problems. Meanwhile, for the unknown ones, we leverage the CEP engine for aggregating and joining streams of real-time data to feed normalized input to FRAHST, our machine learning algorithm that detects anomalous patterns across multivariate numerical streams. We describe how the UI module also operates within the publish/subscribe paradigm to enhance situational awareness. The system had very well acceptance and was successfully implemented at one of the largest Internet Service Providers in South America.",https://doi.org/10.1145/1827418.1827461,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 58}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('IT operations')"", 'index': 3}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 31}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 65}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('IT operations')"", 'index': 4}]",5.0,,,,['failure-prediction'],,Proceedings of the Fourth ACM International Conference on Distributed Event-Based Systems,True,['failure-management'],,,,,,10.1145/1827418.1827461,Conference Paper,"['monitoring', 'anomaly detection', 'complex event processing', 'messaging-oriented middleware', 'data streams']",
535,Scalable near real-time failure localization of data center networks,"['Herodotos Herodotou', 'Bolin Ding', 'Shobana Balakrishnan', 'Geoff Outhred', 'Percy Fitter']",2014,"Large-scale data center networks are complex---comprising several thousand network devices and several hundred thousand links---and form the critical infrastructure upon which all higher-level services depend on. Despite the built-in redundancy in data center networks, performance issues and device or link failures in the network can lead to user-perceived service interruptions. Therefore, determining and localizing user-impacting availability and performance issues in the network in near real time is crucial. Traditionally, both passive and active monitoring approaches have been used for failure localization. However, data from passive monitoring is often too noisy and does not effectively capture silent or gray failures, whereas active monitoring is potent in detecting faults but limited in its ability to isolate the exact fault location depending on its scale and granularity. Our key idea is to use statistical data mining techniques on large-scale active monitoring data to determine a ranked list of suspect causes, which we refine with passive monitoring signals. In particular, we compute a failure probability for devices and links in near real time using data from active monitoring, and look for statistically significant increases in the failure probability. We also correlate the probabilistic output with other failure signals from passive monitoring to increase the confidence of the probabilistic analysis. We have implemented our approach in the Windows Azure production environment and have validated its effectiveness in terms of localization accuracy, precision, and time to localization using known network incidents over the past three months. The correlated ranked list of devices and links is surfaced as a report that is used by network operators to investigate current issues and identify probable root causes.",https://doi.org/10.1145/2623330.2623365,True,"[{'database': 'ACM', 'search_string': ""'classification' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 5}, {'database': 'ACM', 'search_string': ""'classification' AND ('fault localization' OR 'failure localization')"", 'index': 7}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 60}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault localization' OR 'failure localization')"", 'index': 33}]",9.0,['linear-regression'],,,['root-cause-analysis'],,Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining,True,['failure-management'],,,,,,10.1145/2623330.2623365,Conference Paper,"['failure localization', 'data center networks']",
536,A Supervised Learning Model for Identifying Inactive VMs in Private Cloud Data Centers,"['In Kee Kim', 'Sai Zeng', 'Christopher Young', 'Jinho Hwang', 'Marty Humphrey']",2016,"A private cloud has become an essential computing infrastructure for many enterprises. However, according to a recent study, 30% of VMs in data centers are not being used for any productive work. These ""inactive"" (or ""zombie"") VMs can arise from faulty VM management code within the cloud infrastructure but are usually the result of human neglect. Inactive VMs can hurt the performance of productive VMs, can distort internal cost management, and in the extreme can result in the cloud infrastructure being unable to allocate resources for new VMs. Correctly assessing the productivity of a VM can be challenging: e.g., is a VM that has low CPU utilization being used to slowly edit source code or is it an inactive VM that happens to be performing routine maintenance (e.g., virus-scan and software updates)? To address this problem, we develop a supervised learning model that leverages primitive information (e.g., running process, login history, network connections) of VMs periodically collected by a lightweight data collection framework. This model employs a linear support vector machine (SVM) approach that reflects single VM behavior as well as coordinated VM behaviors. We evaluated the identification accuracy of this model with a real-world dataset within IBM of more than 750 VMs. Results show that our model has a 20% higher accuracy (90%) than state-of-the-art approaches. An accurate model is an important first step to enable private cloud infrastructures to achieve better resource management through such actions as suspending or dynamically downsizing inactive VMs.",https://doi.org/10.1145/3007646.3007654,True,"[{'database': 'ACM', 'search_string': ""'classification' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 66}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud computing')"", 'index': 86}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 16}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 8}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 62}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 3}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 48}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 5}]",6.0,,,,['failure-detection'],,Proceedings of the Industrial Track of the 17th International Middleware Conference,True,['failure-management'],,,,,,10.1145/3007646.3007654,Conference Paper,"['Supervised Learning', 'Identifying Inactive VMs', 'IaaS Garbage Collection', 'Data Center Management', 'Cloud Computing']",
537,Fingerprinting the datacenter: automated classification of performance crises,"['Peter Bodik', 'Moises Goldszmidt', 'Armando Fox', 'Dawn B. Woodard', 'Hans Andersen']",2010,"Contemporary datacenters comprise hundreds or thousands of machines running applications requiring high availability and responsiveness. Although a performance crisis is easily detected by monitoring key end-to-end performance indicators (KPIs) such as response latency or request throughput, the variety of conditions that can lead to KPI degradation makes it difficult to select appropriate recovery actions. We propose and evaluate a methodology for automatic classification and identification of crises, and in particular for detecting whether a given crisis has been seen before, so that a known solution may be immediately applied. Our approach is based on a new and efficient representation of the datacenter's state called a fingerprint, constructed by statistical selection and summarization of the hundreds of performance metrics typically collected on such systems. Our evaluation uses 4 months of trouble-ticket data from a production datacenter with hundreds of machines running a 24x7 enterprise-class user-facing application. In experiments in a realistic and rigorous operational setting, our approach provides operators the information necessary to initiate recovery actions with 80% correctness in an average of 10 minutes, which is 50 minutes earlier than the deadline provided to us by the operators. To the best of our knowledge this is the first rigorous evaluation of any such approach on a large-scale production installation.",https://doi.org/10.1145/1755913.1755926,True,"[{'database': 'ACM', 'search_string': ""'classification' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 98}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 78}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 94}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('remediation' OR 'recovery')"", 'index': 64}, {'database': 'ACM', 'search_string': ""'classification' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 56}, {'database': 'ACM', 'search_string': ""'regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 92}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 13}, {'database': 'ACM', 'search_string': ""'logistic regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 5}]",208.0,['similarity-matching'],"['kpis', 'tickets']",,['root-cause-analysis'],,Proceedings of the 5th European conference on Computer systems,True,['failure-management'],True,True,['rca-others'],True,85.0,10.1145/1755913.1755926,Conference Paper,"['web applications', 'performance', 'datacenters']",
538,Taming coincidental correctness: Coverage refinement with context patterns to improve fault localization,"['Xinming Wang', 'S. C. Cheung', 'W. K. Chan', 'Zhenyu Zhang']",2009,"Recent techniques for fault localization leverage code coverage to address the high cost problem of debugging. These techniques exploit the correlations between program failures and the coverage of program entities as the clue in locating faults. Experimental evidence shows that the effectiveness of these techniques can be affected adversely by coincidental correctness, which occurs when a fault is executed but no failure is detected. In this paper, we propose an approach to address this problem. We refine code coverage of test runs using control- and data-flow patterns prescribed by different fault types. We conjecture that this extra information, which we call context patterns, can strengthen the correlations between program failures and the coverage of faulty program entities, making it easier for fault localization techniques to locate the faults. To evaluate the proposed approach, we have conducted a mutation analysis on three real world programs and cross-validated the results with real faults. The experimental results consistently show that coverage refinement is effective in easing the coincidental correctness problem in fault localization techniques.",https://doi.org/10.1109/ICSE.2009.5070507,True,"[{'database': 'ACM', 'search_string': ""'classification' AND ('fault localization' OR 'failure localization')"", 'index': 29}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 78}]",31.0,,,,['root-cause-analysis'],,Proceedings of the 31st International Conference on Software Engineering,True,['failure-management'],,,,,,10.1109/ICSE.2009.5070507,Conference Paper,,
539,Cloud Computing Resource Scheduling and a Survey of Its Evolutionary Approaches,"['Zhi-Hui Zhan', 'Xiao-Fang Liu', 'Yue-Jiao Gong', 'Jun Zhang', 'Henry Shu-Hung Chung', 'Yun Li']",2015,"A disruptive technology fundamentally transforming the way that computing services are delivered, cloud computing offers information and communication technology users a new dimension of convenience of resources, as services via the Internet. Because cloud provides a finite pool of virtualized on-demand resources, optimally scheduling them has become an essential and rewarding topic, where a trend of using Evolutionary Computation (EC) algorithms is emerging rapidly. Through analyzing the cloud computing architecture, this survey first presents taxonomy at two levels of scheduling cloud resources. It then paints a landscape of the scheduling problem and solutions. According to the taxonomy, a comprehensive survey of state-of-the-art approaches is presented systematically. Looking forward, challenges and potential future research directions are investigated and invited, including real-time scheduling, adaptive dynamic scheduling, large-scale scheduling, multiobjective scheduling, and distributed and parallel scheduling. At the dawn of Industry 4.0, cloud computing scheduling for cyber-physical integration with the presence of big data is also discussed. Research in this area is only in its infancy, but with the rapid fusion of information and data technology, more exciting and agenda-setting topics are likely to emerge on the horizon.",https://doi.org/10.1145/2788397,True,"[{'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 90}, {'database': 'ACM', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 60}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 46}, {'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 73}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 87}]",305.0,"['genetic-programming', 'ant-colony', 'particle-swarm']",,['survey'],['scheduling'],,ACM Computing Surveys (CSUR),True,['resource-provisioning'],True,False,,,,10.1145/2788397,Journal Article,"['evolutionary computation', 'Cloud computing', 'resource scheduling', 'particle swarm optimization', 'ant colony optimization', 'genetic algorithm']",
540,Characterizing cloud computing hardware reliability,"['Kashi Venkatesh Vishwanath', 'Nachiappan Nagappan']",2010,"Modern day datacenters host hundreds of thousands of servers that coordinate tasks in order to deliver highly available cloud computing services. These servers consist of multiple hard disks, memory modules, network cards, processors etc., each of which while carefully engineered are capable of failing. While the probability of seeing any such failure in the lifetime (typically 3-5 years in industry) of a server can be somewhat small, these numbers get magnified across all devices hosted in a datacenter. At such a large scale, hardware component failure is the norm rather than an exception. Hardware failure can lead to a degradation in performance to end-users and can result in losses to the business. A sound understanding of the numbers as well as the causes behind these failures helps improve operational experience by not only allowing us to be better equipped to tolerate failures but also to bring down the hardware cost through engineering, directly leading to a saving for the company. To the best of our knowledge, this paper is the first attempt to study server failures and hardware repairs for large datacenters. We present a detailed analysis of failure characteristics as well as a preliminary analysis on failure predictors. We hope that the results presented in this paper will serve as motivation to foster further research in this area.",https://doi.org/10.1145/1807128.1807161,True,"[{'database': 'ACM', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 8}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud computing')"", 'index': 5}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 46}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 7}]",508.0,['decision-tree'],"['configuration', 'host-metrics']",['discussion'],['failure-prediction'],"['server', 'raid', 'memory', 'hard-drive']",Proceedings of the 1st ACM symposium on Cloud computing,True,['failure-management'],True,,['hardware-failure-prediction'],True,25.0,10.1145/1807128.1807161,Conference Paper,"['failures', 'datacenters']",
541,Predicting defect densities in source code files with decision tree learners,"['Patrick Knab', 'Martin Pinzger', 'Abraham Bernstein']",2006,"With the advent of open source software repositories the data available for defect prediction in source files increased tremendously. Although traditional statistics turned out to derive reasonable results the sheer amount of data and the problem context of defect prediction demand sophisticated analysis such as provided by current data mining and machine learning techniques.In this work we focus on defect density prediction and present an approach that applies a decision tree learner on evolution data extracted from the Mozilla open source web browser project. The evolution data includes different source code, modification, and defect measures computed from seven recent Mozilla releases. Among the modification measures we also take into account the change coupling, a measure for the number of change-dependencies between source files. The main reason for choosing decision tree learners, instead of for example neural nets, was the goal of finding underlying rules which can be easily interpreted by humans. To find these rules, we set up a number of experiments to test common hypotheses regarding defects in software entities. Our experiments showed, that a simple tree learner can produce good results with various sets of input data.",https://doi.org/10.1145/1137983.1138012,True,"[{'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault prediction' OR 'failure prediction')"", 'index': 76}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault prediction' OR 'failure prediction')"", 'index': 96}]",63.0,,,,['failure-prediction'],,Proceedings of the 2006 international workshop on Mining software repositories,True,['failure-management'],,,,,,10.1145/1137983.1138012,Conference Paper,"['decision tree learner', 'data mining', 'defect prediction']",
542,Lifelong Anomaly Detection Through Unlearning,"['Min Du', 'Zhi Chen', 'Chang Liu', 'Rajvardhan Oak', 'Dawn Song']",2019,"Anomaly detection is essential towards ensuring system security and reliability. Powered by constantly generated system data, deep learning has been found both effective and flexible to use, with its ability to extract patterns without much domain knowledge. Existing anomaly detection research focuses on a scenario referred to as zero-positive, which means that the detection model is only trained for normal (i.e., negative) data. In a real application scenario, there may be additional manually inspected positive data provided after the system is deployed. We refer to this scenario as lifelong anomaly detection. However, we find that existing approaches are not easy to adopt such new knowledge to improve system performance. In this work, we are the first to explore the lifelong anomaly detection problem, and propose novel approaches to handle corresponding challenges. In particular, we propose a framework called unlearning, which can effectively correct the model when a false negative (or a false positive) is labeled. To this aim, we develop several novel techniques to tackle two challenges referred to as exploding loss and catastrophic forgetting. In addition, we abstract a theoretical framework based on generative models. Under this framework, our unlearning approach can be presented in a generic way to be applied to most zero-positive deep learning-based anomaly detection algorithms to turn them into corresponding lifelong anomaly detection solutions. We evaluate our approach using two state-of-the-art zero-positive deep learning anomaly detection architectures and three real-world tasks. The results show that the proposed approach is able to significantly reduce the number of false positives and false negatives through unlearning.",https://doi.org/10.1145/3319535.3363226,True,"[{'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 8}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('anomaly detection' OR 'outlier detection')"", 'index': 6}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 23}, {'database': 'ACM', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 10}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 12}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 16}, {'database': 'ACM', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 9}]",1.0,['multilayer-perceptron'],,['new-method'],['failure-detection'],,Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security,True,['failure-management'],,,['anomaly-detection'],,,10.1145/3319535.3363226,Conference Paper,"['anomaly detection', 'unlearning', 'online learning']",
543,Recurrent Neural Network Attention Mechanisms for Interpretable System Log Anomaly Detection,"['Andy Brown', 'Aaron Tuor', 'Brian Hutchinson', 'Nicole Nichols']",2018,"Deep learning has recently demonstrated state-of-the art performance on key tasks related to the maintenance of computer systems, such as intrusion detection, denial of service attack detection, hardware and software system failures, and malware detection. In these contexts, model interpretability is vital for administrator and analyst to trust and act on the automated analysis of machine learning models. Deep learning methods have been criticized as black box oracles which allow limited insight into decision factors. In this work we seek to ""bridge the gap"" between the impressive performance of deep learning models and the need for interpretable model introspection. To this end we present recurrent neural network (RNN) language models augmented with attention for anomaly detection in system logs. Our methods are generally applicable to any computer system and logging source. By incorporating attention variants into our RNN language models we create opportunities for model introspection and analysis without sacrificing state-of-the art performance. We demonstrate model performance and illustrate model interpretability on an intrusion detection task using the Los Alamos National Laboratory (LANL) cyber security dataset, reporting upward of 0.99 area under the receiver operator characteristic curve despite being trained only on a single day's worth of data.",https://doi.org/10.1145/3217871.3217872,True,"[{'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 54}, {'database': 'ACM', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 82}, {'database': 'ACM', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 52}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 8}, {'database': 'arxiv', 'search_string': ""('DL' OR 'deep learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 36}, {'database': 'arxiv', 'search_string': ""('DL' OR 'deep learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 17}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 269}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 143}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 117}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 146}]",30.0,['rnn'],,['novel-use'],['failure-detection'],,Proceedings of the First Workshop on Machine Learning for Computing Systems,True,['failure-management'],True,,['anomaly-detection'],True,62.0,10.1145/3217871.3217872,Conference Paper,"['System Log Analysis', 'Online Training', 'Anomaly detection', 'Interpretable Machine Learning', 'Attention', 'Recurrent Neural Networks']",['online']
544,CloudSeer: Workflow Monitoring of Cloud Infrastructures via Interleaved Logs,"['Xiao Yu', 'Pallavi Joshi', 'Jianwu Xu', 'Guoliang Jin', 'Hui Zhang', 'Guofei Jiang']",2016,"Cloud infrastructures provide a rich set of management tasks that operate computing, storage, and networking resources in the cloud. Monitoring the executions of these tasks is crucial for cloud providers to promptly find and understand problems that compromise cloud availability. However, such monitoring is challenging because there are multiple distributed service components involved in the executions. CloudSeer enables effective workflow monitoring. It takes a lightweight non-intrusive approach that purely works on interleaved logs widely existing in cloud infrastructures. CloudSeer first builds an automaton for the workflow of each management task based on normal executions, and then it checks log messages against a set of automata for workflow divergences in a streaming manner. Divergences found during the checking process indicate potential execution problems, which may or may not be accompanied by error log messages. For each potential problem, CloudSeer outputs necessary context information including the affected task automaton and related log messages hinting where the problem occurs to help further diagnosis. Our experiments on OpenStack, a popular open-source cloud infrastructure, show that CloudSeer's efficiency and problem-detection capability are suitable for online monitoring.",https://doi.org/10.1145/2872362.2872407,True,"[{'database': 'ACM', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 56}]",24.0,,,,['failure-detection'],,Proceedings of the Twenty-First International Conference on Architectural Support for Programming Languages and Operating Systems,True,['failure-management'],,,,,,10.1145/2872362.2872407,Conference Paper,"['distributed systems', 'workflow monitoring', 'log analysis', 'cloud infrastructures']",
545,Detecting large-scale system problems by mining console logs,"['Wei Xu', 'Ling Huang', 'Armando Fox', 'David Patterson', 'Michael I. Jordan']",2009,"Surprisingly, console logs rarely help operators detect problems in large-scale datacenter services, for they often consist of the voluminous intermixing of messages from many software components written by independent developers. We propose a general methodology to mine this rich source of information to automatically detect system runtime problems. We first parse console logs by combining source code analysis with information retrieval to create composite features. We then analyze these features using machine learning to detect operational problems. We show that our method enables analyses that are impossible with previous methods because of its superior ability to create sophisticated features. We also show how to distill the results of our analysis to an operator-friendly one-page decision tree showing the critical messages associated with the detected problems. We validate our approach using the Darkstar online game server and the Hadoop File System, where we detect numerous real problems with high accuracy and few false positives. In the Hadoop case, we are able to analyze 24 million lines of console logs in 3 minutes. Our methodology works on textual console logs of any size and requires no changes to the service software, no human input, and no knowledge of the software's internals.",https://doi.org/10.1145/1629575.1629587,True,"[{'database': 'ACM', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 63}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 9}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 85}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 44}, {'database': 'ACM', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 58}]",353.0,"['dimensionality-reduction', 'decision-tree']","['logs', 'source-code']",,['failure-detection'],"['hadoop', 'filesystem']",Proceedings of the ACM SIGOPS 22nd symposium on Operating systems principles,True,['failure-management'],True,,['anomaly-detection'],True,56.0,10.1145/1629575.1629587,Conference Paper,"['tracing', 'problem detection', 'source code analysis', 'monitoring', 'console log analysis', 'statistical learning', 'pca']",
546,VCONF: a reinforcement learning approach to virtual machines auto-configuration,"['Jia Rao', 'Xiangping Bu', 'Cheng-Zhong Xu', 'Leyi Wang', 'George Yin']",2009,"Virtual machine (VM) technology enables multiple VMs to share resources on the same host. Resources allocated to the VMs should be re-configured dynamically in response to the change of application demands or resource supply. Because VM execution involves privileged domain and VM monitor, this causes uncertainties in VMs' resource to performance mapping and poses challenges in online determination of appropriate VM configurations. In this paper, we propose a reinforcement learning (RL) based approach, namely VCONF, to automate the VM configuration process. VCONF employs model-based RL algorithms to address the scalability and adaptability issues in applying RL in systems management. Experimental results on both controlled environments and a testbed of clouds with Xen VMs and representative server workloads demonstrate the effectiveness of VCONF. The approach is able to find optimal (near optimal) configurations in small scale systems and shows good adaptability and scalability.",https://doi.org/10.1145/1555228.1555263,True,"[{'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud computing')"", 'index': 67}]",253.0,['reinforcement-learning'],,['novel-use'],"['resource-consolidation', 'configuration']","['vm', 'xen']",Proceedings of the 6th international conference on Autonomic computing,True,['resource-provisioning'],True,True,,,,10.1145/1555228.1555263,Conference Paper,"['cloud computing', 'reinforcement learning', 'virtual machines', 'autonomic computing']",
547,Automated support for classifying software failure reports,"['Andy Podgurski', 'David Leon', 'Patrick Francis', 'Wes Masri', 'Melinda Minch', 'Jiayang Sun', 'Bin Wang']",2003,This paper proposes automated support for classifying reported software failures in order to facilitate prioritizing them and diagnosing their causes. A classification strategy is presented that involves the use of supervised and unsupervised pattern classification and multivariate visualization. These techniques are applied to profiles of failed executions in order to group together failures with the same or similar causes. The resulting classification is then used to assess the frequency and severity of failures caused by particular defects and to help diagnose those defects. The results of applying the proposed classification strategy to failures of three large subject programs are reported These results indicate that the strategy can be effective.,https://ieeexplore.ieee.org/document/1201224,True,"[{'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('fault localization' OR 'failure localization')"", 'index': 45}, {'database': 'ACM', 'search_string': ""'logistic regression' AND ('fault localization' OR 'failure localization')"", 'index': 20}]",374.0,"['clustering', 'logistic-regression']",['runs'],['novel-use'],['root-cause-analysis'],['software'],Proceedings of the 25th International Conference on Software Engineering,True,['failure-management'],True,,['rca-others'],True,83.0,,Conference Paper,,
548,Semi-supervised network traffic classification,"['Jeffrey Erman', 'Anirban Mahanti', 'Martin Arlitt', 'Ira Cohen', 'Carey Williamson']",2007,"Identifying and categorizing network traffic by application type is challenging because of the continued evolution of applications, especially of those with a desire to be undetectable. The diminished effectiveness of port-based identification and the overheads of deep packet …",https://doi.org/10.1145/1254882.1254934,True,"[{'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 2}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 2}]",65.0,,,,['failure-detection'],,Proceedings of the 2007 ACM SIGMETRICS international conference on Measurement and modeling of computer systems,True,['failure-management'],,,,,,10.1145/1254882.1254934,Conference Paper,"['traffic classification', 'semi-supervised learning']",
549,Failure proximity: a fault localization-based approach,"['Chao Liu', 'Jiawei Han']",2006,"Recent software systems usually feature an automated failure reporting system, with which a huge number of failing traces are collected every day. In order to prioritize fault diagnosis, failing traces due to the same fault are expected to be grouped together. Previous methods, by hypothesizing that similar failing traces imply the same fault, cluster failing traces based on the literal trace similarity, which we call trace proximity. However, since a fault can be triggered in many ways, failing traces due to the same fault can be quite different. Therefore, previous methods actually group together traces exhibiting similar behaviors, like similar branch coverage, rather than traces due to the same fault. In this paper, we propose a new type of failure proximity, called R-Proximity, which regards two failing traces as similar if they suggest roughly the same fault location. The fault location each failing case suggests is automatically obtained with Sober, an existing statistical debugging tool. We show that with R-Proximity, failing traces due to the same fault can be grouped together. In addition, we find that R-Proximity is helpful for statistical debugging: It can help developers interpret and utilize the statistical debugging result. We illustrate the usage of R-Proximity with a case study on the grep program and some experiments on the Siemens suite, and the result clearly demonstrates the advantage of R-Proximity over trace proximity.",https://doi.org/10.1145/1181775.1181782,True,"[{'database': 'ACM', 'search_string': ""'clustering' AND ('fault localization' OR 'failure localization')"", 'index': 30}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 79}]",78.0,,,,['root-cause-analysis'],,Proceedings of the 14th ACM SIGSOFT international symposium on Foundations of software engineering,True,['failure-management'],,,,,,10.1145/1181775.1181782,Conference Paper,"['statistical debugging', 'debugging aids', 'failure proximity']",
550,BALLERINA: automatic generation and clustering of efficient random unit tests for multithreaded code,"['Adrian Nistor', 'Qingzhou Luo', 'Michael Pradel', 'Thomas R. Gross', 'Darko Marinov']",2012,,https://ieeexplore.ieee.org/document/6227145,True,"[{'database': 'ACM', 'search_string': ""'clustering' AND ('fault localization' OR 'failure localization')"", 'index': 56}]",54.0,['clustering'],,,['root-cause-analysis'],,Proceedings of the 34th International Conference on Software Engineering,True,['failure-management'],,,['rca-others'],,,,Conference Paper,,
551,Autonomic Provisioning with Self-Adaptive Neural Fuzzy Control for Percentile-Based Delay Guarantee,"['Palden Lama', 'Xiaobo Zhou']",2013,"Autonomic server provisioning for performance assurance is a critical issue in Internet services. It is challenging to guarantee that requests flowing through a multi-tier system will experience an acceptable distribution of delays. The difficulty is mainly due to highly dynamic workloads, the complexity of underlying computer systems, and the lack of accurate performance models. We propose a novel autonomic server provisioning approach based on a model-independent self-adaptive Neural Fuzzy Control (NFC). Existing model-independent fuzzy controllers are designed manually on a trial-and-error basis, and are often ineffective in the face of highly dynamic workloads. NFC is a hybrid of control-theoretical and machine learning techniques. It is capable of self-constructing its structure and adapting its parameters through fast online learning. We further enhance NFC to compensate for the effect of server switching delays. Extensive simulations demonstrate that, compared to a rule-based fuzzy controller and a Proportional-Integral controller, the NFC-based approach delivers superior performance assurance in the face of highly dynamic workloads. It is robust to variation in workload intensity, characteristics, delay target, and server switching delays. We demonstrate the feasibility and performance of the NFC-based approach with a testbed implementation in virtualized blade servers hosting a multi-tier online auction benchmark.",https://doi.org/10.1145/2491465.2491468,True,"[{'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 97}]",15.0,,,,,,ACM Transactions on Autonomous and Adaptive Systems (TAAS),True,['resource-provisioning'],,True,,,,10.1145/2491465.2491468,Journal Article,"['percentile-based delay guarantee', 'multi-tier internet services', 'server virtualization', 'Resource allocation', 'self-adaptation', 'neural fuzzy control']",
552,Ganesha: blackBox diagnosis of MapReduce systems,"['Xinghao Pan', 'Jiaqi Tan', 'Soila Kavulya', 'Rajeev Gandhi', 'Priya Narasimhan']",2010,"Ganesha aims to diagnose faults transparently (in a black-box manner) in MapReduce systems, by analyzing OS-level metrics. Ganesha's approach is based on peer-symmetry under fault-free conditions, and can diagnose faults that manifest asymmetrically at nodes within a MapReduce system. We evaluate Ganesha by diagnosing Hadoop problems for the Gridmix Hadoop benchmark on 10-node and 50-node MapReduce clusters on Amazon's EC2. We also candidly highlight faults that escape Ganesha's diagnosis.",https://doi.org/10.1145/1710115.1710118,True,"[{'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 43}, {'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 7}, {'database': 'ACM', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 53}]",31.0,,,,['root-cause-analysis'],,ACM SIGMETRICS Performance Evaluation Review,True,['failure-management'],,,,,,10.1145/1710115.1710118,Journal Article,,
553,vPerfGuard: an automated model-driven framework for application performance diagnosis in consolidated cloud environments,"['Pengcheng Xiong', 'Calton Pu', 'Xiaoyun Zhu', 'Rean Griffith']",2013,"Many business customers hesitate to move all their applications to the cloud due to performance concerns. White-box diagnosis relies on human expert experience or performance troubleshooting ""cookbooks"" to find potential performance bottlenecks. Despite wide adoption, the scalability and adaptivity of such approaches remain severely constrained, especially in a highly-dynamic, consolidated cloud environment. Leveraging the rich telemetry collected from applications and systems in the cloud, and the power of statistical learning, vPerfGuard complements the existing approaches with a model-driven framework by: (1) automatically identifying system metrics that are most predictive of application performance, and (2) adaptively detecting changes in the performance and potential shifts in the predictive metrics that may accompany such a change. Although correlation does not imply causation, the predictive system metrics point to potential causes that can guide a cloud service provider to zero in on the root cause. We have implemented vPerfGuard as a combination of three modules: a sensor module, a model building module, and a model updating module. We evaluate its effectiveness using different benchmarks and different workload types, specifically focusing on various resource (CPU, memory, disk I/O) contention scenarios that are caused by workload surges or ""noisy neighbors"". The results show that vPerfGuard automatically points to the correct performance bottleneck in each scenario, including the type of the contended resource and the host where the contention occurred.",https://doi.org/10.1145/2479871.2479909,True,"[{'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 85}, {'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 87}]",31.0,,,,"['failure-detection', 'resource-consolidation', 'anomaly-detection']",,Proceedings of the 4th ACM/SPEC International Conference on Performance Engineering,True,"['resource-provisioning', 'failure-management']",,,,,,10.1145/2479871.2479909,Conference Paper,"['automated framework', 'performance analysis', 'model-driven', 'statistical methods', 'cloud computing', 'performance diagnosis']",
554,Incremental Clustering for Semi-Supervised Anomaly Detection applied on Log Data,"['Markus Wurzenberger', 'Florian Skopik', 'Max Landauer', 'Philipp Greitbauer', 'Roman Fiedler', 'Wolfgang Kastner']",2017,"Anomaly detection based on white-listing and self-learning has proven to be a promising approach to detect customized and advanced cyber attacks. Anomaly detection aims at detecting significant deviations from normal system and network behavior. A well-known method to classify anomalous and normal system behavior is clustering of log lines. However, this approach has been applied for forensic purposes only, where log data dumps are investigated retrospectively. In order to make this concept applicable for on-line anomaly detection, i.e., at the time the log lines are produced, some major extensions to existing approaches are required. Especially distance based clustering approaches usually fail building the required large distance matrices and rely on time-consuming recalculations of the cluster-map on every arriving log line. An incremental clustering approach seems suitable to solve this issues. Thus, we introduce a semi-supervised concept for incremental clustering of log data that builds the basis for a novel on-line anomaly detection solution based on log data streams. Its operation is independent from the syntax and semantics of the processed log lines, which makes it generally applicable. We demonstrate that that the introduced anomaly detection approach allows to achieve both a high recall and a high precision while maintaining linear complexity.",https://doi.org/10.1145/3098954.3098973,True,"[{'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 51}, {'database': 'ACM', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 59}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 2}, {'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 89}]",5.0,,,,['failure-detection'],,"Proceedings of the 12th International Conference on Availability, Reliability and Security",True,['failure-management'],,,,,,10.1145/3098954.3098973,Conference Paper,"['log line clustering', 'intrusion detection', 'anomaly detection']",
555,Eigenspace-based anomaly detection in computer systems,"['Tsuyoshi IDÉ', 'Hisashi KASHIMA']",2004,"We report on an automated runtime anomaly detection method at the application layer of multi-node computer systems. Although several network management systems are available in the market, none of them have sufficient capabilities to detect faults in multi-tier Web-based systems with redundancy. We model a Web-based system as a weighted graph, where each node represents a ""service"" and each edge represents a dependency between services. Since the edge weights vary greatly over time, the problem we address is that of anomaly detection from a time sequence of graphs.In our method, we first extract a feature vector from the adjacency matrix that represents the activities of all of the services. The heart of our method is to use the principal eigenvector of the eigenclusters of the graph. Then we derive a probability distribution for an anomaly measure defined for a time-series of directional data derived from the graph sequence. Given a critical probability, the threshold value is adaptively updated using a novel online algorithm.We demonstrate that a fault in a Web application can be automatically detected and the faulty services are identified without using detailed knowledge of the behavior of the system.",https://doi.org/10.1145/1014052.1014102,True,"[{'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 87}, {'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 75}]",100.0,,,,['failure-detection'],,Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining,True,['failure-management'],,,['anomaly-detection'],,,10.1145/1014052.1014102,Conference Paper,"['Perron-Frobenius theorem', 'time sequence of graphs', 'singular value decomposition', 'principal eigenvector', 'von Mises-Fisher distribution']",
556,Paragon: QoS-aware scheduling for heterogeneous datacenters,"['Christina Delimitrou', 'Christos Kozyrakis']",2013,"Large-scale datacenters (DCs) host tens of thousands of diverse applications each day. However, interference between colocated workloads and the difficulty to match applications to one of the many hardware platforms available can degrade performance, violating the quality of service (QoS) guarantees that many cloud workloads require. While previous work has identified the impact of heterogeneity and interference, existing solutions are computationally intensive, cannot be applied online and do not scale beyond few applications. We present Paragon, an online and scalable DC scheduler that is heterogeneity and interference-aware. Paragon is derived from robust analytical methods and instead of profiling each application in detail, it leverages information the system already has about applications it has previously seen. It uses collaborative filtering techniques to quickly and accurately classify an unknown, incoming workload with respect to heterogeneity and interference in multiple shared resources, by identifying similarities to previously scheduled applications. The classification allows Paragon to greedily schedule applications in a manner that minimizes interference and maximizes server utilization. Paragon scales to tens of thousands of servers with marginal scheduling overheads in terms of time or state. We evaluate Paragon with a wide range of workload scenarios, on both small and large-scale systems, including 1,000 servers on EC2. For a 2,500-workload scenario, Paragon enforces performance guarantees for 91% of applications, while significantly improving utilization. In comparison, heterogeneity-oblivious, interference-oblivious and least-loaded schedulers only provide similar guarantees for 14%, 11% and 3% of workloads. The differences are more striking in oversubscribed scenarios where resource efficiency is more critical.",https://doi.org/10.1145/2451116.2451125,True,"[{'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 37}]",314.0,['collaborative-filtering'],,,['scheduling'],,Proceedings of the eighteenth international conference on Architectural support for programming languages and operating systems,True,['resource-provisioning'],True,,,,,10.1145/2451116.2451125,Conference Paper,"['qos', 'heterogeneity', 'interference', 'cloud computing', 'datacenter', 'scheduling']",
557,Predicting Disk Replacement towards Reliable Data Centers,"['Mirela Madalina Botezatu', 'Ioana Giurgiu', 'Jasmina Bogojeska', 'Dorothea Wiesmann']",2016,"Disks are among the most frequently failing components in today's IT environments. Despite a set of defense mechanisms such as RAID, the availability and reliability of the system are still often impacted severely. In this paper, we present a highly accurate SMART-based analysis pipeline that can correctly predict the necessity of a disk replacement even 10-15 days in advance. Our method has been built and evaluated on more than 30000 disks from two major manufacturers, monitored over 17 months. Our approach employs statistical techniques to automatically detect which SMART parameters correlate with disk replacement and uses them to predict the replacement of a disk with even 98% accuracy.",https://doi.org/10.1145/2939672.2939699,True,"[{'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 79}, {'database': 'ACM', 'search_string': ""'logistic regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 39}]",25.0,,,,['failure-prediction'],,Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,True,['failure-management'],,,,,,10.1145/2939672.2939699,Conference Paper,"['classification', 'changepoint', 'disk replacement', 'time series']",
558,Formal concept analysis applied to fault localization,['Peggy Cellier'],2008,"One time-consuming task in the development of software is debugging. Recent work in fault localization crosschecks traces of correct and failing execution traces, it implicitly searches for association rules which indicate that executing a line will most probably cause the whole execution to fail. This technique has some limitations: it assumes that an error has a single faulty statement origin, and that lines are independent. Our research hypothesis is that using association rules with more expressive premises, some limitations can be alleviated. The solution that we propose combines association rules and formal concept analysis. Our technique is already usable when the size of the execution traces is not too large. We conjecture that the technique can be used to analyze large executions, thanks to the information contained in the Abstract Syntax Tree.",https://doi.org/10.1145/1370175.1370220,True,"[{'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 18}, {'database': 'ACM', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('fault localization' OR 'failure localization')"", 'index': 9}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault localization' OR 'failure localization')"", 'index': 32}]",10.0,,,,['root-cause-analysis'],,Companion of the 30th international conference on Software engineering,True,['failure-management'],,,,,,10.1145/1370175.1370220,Conference Paper,"['association rules', 'data mining', 'formal concept analysis', 'debugging', 'fault localization']",
559,An observation-based model for fault localization,"['Rui Abreu', 'Peter Zoeteweij', 'Arjan J. C. van Gemund']",2008,"Automatic techniques for helping developers in finding the root causes of software failures are extremely important in the development cycle of software. In this paper we study a dynamic modeling approach to fault localization, which is based on logic reasoning over program traces. We present a simple diagnostic performance model to assess the influence of various parameters, such as test set size and coverage, on the debugging effort required to find the root causes of software failures. The model shows that our approach unambiguously reveals the actual faults, provided that sufficient test cases are available. This optimal diagnostic performance is confirmed by numerical experiments. Furthermore, we present preliminary experiments on the diagnostic capabilities of this approach using the single-fault Siemens benchmark set. We show that, for the Siemens set, the approach presented in this paper yields a better diagnostic ranking than other well-known techniques.",https://doi.org/10.1145/1401827.1401841,True,"[{'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 43}, {'database': 'ACM', 'search_string': ""('DL' OR 'deep learning') AND ('fault localization' OR 'failure localization')"", 'index': 86}]",26.0,,,,['root-cause-analysis'],,Proceedings of the 2008 international workshop on dynamic analysis: held in conjunction with the ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA 2008),True,['failure-management'],,,,,,10.1145/1401827.1401841,Conference Paper,"['program spectra', 'test data analysis', 'model-based diagnosis', 'software fault diagnosis']",
560,Combining filtering and statistical methods for anomaly detection,"['Augustin Soule', 'Kavé Salamatian', 'Nina Taft']",2005,"In this work we develop an approach for anomaly detection for large scale networks such as that of an enterprize or an ISP. The traffic patterns we focus on for analysis are that of a network-wide view of the traffic state, called the traffic matrix. In the first step a Kalman filter is used to filter out the ""normal"" traffic. This is done by comparing our future predictions of the traffic matrix state to an inference of the actual traffic matrix that is made using more recent measurement data than those used for prediction. In the second step the residual filtered process is then examined for anomalies. We explain here how any anomaly detection method can be viewed as a problem in statistical hypothesis testing. We study and compare four different methods for analyzing residuals, two of which are new. These methods focus on different aspects of the traffic pattern change. One focuses on instantaneous behavior, another focuses on changes in the mean of the residual process, a third on changes in the variance behavior, and a fourth examines variance changes over multiple timescales. We evaluate and compare all of these methods using ROC curves that illustrate the full tradeoff between false positives and false negatives for the complete spectrum of decision thresholds.",https://dl.acm.org/doi/10.5555/1251086.1251117,True,"[{'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 86}]",375.0,['kalman-filter'],,['new-method'],['failure-detection'],['network'],Proceedings of the 5th ACM SIGCOMM conference on Internet measurement,True,['failure-management'],False,,['anomaly-detection'],,,10.5555/1251086.1251117,Conference Paper,,
561,Using reinforcement learning for controlling an elastic web application hosting platform,"['Han Li', 'Srikumar Venugopal']",2011,"In this paper, we propose and implement a control mechanism that interfaces with Infrastructure as a Service (IaaS) or cloud providers to provision resources and manage instances of web applications in response to volatile and complex request patterns. We use reinforcement learning to orchestrate control actions such as provisioning servers and application placement to meet performance requirements and minimize ongoing costs. The mechanism is incorporated in a distributed, elastic hosting architecture that is evaluated using actual web applications running on resources from Amazon EC2.",https://doi.org/10.1145/1998582.1998630,True,"[{'database': 'ACM', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 98}]",21.0,,,,,,Proceedings of the 8th ACM international conference on Autonomic computing,True,['resource-provisioning'],,,,,,10.1145/1998582.1998630,Conference Paper,"['provisioning', 'reinforcement learning', 'elastic computing']",
562,Optimal cloud resource auto-scaling for web applications,"['Jing Jiang', 'Jie Lu', 'Guangquan Zhang', 'Guodong Long']",2013,"In the on-demand cloud environment, web application providers have the potential to scale virtual resources up or down to achieve cost-effective outcomes. True elasticity and cost-effectiveness in the pay-per-use cloud business model, however, have not yet been achieved. To address this challenge, we propose a novel cloud resource auto-scaling scheme at the virtual machine (VM) level for web application providers. The scheme automatically predicts the number of web requests and discovers an optimal cloud resource demand with cost-latency trade-off. Based on this demand, the scheme makes a resource scaling decision that is up or down or NOP (no operation) in each time-unit re-allocation. We have implemented the scheme on the Amazon cloud platform and evaluated it using three real-world web log datasets. Our experiment results demonstrate that the proposed scheme achieves resource auto-scaling with an optimal cost-latency trade-off, as well as low SLA violations.",https://doi.org/10.1109/CCGrid.2013.73,True,"[{'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 96}, {'database': 'ACM', 'search_string': ""'regression' AND ('cloud')"", 'index': 27}]",14.0,,,,,,"Proceedings of the 13th IEEE/ACM International Symposium on Cluster, Cloud, and Grid Computing",True,['resource-provisioning'],,True,,,,10.1109/CCGrid.2013.73,Conference Paper,"['web services', 'resource scaling', 'resource prediction', 'elastic computing', 'cloud computing']",
563,Pack & Cap: adaptive DVFS and thread packing under power caps,"['Ryan Cochran', 'Can Hankendi', 'Ayse K. Coskun', 'Sherief Reda']",2011,"The ability to cap peak power consumption is a desirable feature in modern data centers for energy budgeting, cost management, and efficient power delivery. Dynamic voltage and frequency scaling (DVFS) is a traditional control knob in the tradeoff between server power and performance. Multi-core processors and the parallel applications that take advantage of them introduce new possibilities for control, wherein workload threads are packed onto a variable number of cores and idle cores enter low-power sleep states. This paper proposes Pack & Cap, a control technique designed to make optimal DVFS and thread packing control decisions in order to maximize performance within a power budget. In order to capture the workload dependence of the performance-power Pareto frontier, a multinomial logistic regression (MLR) classifier is built using a large volume of performance counter, temperature, and power characterization data. When queried during runtime, the classifier is capable of accurately selecting the optimal operating point. We implement and validate this method on a real quad-core system running the PARSEC parallel benchmark suite. When varying the power budget during runtime, Pack & Cap meets power constraints 82% of the time even in the absence of a power measuring device. The addition of thread packing to DVFS as a control knob increases the range of feasible power constraints by an average of 21% when compared to DVFS alone and reduces workload energy consumption by an average of 51.6% compared to existing control techniques that achieve the same power range.",https://doi.org/10.1145/2155620.2155641,True,"[{'database': 'ACM', 'search_string': ""'logistic regression' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 80}, {'database': 'ACM', 'search_string': ""'logistic regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 13}]",267.0,['logistic-regression'],['host-metrics'],,['power-management'],['power'],Proceedings of the 44th Annual IEEE/ACM International Symposium on Microarchitecture,True,['resource-provisioning'],,True,,,,10.1145/2155620.2155641,Conference Paper,,
564,A system software approach to proactive memory-error avoidance,"['Carlos H. A. Costa', 'Yoonho Park', 'Bryan S. Rosenburg', 'Chen-Yong Cher', 'Kyung Dong Ryu']",2014,"Today's HPC systems use two mechanisms to address main-memory errors. Error-correcting codes make correctable errors transparent to software, while checkpoint/restart (CR) enables recovery from uncorrectable errors. Unfortunately, CR overhead will be enormous at exascale due to the high failure rate of memory. We propose a new OS-based approach that proactively avoids memory errors using prediction. This scheme exposes correctable error information to the OS, which migrates pages and offlines unhealthy memory to avoid application crashes. We analyze memory error patterns in extensive logs from a BG/P system and show how correctable error patterns can be used to identify memory likely to fail. We implement a proactive memory management system on BG/Q by extending the firmware and Linux. We evaluate our approach with a realistic workload and compare our overhead against CR. We show improved resilience with negligible performance overhead for applications.",https://doi.org/10.1109/SC.2014.63,True,"[{'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 60}]",14.0,['correlation'],['logs'],,"['failure-prediction', 'failure-prevention']","['memory', 'hpc']","Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis",True,['failure-management'],True,,['hardware-failure-prediction'],True,38.0,10.1109/SC.2014.63,Conference Paper,"['operating systems', 'memory structures', 'reliability', 'fault-tolerance']",
565,Log-based predictive maintenance,"['Ruben Sipos', 'Dmitriy Fradkin', 'Fabian Moerchen', 'Zhuang Wang']",2014,"Success of manufacturing companies largely depends on reliability of their products. Scheduled maintenance is widely used to ensure that equipment is operating correctly so as to avoid unexpected breakdowns. Such maintenance is often carried out separately for every component, based on its usage or simply on some fixed schedule. However, scheduled maintenance is labor-intensive and ineffective in identifying problems that develop between technician's visits. Unforeseen failures still frequently occur. In contrast, predictive maintenance techniques help determine the condition of in-service equipment in order to predict when and what repairs should be performed. The main goal of predictive maintenance is to enable pro-active scheduling of corrective work, and thus prevent unexpected equipment failures.",https://doi.org/10.1145/2623330.2623340,True,"[{'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 82}, {'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('fault prediction' OR 'failure prediction')"", 'index': 97}]",54.0,"['linear-regression', 'support-vector-machine']",['logs'],"['new-method', 'comparison']",['failure-prediction'],[],Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining,True,['failure-management'],,,['system-failure-prediction'],,,10.1145/2623330.2623340,Conference Paper,"['predictive maintenance', 'machine learning', 'crisp-dm', 'log mining']",
566,Adaptive system anomaly prediction for large-scale hosting infrastructures,"['Yongmin Tan', 'Xiaohui Gu', 'Haixun Wang']",2010,"Large-scale hosting infrastructures require automatic system anomaly management to achieve continuous system operation. In this paper, we present a novel adaptive runtime anomaly prediction system, called ALERT, to achieve robust hosting infrastructures. In contrast to traditional anomaly detection schemes, ALERT aims at raising advance anomaly alerts to achieve just-in-time anomaly prevention. We propose a novel context-aware anomaly prediction scheme to improve prediction accuracy in dynamic hosting infrastructures. We have implemented the ALERT system and deployed it on several production hosting infrastructures such as IBM System S stream processing cluster and PlanetLab. Our experiments show that ALERT can achieve high prediction accuracy for a range of system anomalies and impose low overhead to the hosting infrastructure.",https://doi.org/10.1145/1835698.1835741,True,"[{'database': 'ACM', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 88}]",32.0,,,,['failure-detection'],,Proceedings of the 29th ACM SIGACT-SIGOPS symposium on Principles of distributed computing,True,['failure-management'],,,,,,10.1145/1835698.1835741,Conference Paper,"['context-aware prediction model', 'anomaly prediction']",
567,AROMA: automated resource allocation and configuration of mapreduce environment in the cloud,"['Palden Lama', 'Xiaobo Zhou']",2012,"Distributed data processing framework MapReduce is increasingly deployed in Clouds to leverage the pay-per-usage cloud computing model. Popular Hadoop MapReduce environment expects that end users determine the type and amount of Cloud resources for reservation as well as the configuration of Hadoop parameters. However, such resource reservation and job provisioning decisions require in-depth knowledge of system internals and laborious but often ineffective parameter tuning. We propose and develop AROMA, a system that automates the allocation of heterogeneous Cloud resources and configuration of Hadoop parameters for achieving quality of service goals while minimizing the incurred cost. It addresses the significant challenge of provisioning ad-hoc jobs that have performance deadlines in Clouds through a novel two-phase machine learning and optimization framework. Its technical core is a support vector machine based performance model that enables the integration of various aspects of resource provisioning and auto-configuration of Hadoop jobs. It adapts to ad-hoc jobs by robustly matching their resource utilization signature with previously executed jobs and making provisioning decisions accordingly. We implement AROMA as an automated job provisioning system for Hadoop MapReduce hosted in virtualized HP ProLiant blade servers. Experimental results show AROMA's effectiveness in providing performance guarantee of diverse Hadoop benchmark jobs while minimizing the cost of Cloud resource usage.",https://doi.org/10.1145/2371536.2371547,True,"[{'database': 'ACM', 'search_string': ""('support vector machine' OR 'SVM') AND ('cloud')"", 'index': 59}]",89.0,,,,['resource-consolidation'],,Proceedings of the 9th international conference on Autonomic computing,True,['resource-provisioning'],,True,,,,10.1145/2371536.2371547,Conference Paper,"['auto-configuration', 'resource allocation', 'mapreduce']",
568,An empirical evaluation of entropy-based traffic anomaly detection,"['George Nychis', 'Vyas Sekar', 'David G. Andersen', 'Hyong Kim', 'Hui Zhang']",2008,"Entropy-based approaches for anomaly detection are appealing since they provide more fine-grained insights than traditional traffic volume analysis. While previous work has demonstrated the benefits of entropy-based anomaly detection, there has been little effort to comprehensively understand the detection power of using entropy-based analysis of multiple traffic distributions in conjunction with each other. We consider two classes of distributions: flow-header features (IP addresses, ports, and flow-sizes), and behavioral features (degree distributions measuring the number of distinct destination/source IPs that each host communicates with). We observe that the timeseries of entropy values of the address and port distributions are strongly correlated with each other and provide very similar anomaly detection capabilities. The behavioral and flow size distributions are less correlated and detect incidents that do not show up as anomalies in the port and address distributions. Further analysis using synthetically generated anomalies also suggests that the port and address distributions have limited utility in detecting scan and bandwidth flood anomalies. Based on our analysis, we discuss important implications for entropy-based anomaly detection.",https://doi.org/10.1145/1452520.1452539,True,"[{'database': 'ACM', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 80}]",177.0,['entropy-selection'],,['novel-use'],['failure-detection'],['network'],Proceedings of the 8th ACM SIGCOMM conference on Internet measurement,True,['failure-management'],,True,['anomaly-detection'],,,10.1145/1452520.1452539,Conference Paper,"['entropy', 'anomaly detection']",
569,Comparing the use of bayesian networks and neural networks in response time modeling for service-oriented systems,"['Rui Zhang', 'Alan J. Bivens']",2007,"The new paradigm of service-oriented computing facilitates easy construction of dynamic, complex distributed systems. Recent research has shown that machine learning methods can be a promising way to autonomously and accurately derive models to assist autonomic management software or humans in understanding system behaviors and making informed decisions. However, the efficacy of different machine learning techniques in describing various system behaviors and meeting distinct application needs has not been systematically understood. Such an understanding can prove crucial in management infrastructure design and implementation for service-oriented systems. This paper is an initial step to bridge the gap and specifically contrasts the applications of Bayesian networks (BN) and neural networks (NN) in modeling the response time of service-oriented systems. Relatively simple BN and NN models are designed and implemented as a base of the comparison study. As far as model performance is concerned, a wide range of simulations show that BNs offer better accuracy, are less sensitive to small data set size and are therefore more suited for environments that change rapidly and need frequent response time model reconstructions; whereas NNs can achieve faster model evaluation time and support management routines that demand intensive response time predictions. From a non-performance perspective, it is analytically concluded that BNs can be more easily understood by human and support multi-direction evaluation, while NNs provide more flexible response time representation.",https://doi.org/10.1145/1272457.1272467,True,"[{'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('remediation' OR 'recovery')"", 'index': 17}]",25.0,"['multilayer-perceptron', 'bayesian-network']",,['comparison'],"['failure-detection', 'anomaly-detection', 'workload-prediction']",,"Proceedings of the 2007 workshop on Service-oriented computing performance: aspects, issues, and approaches",True,"['resource-provisioning', 'failure-management']",,True,,,,10.1145/1272457.1272467,Conference Paper,"['response time', 'performance modeling', 'Bayesian networks', 'Neural networks', 'comparison study']",
570,Anomaly Detection Using Program Control Flow Graph Mining From Execution Logs,"['Animesh Nandi', 'Atri Mandal', 'Shubham Atreja', 'Gargi B. Dasgupta', 'Subhrajit Bhattacharya']",2016,"We focus on the problem of detecting anomalous run-time behavior of distributed applications from their execution logs. Specifically we mine templates and template sequences from logs to form a control flow graph (cfg) spanning distributed components. This cfg represents the baseline healthy system state and is used to flag deviations from the expected behavior of runtime logs. The novelty in our work stems from the new techniques employed to: (1) overcome the instrumentation requirements or application specific assumptions made in prior log mining approaches, (2) improve the accuracy of mined templates and the cfg in the presence of long parameters and high amount of interleaving respectively, and (3) improve by orders of magnitude the scalability of the cfg mining process in terms of volume of log data that can be processed per day. We evaluate our approach using (a) synthetic log traces and (b) multiple real-world log datasets collected at different layers of application stack. Results demonstrate that our template mining, cfg mining, and anomaly detection algorithms have high accuracy. The distributed implementation of our pipeline is highly scalable and has more than 500 GB/day of log data processing capability even on a 10 low-end VM based (Spark + Hadoop) cluster. We also demonstrate the efficacy of our end-to-end system using a case study with the Openstack VM provisioning system.",https://doi.org/10.1145/2939672.2939712,True,"[{'database': 'ACM', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 97}]",15.0,,,,['failure-detection'],,Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,True,['failure-management'],,,,,,10.1145/2939672.2939712,Conference Paper,"['spark', 'distributed applications', 'application monitoring', 'log analytics', 'execution logs', 'control flow mining', 'template']",
571,Scheduling using improved genetic algorithm in cloud computing for independent tasks,"['Pardeep Kumar', 'Amandeep Verma']",2012,"Cloud computing is a new technology and it is becoming popular day by day because of its great features. In this technology almost everything like hardware, software and platform are provided as a service. These services are charged from users on the pay-per-use bases. A cloud provider in cloud computing provides services on the basis of clients' requests. An important issue in cloud computing is the scheduling of users' requests means how to allocate resources to these requests, so that the requested tasks can be completed in a minimum time according to the user defined time. A good scheduling technique also helps in efficient utilization of the resources. Many scheduling algorithms have been researched like Min-Min, Max-Min, X-Sufferage, Genetic Algorithm, Particle Swarm Optimization etc. In this paper the three scheduling techniques Min-Min, Max-Min and Genetic Algorithm have been discussed and performance metrics of Min-Min and Max-Min have been shown. The performance of the standard Genetic Algorithm and the proposed Improved Genetic Algorithm have been checked against the sample data. A new scheduling idea is also proposed in which Min-Min and Max-Min can be combined in Genetic Algorithm.",https://doi.org/10.1145/2345396.2345420,True,"[{'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 70}]",30.0,,,,['scheduling'],,"Proceedings of the International Conference on Advances in Computing, Communications and Informatics",True,['resource-provisioning'],,True,,,,10.1145/2345396.2345420,Conference Paper,"['min-min', 'max-min', 'cloud computing', 'improved genetic algorithm', 'genetic algorithm']",
572,Renumber Coevolutionary Multiswarm Particle Swarm Optimization for Multi-objective Workflow Scheduling on Cloud Computing Environment,"['Hai-Hao Li', 'Zong-Gan Chen', 'Zhi-Hui Zhan', 'Ke-Jing Du', 'Jun Zhang']",2015,"Resources scheduling is a significant research topic in cloud computing, which is often modeled as a cost-minimization and deadline-constrained workflow scheduling model. This is a constrained single objective problem that to minimize the overall workflow execution cost while meeting deadline constraints. In this paper, we offer a new horizon to convert this single-objective problem to a multi-objective problem and present coevolutionary multiswarm particle swarm optimization (CMPSO) to find the non-dominated solutions with different execute cost and time. Meanwhile, the renumber strategy is adopted in CMPSO to make the learning efficient. CMPSO is compared with a renumber PSO (RNPSO) by setting the execute time in the CMPSO's non-dominated solutions as the deadline constraint of RNPSO. Results show that CMPSO not only offers many non-dominated solutions with different prices and execute time, but also obtains better solution than RNPSO under a same deadline.",https://doi.org/10.1145/2739482.2764632,True,"[{'database': 'ACM', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 90}]",8.0,['particle-swarm'],,,['scheduling'],,Proceedings of the Companion Publication of the 2015 Annual Conference on Genetic and Evolutionary Computation,True,['resource-provisioning'],,,,,,10.1145/2739482.2764632,Conference Paper,"['scheduling', 'renumber', 'particle swarm optimization', 'guiding point', 'cloud computing']",
573,Inferring models of concurrent systems from logs of their behavior with CSight,"['Ivan Beschastnikh', 'Yuriy Brun', 'Michael D. Ernst', 'Arvind Krishnamurthy']",2014,"Concurrent systems are notoriously difficult to debug and understand. A common way of gaining insight into system behavior is to inspect execution logs and documentation. Unfortunately, manual inspection of logs is an arduous process, and documentation is often incomplete and out of sync with the implementation. To provide developers with more insight into concurrent systems, we developed CSight. CSight mines logs of a system's executions to infer a concise and accurate model of that system's behavior, in the form of a communicating finite state machine (CFSM). Engineers can use the inferred CFSM model to understand complex behavior, detect anomalies, debug, and increase confidence in the correctness of their implementations. CSight's only requirement is that the logged events have vector timestamps. We provide a tool that automatically adds vector timestamps to system logs. Our tool prototypes are available at http://synoptic.googlecode.com/. This paper presents algorithms for inferring CFSM models from traces of concurrent systems, proves them correct, provides an implementation, and evaluates the implementation in two ways: by running it on logs from three different networked systems and via a user study that focused on bug finding. Our evaluation finds that CSight infers accurate models that can help developers find bugs.",https://doi.org/10.1145/2568225.2568246,True,"[{'database': 'ACM', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 29}]",118.0,"['rule-mining', 'automaton']",['logs'],['new-method'],['failure-detection'],[],Proceedings of the 36th International Conference on Software Engineering,True,['failure-management'],True,,['anomaly-detection'],True,53.0,10.1145/2568225.2568246,Conference Paper,"['Model inference', 'concurrency', 'CSight', 'log analysis', 'distributed systems']",
574,"Data Storage Management in Cloud Environments: Taxonomy, Survey, and Future Directions","['Yaser Mansouri', 'Adel Nadjaran Toosi', 'Rajkumar Buyya']",2017,"Storage as a Service (StaaS) is a vital component of cloud computing by offering the vision of a virtually infinite pool of storage resources. It supports a variety of cloud-based data store classes in terms of availability, scalability, ACID (Atomicity, Consistency, Isolation, Durability) properties, data models, and price options. Application providers deploy these storage classes across different cloud-based data stores not only to tackle the challenges arising from reliance on a single cloud-based data store but also to obtain higher availability, lower response time, and more cost efficiency. Hence, in this article, we first discuss the key advantages and challenges of data-intensive applications deployed within and across cloud-based data stores. Then, we provide a comprehensive taxonomy that covers key aspects of cloud-based data store: data model, data dispersion, data consistency, data transaction service, and data management cost. Finally, we map various cloud-based data stores projects to our proposed taxonomy to validate the taxonomy and identify areas for future research.",https://doi.org/10.1145/3136623,True,"[{'database': 'ACM', 'search_string': ""((('hidden' AND 'markov') OR ('gaussian' AND 'mixture')) AND 'model') AND ('cloud')"", 'index': 66}]",17.0,,,"['survey', 'discussion']",,,ACM Computing Surveys (CSUR),True,['resource-provisioning'],,,,,,10.1145/3136623,Journal Article,"['data consistency', 'Data management', 'and data management cost', 'data replication', 'data storage', 'transaction service']",
575,SSD Failures in Datacenters: What? When? and Why?,"['Iyswarya Narayanan', 'Di Wang', 'Myeongjae Jeon', 'Bikash Sharma', 'Laura Caulfield', 'Anand Sivasubramaniam', 'Ben Cutler', 'Jie Liu', 'Badriddine Khessib', 'Kushagra Vaid']",2016,"Despite the growing popularity of Solid State Disks (SSDs) in the datacenter, little is known about their reliability characteristics in the field. The little knowledge is mainly vendor supplied, and such information cannot really help understand how SSD failures can manifest and impact the operation of production systems, in order to take appropriate remedial measures. Besides actual failure data and the symptoms exhibited by SSDs before failing, a detailed characterization effort requires wide set of data about factors influencing SSD failures, right from provisioning factors to the operational ones. This paper presents an extensive SSD failure characterization by analyzing a wide spectrum of data from over half a million SSDs that span multiple generations spread across several datacenters which host a wide spectrum of workloads over nearly 3 years. By studying the diverse set of design, provisioning and operational factors on failures, and their symptoms, our work provides the first comprehensive analysis of the what, when and why characteristics of SSD failures in production datacenters.",https://doi.org/10.1145/2928275.2928278,True,"[{'database': 'ACM', 'search_string': ""'logistic regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 81}]",85.0,['random-forest'],,['novel-use'],['failure-prediction'],['ssd'],Proceedings of the 9th ACM International on Systems and Storage Conference,True,['failure-management'],True,,['hardware-failure-prediction'],True,35.0,10.1145/2928275.2928278,Conference Paper,"['solid state drives', 'characterization', 'reliability']",
576,Log Message Anomaly Detection and Classification Using Auto-B/LSTM and Auto-GRU,"['Amir Farzad', 'T. Aaron Gulliver']",2019,"Log messages are now widely used in software systems. They are important for classification as millions of logs are generated each day. Most logs are unstructured which makes classification a challenge. In this paper, Deep Learning (DL) methods called Auto-LSTM, Auto-BLSTM and Auto-GRU are developed for anomaly detection and log classification. These models are used to convert unstructured log data to trained features which is suitable for classification algorithms. They are evaluated using four data sets, namely BGL, Openstack, Thunderbird and IMDB. The first three are popular log data sets while the fourth is a movie review data set which is used for sentiment classification and is used here to show that the models can be generalized to other text classification tasks. The results obtained show that Auto-LSTM, Auto-BLSTM and Auto-GRU perform better than other well-known algorithms.",http://arxiv.org/abs/1911.08744v1,True,"[{'database': 'arxiv', 'search_string': ""'classification' AND ('anomaly detection' OR 'outlier detection')"", 'index': 23}, {'database': 'arxiv', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 6}, {'database': 'arxiv', 'search_string': ""('DL' OR 'deep learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 4}, {'database': 'arxiv', 'search_string': ""('DL' OR 'deep learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 346}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 44}]",0.0,"['autoencoder', 'rnn']",['logs'],['novel-use'],['failure-detection'],['ipenstack'],,True,['failure-management'],,,,,,,,,
577,A Deep Learning based approach to VM behavior identification in cloud systems,"['Matteo Stefanini', 'Riccardo Lancellotti', 'Lorenzo Baraldi', 'Simone Calderara']",2019,"Cloud computing data centers are growing in size and complexity to the point where monitoring and management of the infrastructure become a challenge due to scalability issues. A possible approach to cope with the size of such data centers is to identify VMs exhibiting a similar behavior. Existing literature demonstrated that clustering together VMs that show a similar behavior may improve the scalability of both monitoring andmanagement of a data center. However, available techniques suffer from a trade-off between accuracy and time to achieve this result. Throughout this paper we propose a different approach where, instead of an unsupervised clustering, we rely on classifiers based on deep learning techniques to assigna newly deployed VMs to a cluster of already-known VMs. The two proposed classifiers, namely DeepConv and DeepFFT use a convolution neural network and (in the latter model) exploits Fast Fourier Transformation to classify the VMs. Our proposal is validated using a set of traces describing the behavior of VMs from a realcloud data center. The experiments compare our proposal with state-of-the-art solutions available in literature, demonstrating that our proposal achieve better performance. Furthermore, we show that our solution issignificantly faster than the alternatives as it can produce a perfect classification even with just a few samples of data, making our proposal viable also toclassify on-demand VMs that are characterized by a short life span.",http://arxiv.org/abs/1903.01930v1,True,"[{'database': 'arxiv', 'search_string': ""'classification' AND ('cloud computing')"", 'index': 49}, {'database': 'arxiv', 'search_string': ""'classification' AND ('cloud')"", 'index': 379}, {'database': 'arxiv', 'search_string': ""'classification' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 347}, {'database': 'arxiv', 'search_string': ""'classification' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 5}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 408}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 932}, {'database': 'arxiv', 'search_string': ""'clustering' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 54}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 62}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 173}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 229}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 3}, {'database': 'arxiv', 'search_string': ""('DL' OR 'deep learning') AND ('cloud computing')"", 'index': 27}, {'database': 'arxiv', 'search_string': ""('DL' OR 'deep learning') AND ('cloud')"", 'index': 212}, {'database': 'arxiv', 'search_string': ""('DL' OR 'deep learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 91}, {'database': 'arxiv', 'search_string': ""('DL' OR 'deep learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 3}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud computing')"", 'index': 101}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 392}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 444}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 4}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud computing')"", 'index': 18}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 130}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 120}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 3}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud computing')"", 'index': 41}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 484}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 385}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 3}]",2.0,['cnn'],,"['comparison', 'novel-use']",['failure-detection'],['vm'],,True,['failure-management'],,,['anomaly-detection'],,,,,,
578,Accelerating System Log Processing by Semi-supervised Learning: A Technical Report,"['Guofu Li', 'Pengjia Zhu', 'Zhiyi Chen']",2018,"There is an increasing need for more automated system-log analysis tools for large scale online system in a timely manner. However, conventional way to monitor and classify the log output based on keyword list does not scale well for complex system in which codes contributed by a large group of developers, with diverse ways of encoding the error messages, often with misleading pre-set labels. In this paper, we propose that the design of a large scale online log analysis should follow the ""Least Prior Knowledge Principle"", in which unsupervised or semi-supervised solution with the minimal prior knowledge of the log should be encoded directly. Thereby, we report our experience in designing a two-stage machine learning based method, in which the system logs are regarded as the output of a quasi-natural language, pre-filtered by a perplexity score threshold, and then undergo a fine-grained classification procedure. Tests on empirical data show that our method has obvious advantage regarding to the processing speed and classification accuracy.",http://arxiv.org/abs/1811.01833v1,True,"[{'database': 'arxiv', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 5}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1}]",1.0,,,,['failure-detection'],,,True,['failure-management'],,,['log-enhancement'],,,,,,
579,Deep Convolutional Neural Networks for Anomaly Event Classification on Distributed Systems,"['Jiechao Cheng', 'Rui Ren', 'Lei Wang', 'Jianfeng Zhan']",2017,"The increasing popularity of server usage has brought a plenty of anomaly log events, which have threatened a vast collection of machines. Recognizing and categorizing the anomalous events thereby is a much salient work for our systems, especially the ones generate the massive amount of data and harness it for technology value creation and business development. To assist in focusing on the classification and the prediction of anomaly events, and gaining critical insights from system event records, we propose a novel log preprocessing method which is very effective to filter abundant information and retain critical characteristics. Additionally, a competitive approach for automated classification of anomalous events detected from the distributed system logs with the state-of-the-art deep (Convolutional Neural Network) architectures is proposed in this paper. We measure a series of deep CNN algorithms with varied hyper-parameter combinations by using standard evaluation metrics, the results of our study reveals the advantages and potential capabilities of the proposed deep CNN models for anomaly event classification tasks on real-world systems. The optimal classification precision of our approach is 98.14%, which surpasses the popular traditional machine learning methods.",http://arxiv.org/abs/1710.09052v2,True,"[{'database': 'arxiv', 'search_string': ""'classification' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 64}, {'database': 'arxiv', 'search_string': ""'classification' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 252}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 858}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1435}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 478}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1861}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 166}]",5.0,['cnn'],,,['failure-detection'],,,True,['failure-management'],,,,,,,,,
580,Dogfooding: use IBM Cloud services to monitor IBM Cloud infrastructure,"['William Pourmajidi', 'Andriy Miranskyy', 'John Steinbacher', 'Tony Erwin', 'David Godwin']",2019,"The stability and performance of Cloud platforms are essential as they directly impact customers' satisfaction. Cloud service providers use Cloud monitoring tools to ensure that rendered services match the quality of service requirements indicated in established contracts such as service-level agreements. Given the enormous number of resources that need to be monitored, highly scalable and capable monitoring tools are designed and implemented by Cloud service providers such as Amazon, Google, IBM, and Microsoft. Cloud monitoring tools monitor millions of virtual and physical resources and continuously generate logs for each one of them. Considering that logs magnify any technical issue, they can be used for disaster detection, prevention, and recovery. However, logs are useless if they are not assessed and analyzed promptly. Thus, we argue that the scale of Cloud-generated logs makes it impossible for DevOps teams to analyze them effectively. This implies that one needs to automate the process of monitoring and analysis (e.g., using machine learning and artificial intelligence). If the automation will witness an anomaly in the logs --- it will alert DevOps staff. The automatic anomaly detectors require a reliable and scalable platform for gathering, filtering, and transforming the logs, executing the detector models, and sending out the alerts to the DevOps staff. In this work, we report on implementing a prototype of such a platform based on the 7-layered architecture pattern, which leverages micro-service principles to distribute tasks among highly scalable, resources-efficient modules. The modules interact with each other via an instance of the Publish-Subscribe architectural pattern. The platform is deployed on the IBM Cloud service infrastructure and is used to detect anomalies in logs emitted by the IBM Cloud services, hence the dogfooding.",http://arxiv.org/abs/1907.06094v1,True,"[{'database': 'arxiv', 'search_string': ""'clustering' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 7}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud')"", 'index': 995}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 364}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('remediation' OR 'recovery')"", 'index': 605}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 3}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 101}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 69}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('remediation' OR 'recovery')"", 'index': 202}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 19}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 216}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 436}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 1372}]",0.0,,,,,,,True,['aiops-general'],,,,,,,,,
581,DeCaf: Diagnosing and Triaging Performance Issues in Large-Scale Cloud Services,"['Chetan Bansal', 'Sundararajan Renganathan', 'Ashima Asudani', 'Olivier Midy', 'Mathru Janakiraman']",2019,"Large scale cloud services use Key Performance Indicators (KPIs) for tracking and monitoring performance. They usually have Service Level Objectives (SLOs) baked into the customer agreements which are tied to these KPIs. Dependency failures, code bugs, infrastructure failures, and other problems can cause performance regressions. It is critical to minimize the time and manual effort in diagnosing and triaging such issues to reduce customer impact. Large volume of logs and mixed type of attributes (categorical, continuous) in the logs makes diagnosis of regressions non-trivial. In this paper, we present the design, implementation and experience from building and deploying DeCaf, a system for automated diagnosis and triaging of KPI issues using service logs. It uses machine learning along with pattern mining to help service owners automatically root cause and triage performance issues. We present the learnings and results from case studies on two large scale cloud services in Microsoft where DeCaf successfully diagnosed 10 known and 31 unknown issues. DeCaf also automatically triages the identified issues by leveraging historical data. Our key insights are that for any such diagnosis tool to be effective in practice, it should a) scale to large volumes of service logs and attributes, b) support different types of KPIs and ranking functions, c) be integrated into the DevOps processes.",http://arxiv.org/abs/1910.05339v5,True,"[{'database': 'arxiv', 'search_string': ""'clustering' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 21}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 554}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 2}, {'database': 'arxiv', 'search_string': ""'regression' AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 4}, {'database': 'arxiv', 'search_string': ""'regression' AND ('cloud')"", 'index': 31}, {'database': 'arxiv', 'search_string': ""'regression' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 114}, {'database': 'arxiv', 'search_string': ""'regression' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('DevOps' OR 'site reliability engineering' OR 'SRE')"", 'index': 33}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 568}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 587}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 13}]",0.0,['rule-mining'],"['kpis', 'logs']",,"['root-cause-analysis', 'remediation']",,,True,['failure-management'],,,"['root-cause-diagnosis', 'triage']",,,10.1145/3377813.3381353,,,
582,Real-Time Anomaly Detection for Streaming Analytics,"['Subutai Ahmad', 'Scott Purdy']",2016,"Much of the worlds data is streaming, time-series data, where anomalies give significant information in critical situations. Yet detecting anomalies in streaming data is a difficult task, requiring detectors to process data in real-time, and learn while simultaneously making predictions. We present a novel anomaly detection technique based on an on-line sequence memory algorithm called Hierarchical Temporal Memory (HTM). We show results from a live application that detects anomalies in financial metrics in real-time. We also test the algorithm on NAB, a published benchmark for real-time anomaly detection, where our algorithm achieves best-in-class results.",http://arxiv.org/abs/1607.02480v1,True,"[{'database': 'arxiv', 'search_string': ""'clustering' AND ('anomaly detection' OR 'outlier detection')"", 'index': 29}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 700}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 36}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 193}]",45.0,,,,['failure-detection'],,,True,['failure-management'],,,['anomaly-detection'],,,,,,['real-time']
583,Application of Ontologies in Cloud Computing: The State-Of-The-Art,['Fahim T. Imam'],2016,"This paper presents a systematic survey on existing literature and seminal works relevant to the application of ontologies in different aspects of Cloud computing. Our hypothesis is that ontologies along with their reasoning capabilities can have significant impact on improving various aspects of the Cloud computing phenomena. Ontologies can promote intelligent decision support mechanisms for various Cloud based services. They can also provide effective interoperability among the Cloud based systems and resources. This survey can promote a comprehensive understanding on the roles and significance of ontologies within the overall domain of Cloud Computing. Also, this project can potentially form the basis of new research area and possibilities for both ontology and Cloud computing communities.",http://arxiv.org/abs/1610.02333v1,True,"[{'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 10}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud')"", 'index': 35}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 6}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 50}, {'database': 'arxiv', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud computing')"", 'index': 2}, {'database': 'arxiv', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('cloud')"", 'index': 6}]",2.0,,,['survey'],['resource-consolidation'],,,True,['resource-provisioning'],False,,,,,,,,
584,Predicting Scheduling Failures in the Cloud,"['Mbarka Soualhia', 'Foutse Khomh', 'Sofiene Tahar']",2015,"Cloud Computing has emerged as a key technology to deliver and manage computing, platform, and software services over the Internet. Task scheduling algorithms play an important role in the efficiency of cloud computing services as they aim to reduce the turnaround time of tasks and improve resource utilization. Several task scheduling algorithms have been proposed in the literature for cloud computing systems, the majority relying on the computational complexity of tasks and the distribution of resources. However, several tasks scheduled following these algorithms still fail because of unforeseen changes in the cloud environments. In this paper, using tasks execution and resource utilization data extracted from the execution traces of real world applications at Google, we explore the possibility of predicting the scheduling outcome of a task using statistical models. If we can successfully predict tasks failures, we may be able to reduce the execution time of jobs by rescheduling failed tasks earlier (i.e., before their actual failing time). Our results show that statistical models can predict task failures with a precision up to 97.4%, and a recall up to 96.2%. We simulate the potential benefits of such predictions using the tool kit GloudSim and found that they can improve the number of finished tasks by up to 40%. We also perform a case study using the Hadoop framework of Amazon Elastic MapReduce (EMR) and the jobs of a gene expression correlations analysis study from breast cancer research. We find that when extending the scheduler of Hadoop with our predictive models, the percentage of failed jobs can be reduced by up to 45%, with an overhead of less than 5 minutes.",http://arxiv.org/abs/1507.03562v1,True,"[{'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 179}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud')"", 'index': 1632}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1699}]",5.0,,,,"['failure-prediction', 'scheduling', 'system-failure-prediction']","['hadoop', 'mapreduce']",,True,"['resource-provisioning', 'failure-management']",,,,,,,,,
585,A Self-adaptive Auto-scaling Method for Scientific Applications on HPC Environments and Clouds,"['Kiran Mantripragada', 'Alecio Binotto', 'Leonardo P. Tizzei']",2014,"High intensive computation applications can usually take days to months to finish an execution. During this time, it is common to have variations of the available resources when considering that such hardware is usually shared among a plurality of researchers/departments within an organization. On the other hand, High Performance Clusters can take advantage of Cloud Computing bursting techniques for the execution of applications together with the on-premise resources. In order to meet deadlines, high intensive computational applications can use the Cloud to boost their performance when they are data and task parallel. This article presents an ongoing work towards the use of extended resources of an HPC execution platform together with Cloud. We propose an unified view of such heterogeneous environments and a method that monitors, predicts the application execution time, and dynamically shifts part of the domain -- previously running in local HPC hardware -- to be computed on the Cloud, meeting then a specific deadline. The method is exemplified along with a seismic application that, at runtime, adapts itself to move part of the processing to the Cloud (in a movement called bursting) and also auto-scales (the moved part) over cloud nodes. Our preliminary results show that there is an expected overhead for performing this movement and for synchronizing results, but our outcomes demonstrate it is an important feature for meeting deadlines in the case an on-premise cluster is overloaded or cannot provide the capacity needed for a particular project.",http://arxiv.org/abs/1412.6392v3,True,"[{'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 411}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud')"", 'index': 411}]",7.0,,,,,,,True,['resource-provisioning'],,,,,,,,,
586,Hybrid Genetic Algorithm for Cloud Computing Applications,"['Saeed Javanmardi', 'Mohammad Shojafar', 'Danilo Amendola', 'Nicola Cordeschi', 'Hongbo Liu', 'Ajith Abraham']",2014,"In this paper with the aid of genetic algorithm and fuzzy theory, we present a hybrid job scheduling approach, which considers the load balancing of the system and reduces total execution time and execution cost. We try to modify the standard Genetic algorithm and to reduce the iteration of creating population with the aid of fuzzy theory. The main goal of this research is to assign the jobs to the resources with considering the VM MIPS and length of jobs. The new algorithm assigns the jobs to the resources with considering the job length and resources capacities. We evaluate the performance of our approach with some famous cloud scheduling models. The results of the experiments show the efficiency of the proposed approach in term of execution time, execution cost and average Degree of Imbalance (DI).",http://arxiv.org/abs/1404.5528v1,True,"[{'database': 'arxiv', 'search_string': ""'clustering' AND ('cloud computing')"", 'index': 480}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud computing')"", 'index': 44}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('cloud')"", 'index': 136}]",88.0,['genetic-programming'],,,['scheduling'],,,True,['resource-provisioning'],,,,,,,,,
587,Fault Detection Engine in Intelligent Predictive Analytics Platform for DCIM,"['Bodhisattwa Prasad Majumder', 'Ayan Sengupta', 'Sajal jain', 'Parikshit Bhaduri']",2016,"With the advancement of huge data generation and data handling capability, Machine Learning and Probabilistic modelling enables an immense opportunity to employ predictive analytics platform in high security critical industries namely data centers, electricity grids, utilities, airport etc. where downtime minimization is one of the primary objectives. This paper proposes a novel, complete architecture of an intelligent predictive analytics platform, Fault Engine, for huge device network connected with electrical/information flow. Three unique modules, here proposed, seamlessly integrate with available technology stack of data handling and connect with middleware to produce online intelligent prediction in critical failure scenarios. The Markov Failure module predicts the severity of a failure along with survival probability of a device at any given instances. The Root Cause Analysis model indicates probable devices as potential root cause employing Bayesian probability assignment and topological sort. Finally, a community detection algorithm produces correlated clusters of device in terms of failure probability which will further narrow down the search space of finding route cause. The whole Engine has been tested with different size of network with simulated failure environments and shows its potential to be scalable in real-time implementation.",http://arxiv.org/abs/1610.04872v1,True,"[{'database': 'arxiv', 'search_string': ""'clustering' AND ('fault detection' OR 'failure detection')"", 'index': 22}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 6}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 19}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 6}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('fault detection' OR 'failure detection')"", 'index': 52}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 15}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('fault detection' OR 'failure detection')"", 'index': 23}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 3}]",1.0,,,['new-method'],"['root-cause-analysis', 'failure-detection']",,,True,['failure-management'],,,['fault-localization'],,,,,,
588,Robust Data Preprocessing for Machine-Learning-Based Disk Failure Prediction in Cloud Production Environments,"['Shujie Han', 'Jun Wu', 'Erci Xu', 'Cheng He', 'Patrick P. C. Lee', 'Yi Qiang', 'Qixing Zheng', 'Tao Huang', 'Zixi Huang', 'Rui Li']",2019,"To provide proactive fault tolerance for modern cloud data centers, extensive studies have proposed machine learning (ML) approaches to predict imminent disk failures for early remedy and evaluated their approaches directly on public datasets (e.g., Backblaze SMART logs). However, in real-world production environments, the data quality is imperfect (e.g., inaccurate labeling, missing data samples, and complex failure types), thereby degrading the prediction accuracy. We present RODMAN, a robust data preprocessing pipeline that refines data samples before feeding them into ML models. We start with a large-scale trace-driven study of over three million disks from Alibaba Cloud's data centers, and motivate the practical challenges in ML-based disk failure prediction. We then design RODMAN with three data preprocessing echniques, namely failure-type filtering, spline-based data filling, and automated pre-failure backtracking, that are applicable for general ML models. Evaluation on both the Alibaba and Backblaze datasets shows that RODMAN improves the prediction accuracy compared to without data preprocessing under various settings.",http://arxiv.org/abs/1912.09722v1,True,"[{'database': 'arxiv', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 10}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1798}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('remediation' OR 'recovery')"", 'index': 228}, {'database': 'arxiv', 'search_string': ""'clustering' AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1215}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 6}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('fault prediction' OR 'failure prediction')"", 'index': 2}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 3}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('remediation' OR 'recovery')"", 'index': 1}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('tracing' OR 'trace' OR 'traces')"", 'index': 1}]",1.0,,['host-metrics'],['new-method'],['failure-prediction'],['hard-drive'],,True,['failure-management'],,,['hardware-failure-prediction'],,,,,,
589,Predicting Failures in Multi-Tier Distributed Systems,"['Leonardo Mariani', 'Mauro Pezzè', 'Oliviero Riganelli', 'Rui Xin']",2019,"Many applications are implemented as multi-tier software systems, and are executed on distributed infrastructures, like cloud infrastructures, to benefit from the cost reduction that derives from dynamically allocating resources on-demand. In these systems, failures are becoming the norm rather than the exception, and predicting their occurrence, as well as locating the responsible faults, are essential enablers of preventive and corrective actions that can mitigate the impact of failures, and significantly improve the dependability of the systems. Current failure prediction approaches suffer either from false positives or limited accuracy, and do not produce enough information to effectively locate the responsible faults. In this paper, we present PreMiSE, a lightweight and precise approach to predict failures and locate the corresponding faults in multi-tier distributed systems. PreMiSE blends anomaly-based and signature-based techniques to identify multi-tier failures that impact on performance indicators, with high precision and low false positive rate. The experimental results that we obtained on a Cloud-based IP Multimedia Subsystem indicate that PreMiSE can indeed predict and locate possible failure occurrences with high precision and low overhead.",http://arxiv.org/abs/1911.09561v1,True,"[{'database': 'arxiv', 'search_string': ""'clustering' AND ('fault prediction' OR 'failure prediction')"", 'index': 17}]",1.0,,['kpis'],"['new-method', 'discussion']",['failure-prediction'],,,True,['failure-management'],,,['system-failure-prediction'],,,,,,['online']
590,"Towards Operator-less Data Centers Through Data-Driven, Predictive, Proactive Autonomics","['Alina Sîrbu', 'Ozalp Babaoglu']",2016,"Continued reliance on human operators for managing data centers is a major impediment for them from ever reaching extreme dimensions. Large computer systems in general, and data centers in particular, will ultimately be managed using predictive computational and executable models obtained through data-science tools, and at that point, the intervention of humans will be limited to setting high-level goals and policies rather than performing low-level operations. Data-driven autonomics, where management and control are based on holistic predictive models that are built and updated using live data, opens one possible path towards limiting the role of operators in data centers. In this paper, we present a data-science study of a public Google dataset collected in a 12K-node cluster with the goal of building and evaluating predictive models for node failures. Our results support the practicality of a data-driven approach by showing the effectiveness of predictive models based on data found in typical data center logs. We use BigQuery, the big data SQL platform from the Google Cloud suite, to process massive amounts of data and generate a rich feature set characterizing node state over time. We describe how an ensemble classifier can be built out of many Random Forest classifiers each trained on these features, to predict if nodes will fail in a future 24-hour window. Our evaluation reveals that if we limit false positive rates to 5%, we can achieve true positive rates between 27% and 88% with precision varying between 50% and 72%.This level of performance allows us to recover large fraction of jobs' executions (by redirecting them to other nodes when a failure of the present node is predicted) that would otherwise have been wasted due to failures. [...]",http://arxiv.org/abs/1606.04456v1,True,"[{'database': 'arxiv', 'search_string': ""'clustering' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 48}, {'database': 'arxiv', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 372}, {'database': 'arxiv', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('log' OR 'logs' OR 'log analysis')"", 'index': 1389}, {'database': 'arxiv', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 8}]",12.0,,,,['failure-prediction'],,,True,['failure-management'],,,['system-failure-prediction'],,,10.1007/s10586-016-0564-y,,,
591,Modelling Energy Consumption based on Resource Utilization,"['Lucas Venezian Povoa', 'Cesar Marcondes', 'Hermes Senger']",2017,"Power management is an expensive and important issue for large computational infrastructures such as datacenters, large clusters, and computational grids. However, measuring energy consumption of scalable systems may be impractical due to both cost and complexity for deploying power metering devices on a large number of machines. In this paper, we propose the use of information about resource utilization (e.g. processor, memory, disk operations, and network traffic) as proxies for estimating power consumption. We employ machine learning techniques to estimate power consumption using such information which are provided by common operating systems. Experiments with linear regression, regression tree, and multilayer perceptron on data from different hardware resulted into a model with 99.94\% of accuracy and 6.32 watts of error in the best case.",http://arxiv.org/abs/1709.06076v1,True,"[{'database': 'arxiv', 'search_string': ""'clustering' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 56}, {'database': 'arxiv', 'search_string': ""'regression' AND (('datacenter' OR 'data center') AND 'management')"", 'index': 1}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 2}, {'database': 'arxiv', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 3}]",0.0,,,,['power-management'],,,True,['resource-provisioning'],,,,,,,,,
592,uPredict: A User-Level Profiler-Based Predictive Framework for Single VM Applications in Multi-Tenant Clouds,"['Hamidreza Moradi', 'Wei Wang', 'Amanda Fernandez', 'Dakai Zhu']",2019,"Most existing studies on performance prediction for virtual machines (VMs) in multi-tenant clouds are at system level and generally require access to performance counters in Hypervisors. In this work, we propose uPredict, a user-level profiler-based performance predictive framework for single-VM applications in multi-tenant clouds. Here, three micro-benchmarks are specially devised to assess the contention of CPUs, memory and disks in a VM, respectively. Based on measured performance of an application and micro-benchmarks, the application and VM-specific predictive models can be derived by exploiting various regression and neural network based techniques. These models can then be used to predict the application's performance using the in-situ profiled resource contention with the micro-benchmarks. We evaluated uPredict extensively with representative benchmarks from PARSEC, NAS Parallel Benchmarks and CloudSuite, on both a private cloud and two public clouds. The results show that the average prediction errors are between 9.8% to 17% for various predictive models on the private cloud with high resource contention, while the errors are within 4% on public clouds. A smart load-balancing scheme powered by uPredict is presented and can effectively reduce the execution and turnaround times of the considered application by 19% and 10%, respectively.",http://arxiv.org/abs/1908.04491v1,True,"[{'database': 'arxiv', 'search_string': ""'regression' AND ('cloud')"", 'index': 19}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('cloud')"", 'index': 120}]",0.0,"['multilayer-perceptron', 'linear-regression']",,,['workload-prediction'],['vm'],,True,['resource-provisioning'],,,,,,,,,
593,Feedforward Neural Network for Time Series Anomaly Detection,"['Zhang Rong', 'Dong Shandong', 'Nie Xin', 'Xiao Shiguang']",2018,"Time series anomaly detection is usually formulated as finding outlier data points relative to some usual data, which is also an important problem in industry and academia. To ensure systems working stably, internet companies, banks and other companies need to monitor time series, which is called KPI (Key Performance Indicators), such as CPU used, number of orders, number of online users and so on. However, millions of time series have several shapes (e.g. seasonal KPIs, KPIs of timed tasks and KPIs of CPU used), so that it is very difficult to use a simple statistical model to detect anomaly for all kinds of time series. Although some anomaly detectors have developed many years and some supervised models are also available in this field, we find many methods have their own disadvantages. In this paper, we present our system, which is based on deep feedforward neural network and detect anomaly points of time series. The main difference between our system and other systems based on supervised models is that we do not need feature engineering of time series to train deep feedforward neural network in our system, which is essentially an end-to-end system.",http://arxiv.org/abs/1812.08389v2,True,"[{'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('anomaly detection' OR 'outlier detection')"", 'index': 93}, {'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 4}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 333}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 3}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 161}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 65}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 1}]",1.0,['multilayer-perceptron'],"['kpis', 'host-metrics']",,['failure-detection'],,,True,['failure-management'],,,,,,,,,
594,Autonomous Fault Detection in Self-Healing Systems using Restricted Boltzmann Machines,"['Chris Schneider', 'Adam Barker', 'Simon Dobson']",2015,"Autonomously detecting and recovering from faults is one approach for reducing the operational complexity and costs associated with managing computing environments. We present a novel methodology for autonomously generating investigation leads that help identify systems faults, and extends our previous work in this area by leveraging Restricted Boltzmann Machines (RBMs) and contrastive divergence learning to analyse changes in historical feature data. This allows us to heuristically identify the root cause of a fault, and demonstrate an improvement to the state of the art by showing feature data can be predicted heuristically beyond a single instance to include entire sequences of information.",http://arxiv.org/abs/1501.01501v1,True,"[{'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault detection' OR 'failure detection')"", 'index': 14}]",10.0,,,,['failure-detection'],,,True,['failure-management'],,,,,,,,,
595,Learning Tractable Probabilistic Models for Fault Localization,"['Aniruddh Nath', 'Pedro Domingos']",2015,"In recent years, several probabilistic techniques have been applied to various debugging problems. However, most existing probabilistic debugging systems use relatively simple statistical models, and fail to generalize across multiple programs. In this work, we propose Tractable Fault Localization Models (TFLMs) that can be learned from data, and probabilistically infer the location of the bug. While most previous statistical debugging methods generalize over many executions of a single program, TFLMs are trained on a corpus of previously seen buggy programs, and learn to identify recurring patterns of bugs. Widely-used fault localization techniques such as TARANTULA evaluate the suspiciousness of each line in isolation; in contrast, a TFLM defines a joint probability distribution over buggy indicator variables for each line. Joint distributions with rich dependency structure are often computationally intractable; TFLMs avoid this by exploiting recent developments in tractable probabilistic models (specifically, Relational SPNs). Further, TFLMs can incorporate additional sources of information, including coverage-based features such as TARANTULA. We evaluate the fault localization performance of TFLMs that include TARANTULA scores as features in the probabilistic model. Our study shows that the learned TFLMs isolate bugs more effectively than previous statistical methods or using TARANTULA directly.",http://arxiv.org/abs/1507.01698v1,True,"[{'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('fault localization' OR 'failure localization')"", 'index': 4}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('fault localization' OR 'failure localization')"", 'index': 3}, {'database': 'arxiv', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('fault localization' OR 'failure localization')"", 'index': 5}]",19.0,,,,['root-cause-analysis'],,Thirtieth AAAI Conference on Artificial Intelligence,True,['failure-management'],,,['fault-localization'],,,,,,
596,Using Abduction in Markov Logic Networks for Root Cause Analysis,"['Joerg Schoenfisch', 'Janno von St\x7fulpnagel', 'Jens Ortmann', 'Christian Meilicke', 'Heiner Stuckenschmidt']",2015,"IT infrastructure is a crucial part in most of today's business operations. High availability and reliability, and short response times to outages are essential. Thus a high amount of tool support and automation in risk management is desirable to decrease outages. We propose a new approach for calculating the root cause for an observed failure in an IT infrastructure. Our approach is based on Abduction in Markov Logic Networks. Abduction aims to find an explanation for a given observation in the light of some background knowledge. In failure diagnosis, the explanation corresponds to the root cause, the observation to the failure of a component, and the background knowledge to the dependency graph extended by potential risks. We apply a method to extend a Markov Logic Network in order to conduct abductive reasoning, which is not naturally supported in this formalism. Our approach exhibits a high amount of reusability and enables users without specific knowledge of a concrete infrastructure to gain viable insights in the case of an incident. We implemented the method in a tool and illustrate its suitability for root cause analysis by applying it to a sample scenario.",http://arxiv.org/abs/1511.05719v1,True,"[{'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 2}, {'database': 'arxiv', 'search_string': ""('inference' OR 'logic' OR 'reasoning') AND ('root-cause analysis' OR 'root cause analysis')"", 'index': 2}]",5.0,"['reasoning', 'markov-model']",,,['root-cause-analysis'],,,True,['failure-management'],,,['root-cause-diagnosis'],,,,,,
597,DeepConfig: Automating Data Center Network Topologies Management with Machine Learning,"['Christopher Streiffer', 'Huan Chen', 'Theophilus Benson', 'Asim Kadav']",2017,"In recent years, many techniques have been developed to improve the performance and efficiency of data center networks. While these techniques provide high accuracy, they are often designed using heuristics that leverage domain-specific properties of the workload or hardware. In this vision paper, we argue that many data center networking techniques, e.g., routing, topology augmentation, energy savings, with diverse goals actually share design and architectural similarity. We present a design for developing general intermediate representations of network topologies using deep learning that is amenable to solving classes of data center problems. We develop a framework, DeepConfig, that simplifies the processing of configuring and training deep learning agents that use the intermediate representation to learns different tasks. To illustrate the strength of our approach, we configured, implemented, and evaluated a DeepConfig-Agent that tackles the data center topology augmentation problem. Our initial results are promising --- DeepConfig performs comparably to the optimal.",http://arxiv.org/abs/1712.03890v1,True,"[{'database': 'arxiv', 'search_string': ""('AI' OR 'artificial intelligence') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 1}, {'database': 'arxiv', 'search_string': ""('DL' OR 'deep learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 1}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND (('datacenter' OR 'data center') AND 'management')"", 'index': 1}]",2.0,,,,['configuration'],,,True,['resource-provisioning'],,,,,,,,,
598,Sequential VAE-LSTM for Anomaly Detection on Time Series,"['Run-Qing Chen', 'Guang-Hui Shi', 'Wan-Lei Zhao', 'Chang-Hui Liang']",2019,"In order to support stable web-based applications and services, anomalies on the IT performance status have to be detected timely. Moreover, the performance trend across the time series should be predicted. In this paper, we propose SeqVL (Sequential VAE-LSTM), a neural network model based on both VAE (Variational Auto-Encoder) and LSTM (Long Short-Term Memory). This work is the first attempt to integrate unsupervised anomaly detection and trend prediction under one framework. Moreover, this model performs considerably better on detection and prediction than VAE and LSTM work alone. On unsupervised anomaly detection, SeqVL achieves competitive experimental results compared with other state-of-the-art methods on public datasets. On trend prediction, SeqVL outperforms several classic time series prediction models in the experiments of the public dataset.",http://arxiv.org/abs/1910.03818v2,True,"[{'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 191}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 107}, {'database': 'arxiv', 'search_string': ""(('bayesian' OR 'neural') AND 'network') AND ('anomaly detection' OR 'outlier detection')"", 'index': 72}]",0.0,"['autoencoder', 'rnn']",,,['failure-detection'],['web-server'],,True,['failure-management'],,,['anomaly-detection'],,,,,,
599,Automatic Anomaly Detection in the Cloud Via Statistical Learning,"['Jordan Hochenbaum', 'Owen S. Vallis', 'Arun Kejariwal']",2017,"Performance and high availability have become increasingly important drivers, amongst other drivers, for user retention in the context of web services such as social networks, and web search. Exogenic and/or endogenic factors often give rise to anomalies, making it very challenging to maintain high availability, while also delivering high performance. Given that service-oriented architectures (SOA) typically have a large number of services, with each service having a large set of metrics, automatic detection of anomalies is non-trivial. Although there exists a large body of prior research in anomaly detection, existing techniques are not applicable in the context of social network data, owing to the inherent seasonal and trend components in the time series data. To this end, we developed two novel statistical techniques for automatically detecting anomalies in cloud infrastructure data. Specifically, the techniques employ statistical learning to detect anomalies in both application, and system metrics. Seasonal decomposition is employed to filter the trend and seasonal components of the time series, followed by the use of robust statistical metrics -- median and median absolute deviation (MAD) -- to accurately detect anomalies, even in the presence of seasonal spikes. We demonstrate the efficacy of the proposed techniques from three different perspectives, viz., capacity planning, user behavior, and supervised learning. In particular, we used production data for evaluation, and we report Precision, Recall, and F-measure in each case.",http://arxiv.org/abs/1704.07706v1,True,"[{'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 571}, {'database': 'arxiv', 'search_string': ""('ML' OR 'machine learning') AND ('cloud')"", 'index': 665}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('anomaly detection' OR 'outlier detection')"", 'index': 199}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 135}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 477}]",64.0,,,['discussion'],['failure-detection'],,,True,['failure-management'],,True,['anomaly-detection'],,,,,,
600,WiSeDB: A Learning-based Workload Management Advisor for Cloud Databases,"['Ryan Marcus', 'Olga Papaemmanouil']",2016,"Workload management for cloud databases must deal with the tasks of resource provisioning, query placement and query scheduling in a manner that meets the application's performance goals while minimizing the cost of using cloud resources. Existing solutions have approached these three challenges in isolation, and with only a particular type of performance goal in mind. In this paper, we introduce WiSeDB, a learning-based framework for generating holistic workload management solutions customized to application-defined performance metrics and workload characteristics. Our approach relies on supervised learning to train cost-effective decision tree models for guiding query placement, scheduling, and resource provisioning decisions. Applications can use these models for both batch and online scheduling of incoming workloads. A unique feature of our system is that it can adapt its offline model to stricter/looser performance goals with minimal re-training. This allows us to present alternative workload management strategies that address the typical performance vs. cost trade-off of cloud services. Experimental results show that our approach has very low training overhead while offering low cost strategies for a variety of performance goals and workload characteristics.",http://arxiv.org/abs/1601.08221v3,True,"[{'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('cloud')"", 'index': 57}, {'database': 'arxiv', 'search_string': ""('supervised' OR 'unsupervised' OR 'semi-supervised' OR 'reinforcement') AND ('learning') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 541}, {'database': 'arxiv', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('cloud')"", 'index': 62}, {'database': 'arxiv', 'search_string': ""('tree' OR 'tree-based' OR 'trees' OR 'forest') AND ('metrics' OR 'KPI' OR 'key performance indicator')"", 'index': 971}]",30.0,,,,['workload-prediction'],,,True,['resource-provisioning'],,True,,,,10.14778/2977797.2977804,,,
601,A Review on Software Fault Detection and Prevention Mechanism in Software Development Activities,"['Dhanalaxmi, B', 'Naidu, G Apparao', 'Anuradha, K']",2015,"The need of distributed and complex commercial applications in enterprise demands error free and quality application systems. This makes it extremely important in software development to develop quality and fault free software. It is also very important to design reliable and easy to maintain as it involves a lot of human efforts, cost and time during software life cycle. A software development process performs various activities to minimize the faults such as fault prediction, detection, prevention and correction. This paper presents …",https://www.semanticscholar.org/paper/A-Review-on-Software-Fault-Detection-and-Prevention-Dhanalaxmi-Naidu/234f074baebded71b0413cff8bb738281d8f8b46,True,,8.0,[],"['source-code', 'events']",['survey'],"['failure-detection', 'failure-prevention']",['software'],,True,['failure-management'],True,,"['anomaly-detection', 'software-defect-prediction']",,,,,,
602,A Survey on Software Fault Localization,"['W. Eric Wong', 'Ruizhi Gao', 'Yihao Li', 'Rui Abreu', 'Franz Wotawa']",2016,"Software fault localization, the act of identifying the locations of faults in a program, is widely recognized to be one of the most tedious, time consuming, and expensive – yet equally critical – activities in program debugging. Due to the increasing scale and complexity of software today, manually locating faults when failures occur is rapidly becoming infeasible, and consequently, there is a strong demand for techniques that can guide software developers to the locations of faults in a program with minimal human intervention. This demand in turn has fueled the proposal and development of a broad spectrum of fault localization techniques, each of which aims to streamline the fault localization process and make it more effective by attacking the problem in a unique way. In this article, we catalog and provide a comprehensive overview of such techniques and discuss key issues and concerns that are pertinent to software fault localization as a whole.",https://doi.org/10.1109/TSE.2016.2521368,True,,347.0,[],,['survey'],['root-cause-analysis'],['software'],IEEE Transactions on Software Engineering,True,['failure-management'],True,True,['fault-localization'],,,10.1109/TSE.2016.2521368,,,
603,Magpie: online modelling and performance-aware systems,"['Paul Barham', 'Rebecca Isaacs', 'Richard Mortier', 'Dushyanth Narayanan']",2003,"Understanding the performance of distributed systems requires correlation of thousands of interactions between numerous components -- a task best left to a computer. Today's systems provide voluminous traces from each component but do not synthesise the data into concise models of system performance. We argue that online performance modelling should be a ubiquitous operating system service and outline several uses including performance debugging, capacity planning, system tuning and anomaly detection. We describe the Magpie modelling service which collates detailed traces from multiple machines in an e-commerce site, extracts request-specific audit trails, and constructs probabilistic models of request behaviour. A feasibility study evaluates the approach using an offline demonstrator. Results show that the approach is promising, but that there are many challenges to building a truly ubiquitious, online modelling infrastructure.",https://dl.acm.org/doi/10.5555/1251054.1251069,True,,250.0,"['language-modeling', 'clustering']","['host-metrics', 'logs', 'network-metrics']",,['failure-detection'],,HOTOS'03: Proceedings of the 9th conference on Hot Topics in Operating Systems - Volume 9,True,['failure-management'],True,True,['anomaly-detection'],True,50.0,10.5555/1251054.1251069,,,
604,Improving service availability of cloud systems by predicting disk error,"['Yong Xu', 'Kaixin Sui', 'Randolph Yao', 'Hongyu Zhang', 'Qingwei Lin', 'Yingnong Dang', 'Peng Li', 'Keceng Jiang', 'Wenchi Zhang', 'Jian-Guang Lou', 'Murali Chintalapati', 'Dongmei Zhang']",2018,"High service availability is crucial for cloud systems. A typical cloud system uses a large number of physical hard disk drives. Disk errors are one of the most important reasons that lead to service unavailability. Disk error (such as sector error and latency error) can be seen as a form of gray failure, which are fairly subtle failures that are hard to be detected, even when applications are afflicted by them. In this paper, we propose to predict disk errors proactively before they cause more severe damage to the cloud system. The ability to predict faulty disks enables the live migration of existing virtual machines and allocation of new virtual machines to the healthy disks, therefore improving service availability. To build an accurate online prediction model, we utilize both disk-level sensor (SMART) data as well as system-level signals. We develop a cost-sensitive ranking-based machine learning model that can learn the characteristics of faulty disks in the past and rank the disks based on their error-proneness in the near future. We evaluate our approach using real-world data collected from a production cloud system. The results confirm that the proposed approach is effective and outperforms related methods. Furthermore, we have successfully applied the proposed approach to improve service availability of Microsoft Azure.",https://dl.acm.org/doi/10.5555/3277355.3277402,True,,21.0,,,,['failure-prediction'],['hard-drive'],USENIX ATC '18: Proceedings of the 2018 USENIX Conference on Usenix Annual Technical Conference,True,['failure-management'],False,True,['hardware-failure-prediction'],,,10.5555/3277355.3277402,,,
605,Mining Historical Issue Repositories to Heal Large-Scale Online Service Systems,"['Rui Ding', 'Qiang Fu', 'Jian Guang Lou', 'Qingwei Lin', 'Dongmei Zhang', 'Tao Xie']",2014,"Online service systems have been increasingly popular and important nowadays. Reducing the MTTR (Mean Time to Restore) of a service remains one of the most important steps to assure the user-perceived availability of the service. To reduce the MTTR, a common practice is to restore the service by identifying and applying an appropriate healing action. In this paper, we present an automated mining-based approach for suggesting an appropriate healing action for a given new issue. Our approach suggests an appropriate healing action by adapting healing actions from the retrieved similar historical issues. We have applied our approach to a real-world and large-scale product online service. The studies on 243 real issues of the service show that our approach can effectively suggest appropriate healing actions (with 87% accuracy) to reduce the MTTR of the service. In addition, according to issue characteristics, we further study and categorize issues where automatic healing suggestion faces difficulties.",https://doi.org/10.1109/DSN.2014.39,True,,9.0,,['logs'],,['remediation'],,DSN '14: Proceedings of the 2014 44th Annual IEEE/IFIP International Conference on Dependable Systems and Networks,True,['failure-management'],,True,,,,10.1109/DSN.2014.39,,,
606,Healing online service systems via mining historical issue repositories,"['Rui Ding', 'Qiang Fu', 'Jian-Guang Lou', 'Qingwei Lin', 'Dongmei Zhang', 'Jiajun Shen', 'Tao Xie']",2012,"Online service systems have been increasingly popular and important nowadays, with an increasing demand on the availability of services provided by these systems, while significant efforts have been made to strive for keeping services up continuously. Therefore, reducing the MTTR (Mean Time to Restore) of a service remains the most important step to assure the user-perceived availability of the service. To reduce the MTTR, a common practice is to restore the service by identifying and applying an appropriate healing action (i.e., a temporary workaround action such as rebooting a SQL machine). However, manually identifying an appropriate healing action for a given new issue (such as service down) is typically time consuming and error prone. To address this challenge, in this paper, we present an automated mining-based approach for suggesting an appropriate healing action for a given new issue. Our approach generates signatures of an issue from its corresponding transaction logs and then retrieves historical issues from a historical issue repository. Finally, our approach suggests an appropriate healing action by adapting healing actions for the retrieved historical issues. We have implemented a healing suggestion system for our approach and applied it to a real-world product online service that serves millions of online customers globally. The studies on 77 incidents (severe issues) over 3 months showed that our approach can effectively provide appropriate healing actions to reduce the MTTR of the service.",https://dl.acm.org/doi/10.1145/2351676.2351735,True,,23.0,,['logs'],,['remediation'],,ASE 2012: Proceedings of the 27th IEEE/ACM International Conference on Automated Software Engineering,True,['failure-management'],,True,,,,10.1145/2351676.2351735,,,
607,On Black-Box Monitoring Techniques for Multi-Component Services,"['Filipe, Ricardo', 'Correia, Jaime', 'Araujo, Filipe', 'Cardoso, Jorge']",2018,"Despite the advantages of microservice and function-oriented architectures, there is an increase in complexity to monitor such highly dynamic systems. In this paper, we analyze two distinct methods to tackle the monitoring problem in a system with reduced instrumentation. Our goal is to understand the feasibility of such approach with one specific driver: simplicity. We aim to determine the extent to which it is possible to characterize the state of two generic tandem processes, using as little information as possible. To answer this …",https://doi.org/10.1109/NCA.2018.8548336,True,,1.0,,,,['root-cause-analysis'],,2018 IEEE 17th International Symposium on Network Computing and Applications (NCA),True,['failure-management'],,,,,,10.1109/NCA.2018.8548336,,,
608,Unsupervised Anomaly Event Detection for Cloud Monitoring Using Online Arima,"['Schmidt, Florian', 'Suri-Payer, Florian', 'Gulenko, Anton', 'Wallschl{\\""a}ger, Marcel', 'Acker, Alexander', 'Kao, Odej']",2018,"Virtualization offers cost efficient usage of digital resources. Thus, dedicated hardware solutions are transferred into virtualized services running in the cloud. Such softwarization of hardware is for example the IP multimedia subsystems, which telecommunication system providers currently move to the cloud. The dedicated hardware solutions provided a reliability of 99.999% in the past, but the virtualized services come with higher complexity due to the fragile computation stack and cannot provide such high requirements. Zero touch …",https://doi.org/10.1109/UCC-Companion.2018.00037,True,,1.0,['autoregression'],,,['failure-detection'],,2018 IEEE/ACM International Conference on Utility and Cloud Computing Companion (UCC Companion),True,['failure-management'],,,,,,10.1109/UCC-Companion.2018.00037,,,
609,"Pinpoint: problem determination in large, dynamic Internet services","['Chen, Mike Y', 'Kiciman, Emre', 'Fratkin, Eugene', 'Fox, Armando', 'Brewer, Eric']",2002,"Traditional problem determination techniques rely on static dependency models that are difficult to generate accurately in today's large, distributed, and dynamic application environments such as e-commerce systems. We present a dynamic analysis methodology that automates problem determination in these environments by 1) coarse-grained tagging of numerous real client requests as they travel through the system and 2) using data mining techniques to correlate the believed failures and successes of these requests to determine …",https://doi.org/10.1109/DSN.2002.1029005,True,,957.0,"['similarity-matching', 'clustering']","['traces', 'logs']",['novel-use'],"['root-cause-analysis', 'failure-detection']",['web-server'],Proceedings International Conference on Dependable Systems and Networks,True,['failure-management'],True,True,['anomaly-detection'],True,78.0,10.1109/DSN.2002.1029005,,,
610,Graph-based root cause analysis for service-oriented and microservice architectures,"[""Brand{\\'o}n, {\\'A}lvaro"", ""Sol{\\'e}, Marc"", ""Hu{\\'e}lamo, Alberto"", 'Solans, David', ""P{\\'e}rez, Mar{\\'\\i}a S"", ""Munt{\\'e}s-Mulero, Victor""]",2020,"Service-oriented architectures and microservices define two ways of designing software with the aim of dividing an application into loosely-coupled services that communicate among each other. This translates into rapid development, where each service is developed and deployed by small teams, enabling continuous shipping of new features and fast-evolving applications. However, the underlying complexity of this type of architecture can hinder observability and maintenance by the user. In particular, identifying the root cause of an anomaly detected in the application can be a difficult and time-consuming task, considering the numerous services and connections to be examined. In this work, we present a root cause analysis framework, based on graph representations of these architectures. The graphs can be used to compare any anomalous situation that happens in the system with a library of anomalous graphs that serves as a knowledge base for the user troubleshooting those anomalies. We use the Grid’5000 testbed to deploy three different architectures and inject a set of anomalies. The results show how our graph-based approach is 19.41% more effective than a machine learning method that does not take into account the relationship between elements.",https://doi.org/10.1016/j.jss.2019.110432,True,,0.0,['graph-mining'],,"['new-method', 'comparison']",['root-cause-analysis'],,,True,['failure-management'],,,['root-cause-diagnosis'],,,10.1016/j.jss.2019.110432,,,
611,Execution Anomaly Detection in Distributed Systems through Unstructured Log Analysis,"['Qiang Fu', 'Jian-Guang Lou', 'Yi Wang', 'Jiang Li']",2009,"Detection of execution anomalies is very important for the maintenance, development, and performance refinement of large scale distributed systems. Execution anomalies include both work flow errors and low performance problems. People often use system logs produced by distributed systems for troubleshooting and problem diagnosis. However, manually inspecting system logs to detect anomalies is unfeasible due to the increasing scale and complexity of distributed systems. Therefore, there is a great demand for automatic anomalies detection techniques based on log analysis. In this paper, we propose an unstructured log analysis technique for anomalies detection. In the technique, we propose a novel algorithm to convert free form text messages in log files to log keys without heavily relying on application specific knowledge. The log keys correspond to the log-print statements in the source code which can provide cues of system execution behavior. After converting log messages to log keys, we learn a Finite State Automaton (FSA) from training log sequences to present the normal work flow for each system component. At the same time, a performance measurement model is learned to characterize the normal execution performance based on the log mes-sages’ timing information. With these learned models, we can automatically detect anomalies in newly input log files. Experiments on Hadoop and SILK show that the technique can effectively detect running anomalies.",https://ieeexplore.ieee.org/document/5360240,True,,322.0,['automaton'],['logs'],['novel-use'],['failure-detection'],"['hadoop', 'silk']",ICDM '09: Proceedings of the 2009 Ninth IEEE International Conference on Data Mining,True,['failure-management'],True,,"['anomaly-detection', 'log-enhancement']",True,52.0,10.1109/ICDM.2009.60,,,
612,Algorithms for anomaly detection of traces in logs of process aware information systems,"['FáBio Bezerra', 'Jacques Wainer']",2013,"This paper discusses four algorithms for detecting anomalies in logs of process aware systems. One of the algorithms only marks as potential anomalies traces that are infrequent in the log. The other three algorithms: threshold, iterative and sampling are based on mining a process model from the log, or a subset of it. The algorithms were evaluated on a set of 1500 artificial logs, with different profiles on the number of anomalous traces and the number of times each anomalous traces was present in the log. The sampling algorithm proved to be the most effective solution. We also applied the algorithm to a real log, and compared the resulting detected anomalous traces with the ones detected by a different procedure that relies on manual choices.",https://doi.org/10.1016/j.is.2012.04.004,True,,2.0,,,,['failure-detection'],,Information Systems,True,['failure-management'],,True,,,,10.1016/j.is.2012.04.004,,,
613,A novel technique for long-term anomaly detection in the cloud,"['Owen Vallis', 'Jordan Hochenbaum', 'Arun Kejariwal']",2014,"High availability and performance of a web service is key, amongst other factors, to the overall user experience (which in turn directly impacts the bottom-line). Exogenic and/or endogenic factors often give rise to anomalies that make maintaining high availability and delivering high performance very challenging. Although there exists a large body of prior research in anomaly detection, existing techniques are not suitable for detecting long-term anomalies owing to a predominant underlying trend component in the time series data. To this end, we developed a novel statistical technique to automatically detect long-term anomalies in cloud data. Specifically, the technique employs statistical learning to detect anomalies in both application as well as system metrics. Further, the technique uses robust statistical metrics, viz., median, and median absolute deviation (MAD), and piecewise approximation of the underlying long-term trend to accurately detect anomalies even in the presence of intra-day and/or weekly seasonality. We demonstrate the efficacy of the proposed technique using production data and report Precision, Recall, and F-measure measure. Multiple teams at Twitter are currently using the proposed technique on a daily basis.",https://dl.acm.org/doi/10.5555/2696535.2696550,True,,140.0,,,,['failure-detection'],,HotCloud'14: Proceedings of the 6th USENIX conference on Hot Topics in Cloud Computing,True,['failure-management'],,True,['anomaly-detection'],,,10.5555/2696535.2696550,,,
614,A survey of intelligent network fault diagnosis technology,"['Feng, Lv', 'Xiang, Li', 'Xiu-qing, Wang']",2013,"This paper firstly discusses the common fault types of networks, and the necessity and importance of intelligent fault diagnosis for networks. Secondly, the basic process for network diagnosis systems is put forward. Thirdly, the basic idea and various popular methods of network fault diagnosis based on the intelligent technology, such as: Expert System, Bayesian Network, Rough Set and Neural Networks, are reviewed. Finally, the existing problems and the future research direction for network fault diagnosis are …",https://doi.org/10.1109/CCDC.2013.6561817,True,,4.0,,,,['root-cause-analysis'],,2013 25th Chinese Control and Decision Conference (CCDC),True,['failure-management'],,True,,,,10.1109/CCDC.2013.6561817,,,
615,A Survey of Fault Diagnosis and Fault-Tolerant Techniques—Part I: Fault Diagnosis With Model-Based and Signal-Based Approaches,"['Gao, Zhiwei', 'Cecati, Carlo', 'Ding, Steven X']",2015,"With the continuous increase in complexity and expense of industrial systems, there is less tolerance for performance degradation, productivity decrease, and safety hazards, which greatly necessitates to detect and identify any kinds of potential abnormalities and faults as early as possible and implement real-time fault-tolerant operation for minimizing performance degradation and avoiding dangerous situations. During the last four decades, fruitful results have been reported about fault diagnosis and fault-tolerant control methods …",https://doi.org/10.1109/TIE.2015.2417501,True,,1059.0,[],[],['survey'],['root-cause-analysis'],[],,True,['failure-management'],True,True,"['root-cause-diagnosis', 'fault-localization']",,,10.1109/TIE.2015.2417501,,,
616,Requirements-Driven Root Cause Analysis Using Markov Logic Networks,"['Hamzeh Zawawy', 'Kostas Kontogiannis', 'John Mylopoulos', 'Serge Mankovskii']",2012,"Root cause analysis for software systems is a challenging diagnostic task, due to the complexity emanating from the interactions between system components and the sheer size of logged data. This diagnostic task is usually assisted by human experts who create mental models of the system-at-hand, in order to generate hypotheses and conduct the analysis. In this paper, we propose a root cause analysis framework based on requirement goal models. We consequently use these models to generate a Markov Logic Network that serves as a diagnostic knowledge repository. The network can be trained and used to provide inferences as to why and how a particular failure observation may be explained by collected logged data. The proposed framework improves over existing approaches by handling uncertainty in observations, using natively generated log data, and by providing ranked diagnoses. The framework is illustrated using a test environment based on commercial off-the-shelf software components.",https://doi.org/10.1007/978-3-642-31095-9_23,True,,7.0,['markov-model'],,,['root-cause-analysis'],,CAiSE'12: Proceedings of the 24th international conference on Advanced Information Systems Engineering,True,['failure-management'],,True,,,,10.1007/978-3-642-31095-9_23,,,
617,A Novel Framework for Real-Time Fault Diagnosis Based on Dynamic Fault Tree Analysis,"['Duan, R', 'Ou, X']",2012,,https://doi.org/10.19026/rjaset.5.4744,True,,2.0,"['bayesian-network', 'fault-tree']",,['novel-use'],['root-cause-analysis'],,,True,['failure-management'],,,['root-cause-diagnosis'],,,10.19026/rjaset.5.4744,,,
618,Fault Detection and Diagnosis in Distributed Systems : An Approach by Partially Stochastic Petri Nets,"['Armen Aghasaryan', 'Eric Fabre', 'Albert Benveniste', 'Renée Boubour', 'Claude Jard']",1998,"We address the problem of alarm correlation in large distributed
systems. The key idea is to make use of the concurrence of events in order
to separate and simplify the state estimation in a faulty system. Petri nets
and their causality semantics are used to model concurrency. Special
partially stochastic Petri nets are developed, that establish some kind of
equivalence between concurrence and independence. The diagnosis problem is
defined as the computation of the most likely history of the net given a
sequence of observed alarms. Solutions are provided in four contexts, with a
gradual complexity on the structure of observations.",https://doi.org/10.1023/A:1008241818642,True,,96.0,,,,['failure-detection'],,Discrete Event Dynamic Systems,True,['failure-management'],,True,,,,10.1023/A:1008241818642,,,
619,AIOps: Predictive Analytics & Machine Learning in Operations,"['Masood, Adnan', 'Hashmi, Adnan']",2019,"The operations landscape today is more complex than ever. IT Ops teams have to fight an uphill battle managing the massive amounts of data that is being generated by modern IT systems. They are expected to handle more incidents than ever before with shorter service-level agreements (SLAs), respond to these incidents more quickly, and improve on key metrics, such as mean time to detect (MTTD), mean time to failure (MTTF), mean time between failures (MTBF), and mean time to repair (MTTR). This is not because of lack of tools. Digital enterprise journal research suggests that 41 percent of enterprises use ten or more tools for IT performance monitoring, and downtime can get expensive when companies lose a whopping $5.6 million per outage and MTTR averages 4.2 hours and wastes precious resources. With a hybrid multi-cloud, multi-tenant environment, organizations need even more tools to manage the multiple facets of capacity planning, resource utilization, storage management, anomaly detection, and threat detection and analysis, to name a few.",https://doi.org/10.1007/978-1-4842-4106-6_7,True,,1.0,,,['discussion'],,,Cognitive Computing Recipes,True,['aiops-general'],,True,,,,10.1007/978-1-4842-4106-6_7,,,
620,Guaranteeing High Availability Goals for Virtual Machine Placement,"['Eyal Bin', 'Ofer Biran', 'Odellia Boni', 'Erez Hadad', 'Eliot K. Kolodner', 'Yosef Moatti', 'Dean H. Lorenz']",2011,"The placement of virtual machines (VMs) on a cluster of hosts under multiple constraints, including administrative (security, regulations) resource-oriented (capacity, energy), and QoS-oriented (performance) is a highly complex task. We define a new high-availability property for a VM, when a VM is marked as k-resilient, as long as there are up to k host failures, it should be guaranteed that it can be relocated to a non-failed host without relocating other VMs. Together with Hardware Predictive Failure Analysis and live migration, which enable VMs to be evacuated from a host before it fails, this property allows the continuous running of VMs on the cluster despite host failures. The complexity of the constraints associated with k-resiliency, which are naturally expressed by Second Order logic statements, prevented their integration into the placement computation until now. We present a novel algorithm which enables this integration by transforming the k-resiliency constraints to rules consumable by a generic Constraint Programming engine, prove that it guarantees the required resiliency and describe the implementation. We provide some preliminary results and compare our high availability support with naive solutions.",https://doi.org/10.1109/ICDCS.2011.72,True,,19.0,,,,,,ICDCS '11: Proceedings of the 2011 31st International Conference on Distributed Computing Systems,True,['resource-provisioning'],,True,,,,10.1109/ICDCS.2011.72,,,
621,A Review of Machine Learning and Meta-heuristic Methods for Scheduling Parallel Computing Systems,"['Suejb Memeti', 'Sabri Pllana', 'Alécio Binotto', 'Joanna Kołodziej', 'Ivona Brandic']",2018,"Optimized software execution on parallel computing systems demands consideration of many parameters at run-time. Determining the optimal set of parameters in a given execution context is a complex task, and therefore to address this issue researchers have proposed different approaches that use heuristic search or machine learning. In this paper, we undertake a systematic literature review to aggregate, analyze and classify the existing software optimization methods for parallel computing systems. We review approaches that use machine learning or meta-heuristics for scheduling parallel computing systems. Additionally, we discuss challenges and future research directions. The results of this study may help to better understand the state-of-the-art techniques that use machine learning and meta-heuristics to deal with the complexity of scheduling parallel computing systems. Furthermore, it may aid in understanding the limitations of existing approaches and identification of areas for improvement.",https://doi.org/10.1145/3230905.3230906,True,,3.0,,,['survey'],['scheduling'],,LOPAL '18: Proceedings of the International Conference on Learning and Optimization Algorithms: Theory and Applications,True,['resource-provisioning'],,True,,,,10.1145/3230905.3230906,,,
622,iDice: problem identification for emerging issues,"['Qingwei Lin', 'Jian-Guang Lou', 'Hongyu Zhang', 'Dongmei Zhang']",2016,"One challenge for maintaining a large-scale software system, especially an online service system, is to quickly respond to customer issues. The issue reports typically have many categorical attributes that reflect the characteristics of the issues. For a commercial system, most of the time the volume of reported issues is relatively constant. Sometimes, there are emerging issues that lead to significant volume increase. It is important for support engineers to efficiently and effectively identify and resolve such emerging issues, since they have impacted a large number of customers. Currently, problem identification for an emerging issue is a tedious and error-prone process, because it requires support engineers to manually identify a particular attribute combination that characterizes the emerging issue among a large number of attribute combinations. We call such an attribute combination effective combination, which is important for issue isolation and diagnosis. In this paper, we propose iDice, an approach that can identify the effective combination for an emerging issue with high quality and performance. We evaluate the effectiveness and efficiency of iDice through experiments. We have also successfully applied iDice to several Microsoft online service systems in production. The results confirm that iDice can help identify emerging issues and reduce maintenance effort.",https://doi.org/10.1145/2884781.2884795,True,,6.0,,,,['root-cause-analysis'],,ICSE '16: Proceedings of the 38th International Conference on Software Engineering,True,['failure-management'],,True,,,,10.1145/2884781.2884795,,,
623,An approach for QoS-aware service composition based on genetic algorithms,"['Gerardo Canfora', 'Massimiliano Di Penta', 'Raffaele Esposito', 'Maria Luisa Villani']",2005,"Web services are rapidly changing the landscape of software engineering. One of the most interesting challenges introduced by web services is represented by Quality Of Service (QoS)--aware composition and late--binding. This allows to bind, at run--time, a service--oriented system with a set of services that, among those providing the required features, meet some non--functional constraints, and optimize criteria such as the overall cost or response time. In other words, QoS--aware composition can be modeled as an optimization problem.We propose to adopt Genetic Algorithms to this aim. Genetic Algorithms, while being slower than integer programming, represent a more scalable choice, and are more suitable to handle generic QoS attributes. The paper describes our approach and its applicability, advantages and weaknesses, discussing results of some numerical simulations.",https://doi.org/10.1145/1068009.1068189,True,,576.0,['genetic-programming'],['sla'],['new-method'],['service-composition'],,GECCO '05: Proceedings of the 7th annual conference on Genetic and evolutionary computation,True,['resource-provisioning'],,True,,,,10.1145/1068009.1068189,,,
624,Toward autonomic web services trust and selection,"['E. Michael Maximilien', 'Munindar P. Singh']",2004,"Emerging Web services standards enable the development of large-scale applications in open environments. In particular, they enable services to be dynamically bound. However, current techniques fail to address the critical problem of selecting the right service instances. Service selection should be determined based on user preferences and business policies, and consider the trustworthiness of service instances. We propose a multiagent approach that naturally provides a solution to the selection problem. This approach is based on an architecture and programming model in which agents represent applications and services. The agents support considerations of semantics and quality of service (QoS). They interact and share information, in essence creating an ecosystem of collaborative service providers and consumers. Consequently, our approach enables applications to be dynamically configured at runtime in a manner that continually adapts to the preferences of the participants. Our agents are designed using decision theory and use ontologies. We evaluate our approach through simulation experiments.",https://doi.org/10.1145/1035167.1035198,True,,199.0,"['decision-theory', 'onthology']","['sla', 'structure']",['new-method'],['service-composition'],,ICSOC '04: Proceedings of the 2nd international conference on Service oriented computing,True,['resource-provisioning'],,True,,,,10.1145/1035167.1035198,,,
625,Energy Efficient Resource Management in Virtualized Cloud Data Centers,"['Anton Beloglazov', 'Rajkumar Buyya']",2010,"Rapid growth of the demand for computational power by scientific, business and web-applications has led to the creation of large-scale data centers consuming enormous amounts of electrical power. We propose an energy efficient resource management system for virtualized Cloud data centers that reduces operational costs and provides required Quality of Service (QoS). Energy savings are achieved by continuous consolidation of VMs according to current utilization of resources, virtual network topologies established between VMs and thermal state of computing nodes. We present first results of simulation-driven evaluation of heuristics for dynamic reallocation of VMs using live migration according to current requirements for CPU performance. The results show that the proposed technique brings substantial energy savings, while ensuring reliable QoS. This justifies further investigation and development of the proposed resource management system.",https://doi.org/10.1109/CCGRID.2010.46,True,,95.0,,,['comparison'],"['resource-consolidation', 'power-management']",,"CCGRID '10: Proceedings of the 2010 10th IEEE/ACM International Conference on Cluster, Cloud and Grid Computing",True,['resource-provisioning'],,True,,,,10.1109/CCGRID.2010.46,,,
626,Analytic modeling of multitier Internet applications,"['Bhuvan Urgaonkar', 'Giovanni Pacifici', 'Prashant Shenoy', 'Mike Spreitzer', 'Asser Tantawi']",2007,"Since many Internet applications employ a multitier architecture, in this article, we focus on the problem of analytically modeling the behavior of such applications. We present a model based on a network of queues where the queues represent different tiers of the application. Our model is sufficiently general to capture (i) the behavior of tiers with significantly different performance characteristics and (ii) application idiosyncrasies such as session-based workloads, tier replication, load imbalances across replicas, and caching at intermediate tiers. We validate our model using real multitier applications running on a Linux server cluster. Our experiments indicate that our model faithfully captures the performance of these applications for a number of workloads and configurations. Furthermore, our model successfully handles a comprehensive range of resource utilization---from 0 to near saturation for the CPU---for two separate tiers. For a variety of scenarios, including those with caching at one of the application tiers, the average response times predicted by our model were within the 95% confidence intervals of the observed average response times. Our experiments also demonstrate the utility of the model for dynamic capacity provisioning, performance prediction, bottleneck identification, and session policing. In one scenario, where the request arrival rate increased from less than 1500 to nearly 4200 requests/minute, a dynamic provisioning technique employing our model was able to maintain response time targets by increasing the capacity of two of the tiers by factors of 2 and 3.5, respectively.",https://doi.org/10.1145/1232722.1232724,True,,85.0,,,,['workload-prediction'],,ACM Transactions on the Web (TWEB),True,['resource-provisioning'],,True,,,,10.1145/1232722.1232724,,,
627,An empirical comparison of methods to support QoS-aware service selection,"['Bice Cavallo', 'Massimiliano Di Penta', 'Gerardo Canfora']",2010,"Run-time binding is an important and useful feature of Service Oriented Architectures (SOA), which aims at selecting, among functionally equivalent services, the ones that optimize some QoS objective of the overall application. To this aim, it is particularly relevant to forecast the QoS a service will likely exhibit in future invocations. This paper presents an empirical study aimed at comparing different approaches for QoS forecasting, namely the use of average and current values, linear models, and models based on time series. The study is performed on QoS data obtained by monitoring the execution of 10 real services for 4 months. Results show that, overall, the use of time series forecasting has the best compromise in ensuring a good prediction error, being sensible to outliers, and being able to predict likely violations of QoS constraints.",https://doi.org/10.1145/1808885.1808899,True,,41.0,,,,['service-composition'],,PESOS '10: Proceedings of the 2nd International Workshop on Principles of Engineering Service-Oriented Systems,True,['resource-provisioning'],,,,,,10.1145/1808885.1808899,,,
628,Self-adaptive workload classification and forecasting for proactive resource provisioning,"['Nikolas Roman Herbst', 'Nikolaus Huber', 'Samuel Kounev', 'Erich Amrehn']",2013,"As modern enterprise software systems become increasingly dynamic, workload forecasting techniques are gaining in importance as a foundation for online capacity planning and resource management. Time series analysis offers a broad spectrum of methods to calculate workload forecasts based on history monitoring data. Related work in the field of workload forecasting mostly concentrates on evaluating specific methods and their individual optimisation potential or on predicting Quality-of-Service (QoS) metrics directly. As a basis, we present a survey on established forecasting methods of the time series analysis concerning their benefits and drawbacks and group them according to their computational overheads. In this paper, we propose a novel self-adaptive approach that selects suitable forecasting methods for a given context based on a decision tree and direct feedback cycles together with a corresponding implementation. The user needs to provide only his general forecasting objectives. In several experiments and case studies based on real-world workload traces, we show that our implementation of the approach provides continuous and reliable forecast results at run-time. The results of this extensive evaluation show that the relative error of the individual forecast points is significantly reduced compared to statically applied forecasting methods, e.g. in an exemplary scenario on average by 37%. In a case study, between 55% and 75% of the violations of a given service level agreement can be prevented by applying proactive resource provisioning based on the forecast results of our implementation.",https://doi.org/10.1145/2479871.2479899,True,,37.0,,,,['workload-prediction'],,ICPE '13: Proceedings of the 4th ACM/SPEC International Conference on Performance Engineering,True,['resource-provisioning'],,True,,,,10.1145/2479871.2479899,,,
629,Short term performance forecasting in enterprise systems,"['Rob Powers', 'Moises Goldszmidt', 'Ira Cohen']",2005,"We use data mining and machine learning techniques to predict upcoming periods of high utilization or poor performance in enterprise systems. The abundant data available and complexity of these systems defies human characterization or static models and makes the task suitable for data mining techniques. We formulate the problem as one of classification: given current and past information about the system's behavior, can we forecast whether the system will meet its performance targets over the next hour? Using real data gathered from several enterprise systems in Hewlett-Packard, we compare several approaches ranging from time series to Bayesian networks. Besides establishing the predictive power of these approaches our study analyzes three dimensions that are important for their application as a stand alone tool. First, it quantifies the gain in accuracy of multivariate prediction methods over simple statistical univariate methods. Second, it quantifies the variations in accuracy when using different classes of system and workload features. Third, it establishes that models induced using combined data from various systems generalize well and are applicable to new systems, enabling accurate predictions on systems with insufficient historical data. Together this analysis offers a promising outlook on the development of tools to automate assignment of resources to stabilize performance, (e.g., adding servers to a cluster) and allow opportunistic job scheduling (e.g., backups or virus scans).",https://doi.org/10.1145/1081870.1081976,True,,32.0,,,,['failure-prediction'],,KDD '05: Proceedings of the eleventh ACM SIGKDD international conference on Knowledge discovery in data mining,True,['failure-management'],,,,,,10.1145/1081870.1081976,,,
630,SmartDispatch: enabling efficient ticket dispatch in an IT service environment,"['Shivali Agarwal', 'Renuka Sindhgatta', 'Bikram Sengupta']",2012,"In an IT service delivery environment, the speedy dispatch of a ticket to the correct resolution group is the crucial first step in the problem resolution process. The size and complexity of such environments make the dispatch decision challenging, and incorrect routing by a human dispatcher can lead to significant delays that degrade customer satisfaction, and also have adverse financial implications for both the customer and the IT vendor. In this paper, we present SmartDispatch, a learning-based tool that seeks to automate the process of ticket dispatch while maintaining high accuracy levels. SmartDispatch comes with two classification approaches - the well-known SVM method, and a discriminative term-based approach that we designed to address some of the issues in SVM classification that were empirically observed. Using a combination of these approaches, SmartDispatch is able to automate the dispatch of a ticket to the correct resolution group for a large share of the tickets, while for the rest, it is able to suggest a short list of 3-5 groups that contain the correct resolution group with a high probability. Empirical evaluation of SmartDispatch on data from 3 large service engagement projects in IBM demonstrate the efficacy and practical utility of the approach.",https://doi.org/10.1145/2339530.2339744,True,,25.0,['support-vector-machine'],['tickets'],,['remediation'],[],KDD '12: Proceedings of the 18th ACM SIGKDD international conference on Knowledge discovery and data mining,True,['failure-management'],,True,,,,10.1145/2339530.2339744,,,
631,A Systematic Review of Service Level Management in the Cloud,"['Funmilade Faniyi', 'Rami Bahsoon']",2015,"Cloud computing make it possible to flexibly procure, scale, and release computational resources on demand in response to workload changes. Stakeholders in business and academia are increasingly exploring cloud deployment options for their critical applications. One open problem is that service level agreements (SLAs) in the cloud ecosystem are yet to mature to a state where critical applications can be reliably deployed in clouds. This article systematically surveys the landscape of SLA-based cloud research to understand the state of the art and identify open problems. The survey is particularly aimed at the resource allocation phase of the SLA life cycle while highlighting implications on other phases. Results indicate that (i) minimal number of SLA parameters are accounted for in most studies; (ii) heuristics, policies, and optimisation are the most commonly used techniques for resource allocation; and (iii) the monitor-analysis-plan-execute (MAPE) architecture style is predominant in autonomic cloud systems. The results contribute to the fundamentals of engineering cloud SLA and their autonomic management, motivating further research and industrial-oriented solutions.",https://doi.org/10.1145/2843890,True,,15.0,,,,,,ACM Computing Surveys,True,['resource-provisioning'],,True,,,,10.1145/2843890,,,
632,Automatic request categorization in internet services,"['Abhishek B. Sharma', 'Ranjita Bhagwan', 'Monojit Choudhury', 'Leana Golubchik', 'Ramesh Govindan', 'Geoffrey M. Voelker']",2008,"Modeling system performance and workload characteristics has become essential for efficiently provisioning Internet services and for accurately predicting future resource requirements on anticipated workloads. The accuracy of these models benefits substantially by differentiating among categories of requests based on their resource usage characteristics. However, categorizing requests and their resource demands often requires significantly more monitoring infrastructure. In this paper, we describe a method to automatically differentiate and categorize requests without requiring sophisticated monitoring techniques. Using machine learning, our method requires only aggregate measures such as total number of requests and the total CPU and network demands, and does not assume prior knowledge of request categories or their individual resource demands. We explore the feasibility of our method on the .Net PetShop 4.0 benchmark application, and show that it works well while being lightweight, generic, and easily deployable.",https://doi.org/10.1145/1453175.1453179,True,,13.0,,,,,,ACM SIGMETRICS Performance Evaluation Review,True,['resource-provisioning'],,,,,,10.1145/1453175.1453179,,,
633,SLA-Driven Dynamic Resource Management for Multi-tier Web Applications in a Cloud,"['Waheed Iqbal', 'Matthew N. Dailey', 'David Carrera']",2010,"Current service-level agreements (SLAs) offered by cloud providers do not make guarantees about response time of Web applications hosted on the cloud. Satisfying a maximum average response time guarantee for Web applications is difficult due to unpredictable traffic patterns. The complex nature of multi-tier Web applications increases the difficulty of identifying bottlenecks and resolving them automatically. It may be possible to minimize the probability that tiers (hosted on virtual machines) become bottlenecks by optimizing the placement of the virtual machines in a cloud. This research focuses on enabling clouds to offer multi-tier Web application owners maximum response time guarantees while minimizing resource utilization. We present our basic approach, preliminary experiments, and results on a EUCALYPTUS-based testbed cloud. Our preliminary results shows that dynamic bottleneck detection and resolution for multi-tier Web application hosted on the cloud will help to offer SLAs that can offer response time guarantees.",https://doi.org/10.1109/CCGRID.2010.59,True,,9.0,,,,,,"CCGRID '10: Proceedings of the 2010 10th IEEE/ACM International Conference on Cluster, Cloud and Grid Computing",True,['resource-provisioning'],,True,,,,10.1109/CCGRID.2010.59,,,
634,Reinforcement learning for autonomic network repair,"['M.L. Littman', 'N. Ravi', 'E. Fenson', 'R. Howard']",2004,We report on our efforts to formulate autonomic network repair as a reinforcement-learning problem. Our implemented system is able to learn to efficiently restore network connectivity after a failure.,https://ieeexplore.ieee.org/document/1301380,True,,11.0,,,,['remediation'],,"International Conference on Autonomic Computing, 2004. Proceedings.",True,['failure-management'],,,,,,10.1109/ICAC.2004.1301380,,,
635,Building autonomic systems using collaborative reinforcement learning,"['Jim Dowling', 'Raymond Cunningham', 'Eoin Curran', 'Vinny Cahill']",2006,"This paper presents Collaborative Reinforcement Learning (CRL), a coordination model for online system optimization in decentralized multi-agent systems. In CRL system optimization problems are represented as a set of discrete optimization problems, each of whose solution cost is minimized by model-based reinforcement learning agents collaborating on their solution. CRL systems can be built to provide autonomic behaviours such as optimizing system performance in an unpredictable environment and adaptation to partial failures. We evaluate CRL using an ad hoc routing protocol that optimizes system routing performance in an unpredictable network environment.",https://dl.acm.org/doi/10.1017/S0269888906000956,True,,33.0,,,,['remediation'],,The Knowledge Engineering Review,True,['failure-management'],,,,,,10.1017/S0269888906000956,,,
636,Predictive models for proactive network management: application to a production Web server,"['D. Shen', 'J.L. Hellerstein']",2000,"Proactive management holds the promise of taking corrective actions in advance of service disruptions. Achieving this goal requires predictive models so that potential problems can be anticipated. Our approach builds on previous research in which HTTP operations per second are studied in a Web server. As in this prior work, we model HTTP operations as two subprocesses, a (deterministic) trend subprocess and a (random but stationary) residual subprocess. Herein, the trend model is enhanced by using a low-pass filter. Further, we employ techniques that reduce the required data history, thereby reducing the impact of changes in the trend process. As in the prior work, an autoregressive model is used for the residual process. We study the limits of the autoregressive model in the prediction of network traffic. Then we demonstrate that long-range dependencies remain in the residual process even after autoregressive components are removed, which impacts our ability to predict future observations. Last, we analyze the validity of assumptions employed, especially the normality assumption.",https://ieeexplore.ieee.org/document/830432,True,,21.0,,,,,,NOMS 2000. 2000 IEEE/IFIP Network Operations and Management Symposium'The Networked Planet: Management Beyond 2000'(Cat. No. 00CB37074),True,['resource-provisioning'],,,,,,,,,
637,An approach to on-line predictive detection,"['Fan Zhang', 'J.L. Hellerstein']",2000,"Predicting network performance problems enables network operators to take corrective actions in advance of service disruptions. Typically, service problems are detected by tests that compare a metric (e.g., response time) to a threshold. The authors present an online algorithm for predicting the probability of threshold violations over a time horizon. The algorithm uses two cascaded submodels. The first removes non-stationarities by employing a discrete time Kalman filter in combination with analysis of variance. We derive parameters of the Kalman filter from differential equations that describe characteristics of the data. The second submodel estimates the probability of threshold violations by using a second order autoregressive model in combination with change-point detection. Using data from a production Web server, we evaluate our approach and show that it produces average accuracies that are comparable to those of an offline algorithm. However, our online algorithm produces predictions with considerably smaller variances. Further advantages of our approach are: (a) requiring much less data than the offline technique, one day versus multiple months; and (b) adapting to changes in the system and workloads since parameters are estimated online.",https://ieeexplore.ieee.org/document/876583,True,,8.0,,,,['failure-prediction'],,,True,['failure-management'],,,,,,10.1109/MASCOT.2000.876583,,,
638,Log File Categorization and Anomaly Analysis Using Grammar Inference,"['Memon, Ahmed Umar', 'Cordy, James R', 'Dean, Thomas']",2008,"In the information age of today, vast amounts of sensitive and confidential data is exchanged over an array of different mediums. Accompanied with this phenomenon is a comparable increase in the number and types of attacks to acquire this information. Information security and data consistency have hence, become quintessentially important. Log file analysis has proven to be a good defense mechanism as logs provide an accessible record of network activities in the form of server generated messages. However, manual analysis is tedious …",http://hdl.handle.net/1974/1217,True,,10.0,,,,['failure-detection'],,,True,['failure-management'],,,,,,,,,
639,Unsupervised Anomaly Detection in Noisy Business Process Event Logs Using Denoising Autoencoders,"['Nolle, Timo', 'Seeliger, Alexander', 'M{\\""u}hlh{\\""a}user, Max']",2016,"Business processes are prone to subtle changes over time, as unwanted behavior manifests in the execution over time. This problem is related to anomaly detection, as these subtle changes start of as anomalies at first, and thus it is important to detect them early. However, the necessary process documentation is often outdated, and thus not usable. Moreover, the only way of analyzing a process in execution is the use of event logs coming from process-aware information systems, but these event logs already contain anomalous behavior and …",https://link.springer.com/chapter/10.1007/978-3-319-46307-0_28,True,,20.0,"['multilayer-perceptron', 'autoencoder']",['logs'],,['failure-detection'],,International conference on discovery science,True,['failure-management'],,,['anomaly-detection'],,,,,,
640,Anomaly Detection for Application Log Data,"['Grover, Aarish']",2018,"In software development, there is an absolute requirement to ensure that a system once developed, functions at its best throughout its lifetime. Application log data is critical to maintaining application performance and thus techniques to parse, understand and detect anomalies in application log data are critical to ensuring efficiency in software development. While initially hampered by limited hardware and lack of quality datasets, anomaly detection techniques have recently received a surge of interest with advancements in machine learning technology and especially neural networks. In this paper, we explore anomaly detection, historical techniques to detect anomalies and recent advancements in neural networks, which promise to revolutionize anomaly detection in application log data. Further, we analyze the most promising anomaly detection techniques and propose a hybrid model combining LSTM Neural Network and Auto Encoder which improves upon existing techniques.",https://scholarworks.sjsu.edu/etd_projects/635/,True,,2.0,"['autoencoder', 'rnn']",['logs'],,['failure-detection'],,,True,['failure-management'],,,['anomaly-detection'],,,,,,
641,Anomaly Detection in Unstructured Time Series Data using an LSTM Autoencoder,"['Wolpher, Maxim']",2018,"An exploration of anomaly detection. Much work has been done on the topic of anomaly detection, but what seems to be lacking is a dive into anomaly detection of unstructured and unlabeled data. This thesis aims to determine the effectiveness of combining recurrent neural networks with autoencoder structures for sequential anomaly detection. The use of an LSTM autoencoder will be detailed, but along the way there will also be background on time-independent anomaly detection using Isolation Forests and Replicator Neural Networks on …",http://www.diva-portal.org/smash/get/diva2:1225367/FULLTEXT01.pdf,True,,1.0,,,,['failure-detection'],,,True,['failure-management'],,,,,,,,,
642,adaptive event prediction strategy with dynamic time window for large-scale HPC systems,"['Ana Gainaru', 'Franck Cappello', 'Joshi Fullop', 'Stefan Trausan-Matu', 'William Kramer']",2011,"In this paper, we analyse messages generated by different HPC large-scale systems in order to extract sequences of correlated events which we lately use to predict the normal and faulty behaviour of the system. Our method uses a dynamic window strategy that is able to find frequent sequences of events regardless on the time delay between them. Most of the current related research narrows the correlation extraction to fixed and relatively small time windows that do not reflect the whole behaviour of the system. The generated events are in constant change during the lifetime of the machine. We consider that it is important to update the sequences at runtime by applying modifications after each prediction phase according to the forecast's accuracy and the difference between what was expected and what really happened. Our experiments show that our analysing system is able to predict around 60% of events with a precision of around 85% at a lower event granularity than before.",https://dl.acm.org/doi/10.1145/2038633.2038637,True,,35.0,,,,['failure-detection'],,SLAML '11: Managing Large-scale Systems via the Analysis of System Logs and the Application of Machine Learning Techniques,True,['failure-management'],,True,['anomaly-detection'],,,10.1145/2038633.2038637,,,
643,Managing power consumption and performance of computing systems using reinforcement learning,"['Gerald Tesauro', 'Rajarshi Das', 'Hoi Chan', 'Jeffrey O. Kephart', 'Charles Lefurgy', 'David W. Levine', 'Freeman Rawson']",2007,"Electrical power management in large-scale IT systems such as commercial data-centers is an application area of rapidly growing interest from both an economic and ecological perspective, with billions of dollars and millions of metric tons of CO2 emissions at stake annually. Businesses want to save power without sacrificing performance. This paper presents a reinforcement learning approach to simultaneous online management of both performance and power consumption. We apply RL in a realistic laboratory testbed using a Blade cluster and dynamically varying HTTP workload running on a commercial web applications middleware platform. We embed a CPU frequency controller in the Blade servers' firmware, and we train policies for this controller using a multi-criteria reward signal depending on both application performance and CPU power consumption. Our testbed scenario posed a number of challenges to successful use of RL, including multiple disparate reward functions, limited decision sampling rates, and pathologies arising when using multiple sensor readings as state variables. We describe innovative practical solutions to these challenges, and demonstrate clear performance improvements over both hand-designed policies as well as obvious ""cookbook"" RL implementations.",https://dl.acm.org/doi/10.5555/2981562.2981750,True,,109.0,,,,"['power-management', 'resource-consolidation']",,NIPS'07: Proceedings of the 20th International Conference on Neural Information Processing Systems,True,['resource-provisioning'],,False,,,,10.5555/2981562.2981750,,,
644,Reinforcement Learning in Autonomic Computing: A Manifesto and Case Studies,['Gerald Tesauro'],2007,"Reinforcement learning is a promising new approach for automatically developing effective policies for real-time self-management. RL can achieve superior performance to traditional methods, while requiring less built-in domain knowledge. Several case studies from real and simulated systems management applications demonstrate RL's promises and challenges. These studies show that standard online RL can learn effective policies in feasible training times. Moreover, a Hybrid RL approach can profit from any knowledge contained in an existing policy by training on the policy's observable behavior, without needing to interface directly to such knowledge",https://ieeexplore.ieee.org/document/4061117,True,,58.0,,,,['resource-consolidation'],,,True,['resource-provisioning'],,False,,,,,,,
645,A reinforcement learning approach to dynamic resource allocation,"['Vengerov, David']",2007,"This paper presents a general framework for performing adaptive reconfiguration of a distributed system based on maximizing the long-term business value, defined as the discounted sum of all future rewards and penalties. The problem of dynamic resource allocation among multiple entities sharing a common set of resources is used as an example. A specific architecture (DRA-FRL) is presented, which uses the emerging methodology of reinforcement learning in conjunction with fuzzy rulebases to achieve the …",https://dl.acm.org/doi/book/10.5555/1698183,True,,77.0,,,,['resource-consolidation'],,,True,['resource-provisioning'],,True,,,,10.5555/1698183,,,
646,Online resource allocation using decompositional reinforcement learning,['Gerald Tesauro'],2005,"This paper considers a novel application domain for reinforcement learning: that of ""autonomic computing,"" i.e. selfmanaging computing systems. RL is applied to an online resource allocation task in a distributed multi-application computing environment with independent time-varying load in each application. The task is to allocate servers in real time so as to maximize the sum of performance-based expected utility in each application. This task may be treated as a composite MDP, and to exploit the problem structure, a simple localized RL approach is proposed, with better scalability than previous approaches. The RL approach is tested in a realistic prototype data center comprising real servers, real HTTP requests, and realistic time-varying demand. This domain poses a number of major challenges associated with live training in a real system, including: the need for rapid training, exploration that avoids excessive penalties, and handling complex, potentially non-Markovian system effects. The early results are encouraging: in overnight training, RL performs as well as or slightly better than heavily researched model-based approaches derived from queuing theory.",https://dl.acm.org/doi/10.5555/1619410.1619475,True,,117.0,['reinforcement-learning'],,,['resource-consolidation'],,AAAI'05: Proceedings of the 20th national conference on Artificial intelligence - Volume 2,True,['resource-provisioning'],,False,,,,10.5555/1619410.1619475,,,
647,Self-Optimizing Memory Controllers: A Reinforcement Learning Approach,"['Engin Ipek', 'Onur Mutlu', 'José F. Martínez', 'Rich Caruana']",2008,"Efficiently utilizing off-chip DRAM bandwidth is a critical issuein designing cost-effective, high-performance chip multiprocessors(CMPs). Conventional memory controllers deliver relativelylow performance in part because they often employ fixed,rigid access scheduling policies designed for average-case applicationbehavior. As a result, they cannot learn and optimizethe long-term performance impact of their scheduling decisions,and cannot adapt their scheduling policies to dynamic workloadbehavior.We propose a new, self-optimizing memory controller designthat operates using the principles of reinforcement learning (RL)to overcome these limitations. Our RL-based memory controllerobserves the system state and estimates the long-term performanceimpact of each action it can take. In this way, the controllerlearns to optimize its scheduling policy on the fly to maximizelong-term performance. Our results show that an RL-basedmemory controller improves the performance of a set of parallelapplications run on a 4-core CMP by 19% on average (upto 33%), and it improves DRAM bandwidth utilization by 22%compared to a state-of-the-art controller.",https://dl.acm.org/doi/10.1145/1394608.1382172,True,,177.0,['reinforcement-learning'],,,['resource-consolidation'],['memory'],ACM SIGARCH Computer Architecture News,True,['resource-provisioning'],,True,,,,10.1145/1394608.1382172,,,
648,On the use of hybrid reinforcement learning for autonomic resource allocation,"['Tesauro, Gerald', 'Jong, Nicholas K', 'Das, Rajarshi', 'Bennani, Mohamed N']",2007,"Reinforcement Learning (RL) provides a promising new approach to systems performance management that differs radically from standard queuing-theoretic approaches making use of explicit system performance models. In principle, RL can automatically learn high-quality management policies without an explicit performance model or traffic model, and with little or no built-in system specific knowledge. In our original work (Das, R., Tesauro, G., Walsh, W.E.: IBM Research, Tech. Rep. RC23802 (2005), Tesauro, G.: In: Proc. of AAAI-05, pp. 886–891 (2005), Tesauro, G., Das, R., Walsh, W.E., Kephart, J.O.: In: Proc. of ICAC-05, pp. 342–343 (2005)) we showed the feasibility of using online RL to learn resource valuation estimates (in lookup table form) which can be used to make high-quality server allocation decisions in a multi-application prototype Data Center scenario. The present work shows how to combine the strengths of both RL and queuing models in a hybrid approach, in which RL trains offline on data collected while a queuing model policy controls the system. By training offline we avoid suffering potentially poor performance in live online training. We also now use RL to train nonlinear function approximators (e.g. multi-layer perceptrons) instead of lookup tables; this enables scaling to substantially larger state spaces. Our results now show that, in both open-loop and closed-loop traffic, hybrid RL training can achieve significant performance improvements over a variety of initial model-based policies. We also find that, as expected, RL can deal effectively with both transients and switching delays, which lie outside the scope of traditional steady-state queuing theory.",https://link.springer.com/article/10.1007%2Fs10586-007-0035-6,True,,120.0,"['multilayer-perceptron', 'reinforcement-learning']",,['novel-use'],['resource-consolidation'],,,True,['resource-provisioning'],,True,,,,,,,
649,Profiling and modeling resource usage of virtualized applications,"['Timothy Wood', 'Ludmila Cherkasova', 'Kivanc Ozonat', 'Prashant Shenoy']",2008,"Next Generation Data Centers are transforming labor-intensive, hard-coded systems into shared, virtualized, automated, and fully managed adaptive infrastructures. Virtualization technologies promise great opportunities for reducing energy and hardware costs through server consolidation. However, to safely transition an application running natively on real hardware to a virtualized environment, one needs to estimate the additional resource requirements incurred by virtualization overheads. In this work, we design a general approach for estimating the resource requirements of applications when they are transferred to a virtual environment. Our approach has two key components: a set of microbench-marks to profile the different types of virtualization overhead on a given platform, and a regression-based model that maps the native system usage profile into a virtualized one. This derived model can be used for estimating resource requirements of any application to be virtualized on a given platform. Our approach aims to eliminate error-prone manual processes and presents a fully automated solution. We illustrate the effectiveness of our methodology using Xen virtual machine monitor. Our evaluation shows that our automated model generation procedure effectively characterizes the different virtualization overheads of two diverse hardware platforms and that the models have median prediction error of less than 5% for both the RUBiS and TPC-W benchmarks.",https://dl.acm.org/doi/10.5555/1496950.1496973,True,,36.0,,,,,,Middleware '08: Proceedings of the 9th ACM/IFIP/USENIX International Conference on Middleware,True,['resource-provisioning'],,,,,,10.5555/1496950.1496973,,,
650,Application performance modeling in a virtualized environment,"['Sajib Kundu', 'Raju Rangaswami', 'Kaushik Dutta', 'Ming Zhao']",2010,"Performance models provide the ability to predict application performance for a given set of hardware resources and are used for capacity planning and resource management. Traditional performance models assume the availability of dedicated hardware for the application. With growing application deployment on virtualized hardware, hardware resources are increasingly shared across multiple virtual machines. In this paper, we build performance models for applications in virtualized environments. We identify a key set of virtualization architecture independent parameters that influence application performance for a diverse and representative set of applications. We explore several conventional modeling techniques and evaluate their effectiveness in modeling application performance in a virtualized environment. We propose an iterative model training technique based on artificial neural networks which is found to be accurate across a range of applications. The proposed approach is implemented as a prototype in Xen-based virtual machine environments and evaluated for accuracy, sensitivity to the training process, and overhead. Median modeling error in the range 1.16-6.65% across a diverse application set and low modeling overhead suggest the suitability of our approach in production virtualized environments.",https://ieeexplore.ieee.org/document/5463058,True,,68.0,,,,['workload-prediction'],,HPCA-16 2010 The Sixteenth International Symposium on High-Performance Computer Architecture,True,['resource-provisioning'],,True,,,,,,,
651,Correlating instrumentation data to system states: a building block for automated diagnosis and control,"['Ira Cohen', 'Moises Goldszmidt', 'Terence Kelly', 'Julie Symons', 'Jeffrey S. Chase']",2004,"This paper studies the use of statistical induction techniques as a basis for automated performance diagnosis and performance management. The goal of the work is to develop and evaluate tools for offline and online analysis of system metrics gathered from instrumentation in Internet server platforms. We use a promising class of probabilistic models (Tree-Augmented Bayesian Networks or TANs) to identify combinations of system-level metrics and threshold values that correlate with high-level performance states--compliance with Service Level Objectives (SLOs) for average-case response time--in a three-tier Web service under a variety of conditions. Experimental results from a testbed show that TAN models involving small subsets of metrics capture patterns of performance behavior in a way that is accurate and yields insights into the causes of observed performance effects. TANs are extremely efficient to represent and evaluate, and they have interpretability properties that make them excellent candidates for automated diagnosis and control. We explore the use of TAN models for offline forensic diagnosis, and in a limited online setting for performance forecasting with stable workloads.",https://dl.acm.org/doi/10.5555/1251254.1251270,True,,375.0,"['bayesian-network', 'decision-tree']","['kpis', 'host-metrics']",,"['failure-prediction', 'failure-detection']",['web-server'],OSDI'04: Proceedings of the 6th conference on Symposium on Operating Systems Design & Implementation - Volume 6,True,['failure-management'],True,,"['system-failure-prediction', 'anomaly-detection']",True,43.0,10.5555/1251254.1251270,,,
652,Autonomic resource management in virtualized data centers using fuzzy logic-based approaches,"['Xu, Jing', 'Zhao, Ming', ""Fortes, Jos{\\'e}"", 'Carpenter, Robert', 'Yousif, Mazin']",2008,"Data centers, as resource providers, are expected to deliver on performance guarantees while optimizing resource utilization to reduce cost. Virtualization techniques provide the opportunity of consolidating multiple separately managed containers of virtual resources on underutilized physical servers. A key challenge that comes with virtualization is the simultaneous on-demand provisioning of shared physical resources to virtual containers and the management of their capacities to meet service-quality targets at the least cost. This paper proposes a two-level resource management system to dynamically allocate resources to individual virtual containers. It uses local controllers at the virtual-container level and a global controller at the resource-pool level. An important advantage of this two-level control architecture is that it allows independent controller designs for separately optimizing the performance of applications and the use of resources. Autonomic resource allocation is realized through the interaction of the local and global controllers. A novelty of the local controller designs is their use of fuzzy logic-based approaches to efficiently and robustly deal with the complexity and uncertainties of dynamically changing workloads and resource usage. The global controller determines the resource allocation based on a proposed profit model, with the goal of maximizing the total profit of the data center. Experimental results obtained through a prototype implementation demonstrate that, for the scenarios under consideration, the proposed resource management system can significantly reduce resource consumption while still achieving application performance targets.",https://link.springer.com/article/10.1007/s10586-008-0060-0,True,,157.0,,,,['resource-consolidation'],,,True,['resource-provisioning'],,True,,,,,,,
653,"CHAMELEON: a self-evolving, fully-adaptive resource arbitrator for storage systems","['Sandeep Uttamchandani', 'Li Yin', 'Guillermo A. Alvarez', 'John Palmer', 'Gul Agha']",2005,"Enterprise applications typically depend on guaranteed performance from the storage subsystem, lest they fail. However, unregulated competition is unlikely to result in a fair, predictable apportioning of resources. Given that widespread access protocols and scheduling policies are largely best-effort, the problem of providing performance guarantees on a shared system is a very difficult one. Clients typically lack accurate information on the storage system's capabilities and on the access patterns of the workloads using it, thereby compounding the problem. CHAMELEON is an adaptive arbitrator for shared storage resources; it relies on a combination of self-refining models and constrained optimization to provide performance guarantees to clients. This process depends on minimal information from clients, and is fully adaptive; decisions are based on device and workload models automatically inferred, and continuously refined, at run-time. Corrective actions taken by CHAMELEON are only as radical as warranted by the current degree of knowledge about the system's behavior. In our experiments on a real storage system CHAMELEON identified, analyzed, and corrected performance violations in 3-14 minutes--which compares very favorably with the time a human administrator would have needed. Our learning-based paradigm is a most promising way of deploying large-scale storage systems that service variable workloads on an ever-changing mix of device types.",https://dl.acm.org/doi/10.5555/1247360.1247366,True,,65.0,,,,['resource-consolidation'],,ATEC '05: Proceedings of the annual conference on USENIX Annual Technical Conference,True,['resource-provisioning'],,,,,,10.5555/1247360.1247366,,,
654,Model-based and model-free approaches to autonomic resource allocation,"['Das, Rajarshi', 'Tesauro, Gerald', 'Walsh, William E']",2005,"A major goal of autonomic computing is to dynamically allocate computational resources so as to continually optimize high-level policy objectives. A key challenge to achieving this goal is to accurately estimate the impact of resource-level changes on application performance with respect to a Service Level Ageement (SLA). We compare two methodologies for accomplishing this:(i) developing a queuing-theoretic performance model for an application, and fitting its parameters online based on current state;(ii) using modelfree reinforcement …",https://www.researchgate.net/profile/Rajarshi_Das/publication/246843416_Model-Based_and_Model-Free_Approaches_to_Autonomic_Resource_Allocation/links/57fd9dd308ae6750f80662fb.pdf,True,,21.0,,,,,,,True,['resource-provisioning'],,,,,,,,,
655,Google hostload prediction based on Bayesian model with optimized feature combination,"['Sheng Di', 'Derrick Kondo', 'Walfredo Cirne']",2014,"We design a novel prediction method with Bayes model to predict a load fluctuation pattern over a long-term interval, in the context of Google data centers. We exploit a set of features that capture the expectation, trend, stability and patterns of recent host loads. We also investigate the correlations among these features and explore the most effective combinations of features with various training periods. All of the prediction methods are evaluated using Google trace with 10,000+ heterogeneous hosts. Experiments show that our Bayes method improves the long-term load prediction accuracy by 5.6%-50%, compared to other state-of-the-art methods based on moving average, auto-regression, and/or noise filters. Mean squared error of pattern prediction with Bayes method can be approximately limited in 1 0 - 8 , 1 0 - 5 . Through a load balancing scenario, we confirm the precision of pattern prediction in finding a set of idlest/busiest hosts from among 10,000+ hosts can be improved by about 7% on average. We devise an exponentially segmented pattern model for the hostload prediction.We devise a Bayes method and exploit 10 features to find the best-fit combination.We evaluate the Bayes method and 8 other well-known load prediction methods.The experiment is based on Google trace with over 10 k hosts and millions of jobs.The pattern prediction with Bayes method has much higher precision than others.",https://dl.acm.org/doi/10.1016/j.jpdc.2013.10.001,True,,49.0,,,,['workload-prediction'],,Journal of Parallel and Distributed Computing,True,['resource-provisioning'],,,,,,10.1016/j.jpdc.2013.10.001,,,
656,Workload Characterization for Capacity Planning and Performance Management in IaaS Cloud,"['Shruti Mahambre', 'Purushottam Kulkarni', 'Umesh Bellur', 'Girish Chafle', 'Deepak Deshpande']",2012,"Effective characterization of workload could be used to drive Capacity Planning and Performance Management in IaaS Cloud. There are different workload metrics (e.g. CPU, memory usage, throughput, response time) which could be modeled along with relationships between them. Similarly, we could model relationships across a set of workloads. Analyzing and characterizing this would enable decision making for various scenarios such as migration, re-provisioning, load balancing, resource management, initial placement. In this paper, we study workload running in IaaS cloud and categorize into patterns, based on their behavioral characteristics. We define different types of behavioral patterns and outline statistical techniques to be used in determining these patterns. We present initial results for development workload data collected in the lab.",https://ieeexplore.ieee.org/document/6354624,True,,18.0,,,,,,2012 IEEE International Conference on Cloud Computing in Emerging Markets (CCEM),True,['resource-provisioning'],,,,,,,,,
657,Efficient Autoscaling in the Cloud using Predictive Models for Workload Forecasting,"['Nilabja Roy', 'Abhishek Dubey', 'Aniruddha Gokhale']",2011,"Large-scale component-based enterprise applications that leverage Cloud resources expect Quality of Service(QoS) guarantees in accordance with service level agreements between the customer and service providers. In the context of Cloud computing, auto scaling mechanisms hold the promise of assuring QoS properties to the applications while simultaneously making efficient use of resources and keeping operational costs low for the service providers. Despite the perceived advantages of auto scaling, realizing the full potential of auto scaling is hard due to multiple challenges stemming from the need to precisely estimate resource usage in the face of significant variability in client workload patterns. This paper makes three contributions to overcome the general lack of effective techniques for workload forecasting and optimal resource allocation. First, it discusses the challenges involved in auto scaling in the cloud. Second, it develops a model-predictive algorithm for workload forecasting that is used for resource auto scaling. Finally, empirical results are provided that demonstrate that resources can be allocated and deal located by our algorithm in a way that satisfies both the application QoS while keeping operational costs low.",https://ieeexplore.ieee.org/document/6008748,True,,502.0,,,['new-method'],"['resource-consolidation', 'workload-prediction']",,2011 IEEE 4th International Conference on Cloud Computing,True,['resource-provisioning'],,,,,,,,,
658,"Mistral: Dynamically Managing Power, Performance, and Adaptation Cost in Cloud Infrastructures","['Gueyoung Jung', 'Matti A. Hiltunen', 'Kaustubh R. Joshi', 'Richard D. Schlichting', 'Calton Pu']",2010,"Server consolidation based on virtualization is an important technique for improving power efficiency and resource utilization in cloud infrastructures. However, to ensure satisfactory performance on shared resources under changing application workloads, dynamic management of the resource pool via online adaptation is critical. The inherent tradeoffs between power and performance as well as between the cost of an adaptation and its benefits make such management challenging. In this paper, we present Mistral, a holistic controller framework that optimizes power consumption, performance benefits, and the transient costs incurred by various adaptations and the controller itself to maximize overall utility. Mistral can handle multiple distributed applications and large-scale infrastructures through a multi-level adaptation hierarchy and scalable optimization algorithm. We show that our approach outstrips other strategies that address the tradeoff between only two of the objectives (power, performance, and transient costs).",https://ieeexplore.ieee.org/document/5541703,True,,138.0,['optimization'],,,"['resource-consolidation', 'power-management']",,2010 IEEE 30th International Conference on Distributed Computing Systems,True,['resource-provisioning'],,,,,,,,,
659,Dynamic Placement of Virtual Machines for Managing SLA Violations,"['Norman Bobroff', 'Andrzej Kochut', 'Kirk Beaty']",2007,"A dynamic server migration and consolidation algorithm is introduced. The algorithm is shown to provide substantial improvement over static server consolidation in reducing the amount of required capacity and the rate of service level agreement violations. Benefits accrue for workloads that are variable and can be forecast over intervals shorter than the time scale of demand variability. The management algorithm reduces the amount of physical capacity required to support a specified rate of SLA violations for a given workload by as much as 50% as compared to static consolidation approach. Another result is that the rate of SLA violations at fixed capacity may be reduced by up to 20%. The results are based on hundreds of production workload traces across a variety of operating systems, applications, and industries.",https://ieeexplore.ieee.org/document/4258528,True,,996.0,['autoregression'],['host-metrics'],['new-method'],['workload-prediction'],"['server', 'vm']",2007 10th IFIP/IEEE International Symposium on Integrated Network Management,True,['resource-provisioning'],True,,,,,,,,
660,Optimizing Resource Consumptions in Clouds,"['Ligang He', 'Deqing Zou', 'Zhang Zhang', 'Kai Yang', 'Hai Jin', 'Stephen A. Jarvis']",2011,"This paper considers the scenario where multiple clusters of Virtual Machines (i.e., termed as Virtual Clusters) are hosted in a Cloud system consisting of a cluster of physical nodes. Multiple Virtual Clusters (VCs) cohabit in the physical cluster, with each VC offering a particular type of service for the incoming requests. In this context, VM consolidation, which strives to use a minimal number of nodes to accommodate all VMs in the system, plays an important role in saving resource consumption. Most existing consolidation methods proposed in the literature regard VMs as ""rigid"" during consolidation, i.e., VMs' resource capacities remain unchanged. In VC environments, QoS is usually delivered by a VC as a single entity. Therefore, there is no reason why VMs' resource capacity cannot be adjusted as long as the whole VC is still able to maintain the desired QoS. Treating VMs as being ""mouldable"" during consolidation may be able to further consolidate VMs into an even fewer number of nodes. This paper investigates this issue and develops a Genetic Algorithm (GA) to consolidate mouldable VMs. The GA is able to evolve an optimized system state, which represents the VM-to-node mapping and the resource capacity allocated to each VM. After the new system state is calculated by the GA, the Cloud will transit from the current system state to the new one. The transition time represents overhead and should be minimized. In this paper, a cost model is formalized to capture the transition overhead, and a reconfiguration algorithm is developed to transit the Cloud to the optimized system state at the low transition overhead. Experiments have been conducted in this paper to evaluate the performance of the GA and the reconfiguration algorithm.",https://ieeexplore.ieee.org/document/6076497,True,,18.0,,,,,,2011 IEEE/ACM 12th International Conference on Grid Computing,True,['resource-provisioning'],,,,,,,,,
661,Energy-aware service allocation,"['Borgetto, Damien', 'Casanova, Henri', 'Da Costa, Georges', 'Pierson, Jean-Marc']",2012,"In this paper we study the problem of energy-aware resource allocation for hosting long-term services or on-demand computing jobs in clusters, eg, deployed as part of computing infrastructures. We formalize the problem as three constrained optimization problems …",https://www.sciencedirect.com/science/article/abs/pii/S0167739X11000690?via%3Dihub,True,,77.0,,,,['power-management'],,,True,['resource-provisioning'],,,,,,,,,
662,Algorithms for Event-Driven Application Brownout,"['Desmeurs, David']",2015,"Existing problems in cloud data centers include hardware failures, unexpected peaks of incoming requests, or waste of energy due to low utilization and lack of energy proportionality, which all lead to resource shortages and as a result, application problems such as delays or crashes. A paradigm called Brownout has been designed to counteract these problems by automatically activating or deactivating optional computations in cloud applications. When optional computations are deactivated, the capacity requirement is …",http://umu.diva-portal.org/smash/record.jsf?pid=diva2%3A801818&dswid=-2932,True,,2.0,,,,,,,True,['resource-provisioning'],,,,,,,,,
663,Efficient Decision-Making under Uncertainty for Proactive Self-Adaptation,"['Gabriel A. Moreno', 'Javier Cámara', 'David Garlan', 'Bradley Schmerl']",2016,"Proactive latency-aware adaptation is an approach for self-adaptive systems that improves over reactive adaptation by considering both the current and anticipated adaptation needs of the system, and taking into account the latency of adaptation tactics so that they can be started with the necessary lead time. Making an adaptation decision with these characteristics requires solving an optimization problem to select the adaptation path that maximizes an objective function over a finite look-ahead horizon. Since this is a problem of selecting adaptation actions in the context of the probabilistic behavior of the environment, Markov decision processes (MDP) are a suitable approach. However, given all the possible interactions between the different and possibly concurrent adaptation tactics, the system, and the environment, constructing the MDP is a complex task. Probabilistic model checking can be used to deal with this problem since it takes as input a formal specification of the stochastic system, which is internally translated into an MDP, and solved. One drawback of this solution is that the MDP has to be constructed every time an adaptation decision has to be made to incorporate the latest predictions of the environment behavior. In this paper we present an approach that eliminates that run-time overhead by constructing most of the MDP offline, also using formal specification. At run time, the adaptation decision is made by solving the MDP through stochastic dynamic programming, weaving in the stochastic environment model as the solution is computed. Our experimental results show that this approach reduces the adaptation decision time by an order of magnitude compared to the probabilistic model checking approach, while producing the same results.",https://ieeexplore.ieee.org/document/7573126,True,,13.0,,,,,,2016 IEEE International Conference on Autonomic Computing (ICAC),True,['resource-provisioning'],,,,,,,,,
664,Applying reinforcement learning towards automating resource allocation and application scalability in the cloud,"['Barrett, Enda', 'Howley, Enda', 'Duggan, Jim']",2012,"Public Infrastructure as a Service (IaaS) clouds such as Amazon, GoGrid and Rackspace deliver computational resources by means of virtualisation technologies. These technologies allow multiple independent virtual machines to reside in apparent isolation on the same physical host. Dynamically scaling applications running on IaaS clouds can lead to varied and unpredictable results because of the performance interference effects associated with co‐located virtual machines. Determining appropriate scaling policies in a dynamic non‐stationary environment is non‐trivial. One principle advantage exhibited by IaaS clouds over their traditional hosting counterparts is the ability to scale resources on‐demand. However, a problem arises concerning resource allocation as to which resources should be added and removed when the underlying performance of the resource is in a constant state of flux. Decision theoretic frameworks such as Markov Decision Processes are particularly suited to decision making under uncertainty. By applying a temporal difference, reinforcement learning algorithm known as Q‐learning, optimal scaling policies can be determined. Additionally, reinforcement learning techniques typically suffer from curse of dimensionality problems, where the state space grows exponentially with each additional state variable. To address this challenge, we also present a novel parallel Q‐learning approach aimed at reducing the time taken to determine optimal policies whilst learning online.",https://onlinelibrary.wiley.com/doi/abs/10.1002/cpe.2864,True,,168.0,['reinforcement-learning'],,['novel-use'],['resource-consolidation'],,,True,['resource-provisioning'],True,,,,,,,,
665,Using Reinforcement Learning for Autonomic Resource Allocation in Clouds: Towards a Fully Automated Workflow,"['Dutreilh, Xavier', 'Kirgizov, Sergey', 'Melekhova, Olga', 'Malenfant, Jacques', 'Rivierre, Nicolas', 'Truck, Isis']",2011,"Dynamic and appropriate resource dimensioning is a crucial issue in cloud computing. As applications go more and more 24/7, online policies must be sought to balance performance with the cost of allocated virtual machines. Most industrial approaches to date use ad hoc manual policies, such as thresholdbased ones. Providing good thresholds proved to be tricky and hard to automatize to fit every application requirement. Research is being done to apply automatic decision-making approaches, such as reinforcement learning. Yet, they face …",https://www.researchgate.net/publication/267990933_Using_Reinforcement_Learning_for_Autonomic_Resource_Allocation_in_Clouds_Towards_a_Fully_Automated_Workflow,True,,128.0,['reinforcement-learning'],,,['resource-consolidation'],,"ICAS 2011, The Seventh International Conference on Autonomic and Autonomous Systems",True,['resource-provisioning'],,,,,,,,,
666,NASLA: Novel Auto Scaling Approach based on Learning Automata for Web Application in Cloud Computing Environment,"['Fallah, Monireh', 'Arani, Mostafa Ghobaei', 'Maeen, Mehrdad']",2015,"Considering the growing interest in using cloud services, the accessibility and the effective management of the required resources, irrespective of the time and place, seems to be of great importance both to the service providers and users. One of the best ways for increasing utilization and improving the performance of the cloud systems is the auto-scaling of the applications; this is because of the fact that, due to the scalability of cloud computing, on the one hand, cloud providers believe that sufficient resources have to be prepared for …",https://www.researchgate.net/publication/273758585_NASLA_Novel_Auto_Scaling_Approach_based_on_Learning_Automata_for_Web_Application_in_Cloud_Computing_Environment,True,,25.0,,,,['resource-consolidation'],,,True,['resource-provisioning'],,,,,,,,,
667,An autonomic resource provisioning approach for service-based cloud applications: A hybrid approach,"['Ghobaei-Arani, Mostafa', 'Jabbehdari, Sam', 'Pourmina, Mohammad Ali']",2018,"In cloud computing environment, resources can be dynamically provisioned on deman for cloud services The amount of the resources to be provisioned is determined during runtime according to the workload changes. Deciding the right amount of resources required to run the cloud services is not trivial, and it depends on the current workload of the cloud services. Therefore, it is necessary to predict the future demands to automatically provision resources in order to deal with fluctuating demands of the cloud services. In this paper, we propose a hybrid resource provisioning approach for cloud services that is based on a combination of the concept of the autonomic computing and the reinforcement learning (RL). Also, we present a framework for autonomic resource provisioning which is inspired by the cloud layer model. Finally, we evaluate the effectiveness of our approach under two real world workload traces. The experimental results show that the proposed approach reduces the total cost by up to 50%, and increases the resource utilization by up to 12% compared with the other approaches.",https://www.sciencedirect.com/science/article/abs/pii/S0167739X17302327?via%3Dihub,True,,62.0,['reinforcement-learning'],,,['resource-consolidation'],,,True,['resource-provisioning'],,,,,,,,,
668,Energy-aware server provisioning and load dispatching for connection-intensive internet services,"['Gong Chen', 'Wenbo He', 'Jie Liu', 'Suman Nath', 'Leonidas Rigas', 'Lin Xiao', 'Feng Zhao']",2008,"Energy consumption in hosting Internet services is becoming a pressing issue as these services scale up. Dynamic server provisioning techniques are effective in turning off unnecessary servers to save energy. Such techniques, mostly studied for request-response services, face challenges in the context of connection servers that host a large number of long-lived TCP connections. In this paper, we characterize unique properties, performance, and power models of connection servers, based on a real data trace collected from the deployed Windows Live Messenger. Using the models, we design server provisioning and load dispatching algorithms and study subtle interactions between them. We show that our algorithms can save a significant amount of energy without sacrificing user experiences.",https://dl.acm.org/doi/10.5555/1387589.1387613,True,,130.0,['autoregression'],"['requests', 'sla', 'host-metrics']",['new-method'],"['resource-consolidation', 'power-management', 'workload-prediction']",,NSDI'08: Proceedings of the 5th USENIX Symposium on Networked Systems Design and Implementation,True,['resource-provisioning'],,,,,,10.5555/1387589.1387613,,,
669,Dynamic Provisioning Modeling for Virtualized Multi-tier Applications in Cloud Data Center,"['Jing Bi', 'Zhiliang Zhu', 'Ruixiong Tian', 'Qingbo Wang']",2010,"Dynamic provisioning is a useful technique for handling the virtualized multi-tier applications in cloud environment. Understanding the performance of virtualized multi-tier applications is crucial for efficient cloud infrastructure management. In this paper, we present a novel dynamic provisioning technique for a cluster-based virtualized multi-tier application that employ a flexible hybrid queueing model to determine the number of virtual machines at each tier in a virtualized application. We present a cloud data center based on virtual machine to optimize resources provisioning. Using simulation experiments of three-tier application, we adopt an optimization model to minimize the total number of virtual machines while satisfying the customer average response time constraint and the request arrival rate constraint. Our experiments show that cloud data center resources can be allocated accurately with these techniques, and the extra cost can be effectively reduced.",https://ieeexplore.ieee.org/document/5557972,True,,72.0,,,,['resource-consolidation'],['vm'],,True,['resource-provisioning'],,,,,,,,,
670,Kriging Controllers for Cloud Applications,"['Alessio Gambi', 'Giovanni Toffetti', 'Cesare Pautasso', 'Mauro Pezzè']",2012,"The infrastructure-as-a-service paradigm for cloud computing lets service providers execute applications on third-party infrastructures with a pay-as-you-go billing model. Providers can balance operational costs and quality of service by monitoring application behavior and changing the deployed configuration at runtime as operating conditions change. Current approaches for automatically scaling cloud applications exploit user-defined rules that respond well to predictable events but don't react adequately to unexpected execution conditions. The authors' autonomic controllers, designed using Kriging models, automatically adapt to unpredicted conditions by dynamically updating a model of the system's behavior.",https://ieeexplore.ieee.org/document/6357171,True,,25.0,,,,,,,True,['resource-provisioning'],,,,,,,,,
671,"Adaptive, Model-driven Autoscaling for Cloud Applications","['Gandhi, Anshul', 'Dube, Parijat', 'Karve, Alexei', 'Kochut, Andrzej', 'Zhang, Li']",2014,"Applications with a dynamic workload demand need access to a flexible infrastructure to meet performance guarantees and minimize resource costs. While cloud computing provides the elasticity to scale the infrastructure on demand, cloud service providers lack control and visibility of user space applications, making it difficult to accurately scale the underlying infrastructure. Thus, the burden of scaling falls on the user. In this paper, we propose a new cloud service, Dependable Compute Cloud (DC2), that automatically scales the infrastructure to meet the user-specified performance requirements. DC2 employs Kalman filtering to automatically learn the (possibly changing) system parameters for each application, allowing it to proactively scale the infrastructure to meet performance guarantees. DC2 is designed for the cloud it is application-agnostic and does not require any offline application profiling or benchmarking. Our implementation results on OpenStack using a multi-tier application under a range of workload traces demonstrate the robustness and superiority of DC2 over existing rule-based approaches.",https://www.semanticscholar.org/paper/Adaptive%2C-Model-driven-Autoscaling-for-Cloud-Gandhi-Dube/e8202368a8143f295111f0a834c54977feeae3bf,True,,111.0,['kalman-filter'],,,['resource-consolidation'],['openstack'],11th International Conference on Autonomic Computing ($\{$ICAC$\}$ 14),True,['resource-provisioning'],,,,,,,,,
672,A Regression-Based Analytic Model for Dynamic Resource Provisioning of Multi-Tier Applications,"['Qi Zhang', 'Ludmila Cherkasova', 'Evgenia Smirni']",2007,"The multi-tier implementation has become the industry standard for developing scalable client-server enterprise applications. Since these applications are performance sensitive, effective models for dynamic resource provisioning and for delivering quality of service to these applications become critical. Workloads in such environments are characterized by client sessions of interdependent requests with changing transaction mix and load over time, making model adaptivity to the observed workload changes a critical requirement for model effectiveness. In this work, we apply a regression-based approximation of the CPU demand of client transactions on a given hardware. Then we use this approximation in an analytic model of a simple network of queues, each queue representing a tier, and show the approximation's effectiveness for modeling diverse workloads with a changing transaction mix over time. Using the TPCW benchmark and its three different transaction mixes we investigate factors that impact the efficiency and accuracy of the proposed performance prediction models. Experimental results show that this regression-based approach provides a simple and powerful solution for efficient capacity planning and resource provisioning of multi-tier applications under changing workload conditions.",https://dl.acm.org/doi/10.1109/ICAC.2007.1,True,,65.0,,,,['resource-consolidation'],,ICAC '07: Proceedings of the Fourth International Conference on Autonomic Computing,True,['resource-provisioning'],,,,,,10.1109/ICAC.2007.1,,,
673,Kriging-Based Self-Adaptive Cloud Controllers,"['Alessio Gambi', 'Mauro Pezzè', 'Giovanni Toffetti']",2015,"Cloud technology is rapidly substituting classic computing solutions, and challenges the community with new problems. In this paper we focus on controllers for cloud application elasticity, and propose a novel solution for self-adaptive cloud controllers based on Kriging models. Cloud controllers are application specific schedulers that allocate resources to applications running in the cloud, aiming to meet the quality of service requirements while optimizing the execution costs. General-purpose cloud resource schedulers provide sub-optimal solutions to the problem with respect to application-specific solutions that we call cloud controllers. In this paper we discuss a general way to design self-adaptive cloud controllers based on Kriging models. We present Kriging models, and show how they can be used for building efficient controllers thanks to their unique characteristics. We report experimental data that confirm the suitability of Kriging models to support efficient cloud control and open the way to the development of a new generation of cloud controllers.",https://ieeexplore.ieee.org/document/7004890,True,,20.0,,,,,,,True,['resource-provisioning'],,,,,,,,,
674,Integrated estimation and tracking of performance model parameters with autoregressive trends,"['Tao Zheng', 'Marin Litoiu', 'Murray Woodside']",2011,"Adaptive management of a software service system can take advantage of a performance model which can predict the effect of proposed changes, before they are deployed. As the system varies over time the model parameters can be tracked by an estimator such as a Kalman Filter, so that decisions can be updated. The filter is valuable when parameters are 'hidden' and cannot be directly measured without excessive cost (as is usually the case for the CPU time of a service). Because there may be significant delays in some management control actions (especially in deploying a new replica of a service), it is also important to be able to predict the changes ahead somewhat in time, that is, to predict the trends. The trend predictor itself needs to be estimated from observed trends in the model parameters. This work uses an autoregressive model for trend prediction and integrates it with the parameter estimator, in a single Kalman Filter, using auxiliary states for the parameter evolution process. This paper describes how the trend model is constructed, and evaluates its effectiveness. It compares the overall performance predictions to a simpler trend predictor using linear extrapolation of the fitted parameter time-series, which turns out to be almost as good. The approach is validated on a real system running a benchmark web application.",https://dl.acm.org/doi/10.1145/1958746.1958772,True,,21.0,,,,,,ICPE '11: Proceedings of the 2nd ACM/SPEC International Conference on Performance engineering,True,['resource-provisioning'],,,,,,10.1145/1958746.1958772,,,
675,Cloud Application Resource Mapping and Scaling Based on Monitoring of QoS Constraints,"['Collazo-Mojica, Xabriel J', 'Sadjadi, Seyed Masoud', 'Ejarque, Jorge', 'Badia, Rosa M']",2012,"Infrastructure as a Service (IaaS) clouds promise unlimited raw computing resources on-demand. However, the performance and granularity of these resources can vary widely between providers. Cloud computing users, such as Web developers, can benefit from a service which automatically maps performance non-functional requirements to these resources. We propose a SOA API, in which users provide a cloud application model and get back possible resource allocations in an IaaS provider. The solution emphasizes the …",https://www.semanticscholar.org/paper/Cloud-Application-Resource-Mapping-and-Scaling-on-Collazo-Mojica-Sadjadi/be202ef18fcd43f7fd0fe8153525a7f361cbf07d,True,,13.0,,,,,,SEKE,True,['resource-provisioning'],,,,,,,,,
676,Virtual Machine Resource Allocation in Cloud Computing via Multi-Agent Fuzzy Control,"['Dorian Minarolli', 'Bernd Freisleben']",2013,"Dynamic resource (re)-allocation for virtual machines in cloud computing is important to guarantee application performance and to reduce operating costs. The problem is to find an adequate trade-off between these two conflicting goals. An approach is presented to support the Virtual Machine Monitor in performing resource allocation of VMs running on a physical machine of a cloud provider by expressing the two objectives in a utility function and optimizing this function using fuzzy control. To potentially work for an increased number of virtual machines, a multi-agent fuzzy controller is realized where each agent optimizes its own local utility function. The multi-agent fuzzy controller is empirically compared to a centralized fuzzy controller and an adaptive optimal control approach. Experimental results show the effectiveness of the multi-agent fuzzy controller in finding an adequate trade-off between performance and cost.",https://ieeexplore.ieee.org/document/6686027?section=abstract,True,,23.0,,,,,,2013 International Conference on Cloud and Green Computing,True,['resource-provisioning'],,,,,,,,,
677,Adaptive Resource Provisioning for Virtualized Servers Using Kalman Filters,"['Evangelia Kalyvianaki', 'Themistoklis Charalambous', 'Steven Hand']",2014,"Resource management of virtualized servers in data centers has become a critical task, since it enables cost-effective consolidation of server applications. Resource management is an important and challenging task, especially for multitier applications with unpredictable time-varying workloads. Work in resource management using control theory has shown clear benefits of dynamically adjusting resource allocations to match fluctuating workloads. However, little work has been done toward adaptive controllers for unknown workload types. This work presents a new resource management scheme that incorporates the Kalman filter into feedback controllers to dynamically allocate CPU resources to virtual machines hosting server applications. We present a set of controllers that continuously detect and self-adapt to unforeseen workload changes. Furthermore, our most advanced controller also self-configures itself without any a priori information and with a small 4.8% performance penalty in the case of high-intensity workload changes. In addition, our controllers are enhanced to deal with multitier server applications: by using the pair-wise resource coupling between tiers, they improve server response to large workload increases as compared to controllers with no such resource-coupling mechanism. Our approaches are evaluated and their performance is illustrated on a 3-tier Rubis benchmark website deployed on a prototype Xen-virtualized cluster.",https://dl.acm.org/doi/10.1145/2626290,True,,32.0,,,,['resource-consolidation'],,ACM Transactions on Autonomous and Adaptive Systems,True,['resource-provisioning'],,,,,,10.1145/2626290,,,
678,An Analysis of Performance Interference Effects in Virtual Environments,"['Younggyun Koh', 'Rob Knauerhase', 'Paul Brett', 'Mic Bowman', 'Zhihua Wen', 'Calton Pu']",2007,"Virtualization is an essential technology in modern datacenters. Despite advantages such as security isolation, fault isolation, and environment isolation, current virtualization techniques do not provide effective performance isolation between virtual machines (VMs). Specifically, hidden contention for physical resources impacts performance differently in different workload configurations, causing significant variance in observed system throughput. To this end, characterizing workloads that generate performance interference is important in order to maximize overall utility. In this paper, we study the effects of performance interference by looking at system-level workload characteristics. In a physical host, we allocate two VMs, each of which runs a sample application chosen from a wide range of benchmark and real-world workloads. For each combination, we collect performance metrics and runtime characteristics using an instrumented Ken hypervisor. Through subsequent analysis of collected data, we identify clusters of applications that generate certain types of performance interference. Furthermore, we develop mathematical models to predict the performance of a new application from its workload characteristics. Our evaluation shows our techniques were able to predict performance with average error of approximately 5%",https://ieeexplore.ieee.org/document/4211036,True,,378.0,"['linear-regression', 'clustering']","['kpis', 'host-metrics']",,['workload-prediction'],['vm'],2007 IEEE International Symposium on Performance Analysis of Systems \& Software,True,['resource-provisioning'],,,,,,,,,
679,Q-clouds: managing performance interference effects for QoS-aware clouds,"['Ripal Nathuji', 'Aman Kansal', 'Alireza Ghaffarkhah']",2010,"Cloud computing offers users the ability to access large pools of computational and storage resources on demand. Multiple commercial clouds already allow businesses to replace, or supplement, privately owned IT assets, alleviating them from the burden of managing and maintaining these facilities. However, there are issues that must be addressed before this vision of utility computing can be fully realized. In existing systems, customers are charged based upon the amount of resources used or reserved, but no guarantees are made regarding the application level performance or quality-of-service (QoS) that the given resources will provide. As cloud providers continue to utilize virtualization technologies in their systems, this can become problematic. In particular, the consolidation of multiple customer applications onto multicore servers introduces performance interference between collocated workloads, significantly impacting application QoS. To address this challenge, we advocate that the cloud should transparently provision additional resources as necessary to achieve the performance that customers would have realized if they were running in isolation. Accordingly, we have developed Q-Clouds, a QoS-aware control framework that tunes resource allocations to mitigate performance interference effects. Q-Clouds uses online feedback to build a multi-input multi-output (MIMO) model that captures performance interference interactions, and uses it to perform closed loop resource management. In addition, we utilize this functionality to allow applications to specify multiple levels of QoS as application Q-states. For such applications, Q-Clouds dynamically provisions underutilized resources to enable elevated QoS levels, thereby improving system efficiency. Experimental evaluations of our solution using benchmark applications illustrate the benefits: performance interference is mitigated completely when feasible, and system utilization is improved by up to 35% using Q-states.",https://dl.acm.org/doi/10.1145/1755913.1755938,True,,357.0,,,,['resource-consolidation'],,EuroSys '10: Proceedings of the 5th European conference on Computer systems,True,['resource-provisioning'],,,,,,10.1145/1755913.1755938,,,
680,Mitigating interference in cloud services by middleware reconfiguration,"['Amiya K. Maji', 'Subrata Mitra', 'Bowen Zhou', 'Saurabh Bagchi', 'Akshat Verma']",2014,"Application performance has been and remains one of top five concerns since the inception of cloud computing. A primary determinant of application performance is multi-tenancy or sharing of hardware resources in clouds. While some hardware resources can be partitioned well among VMs (such as CPUs), many others cannot (such as memory bandwidth). In this paper, we focus on understanding the variability in application performance on a cloud and explore ways for an end customer to deal with it. Based on rigorous experiments using CloudSuite, a popular Web2.0 benchmark, running on EC2, we found that interference-induced performance degradation is a reality. On a private cloud testbed, we also observed that interference impacts the choice of best configuration values for applications and middleware. We posit that intelligent reconfiguration of application parameters presents a way for an end customer to reduce the impact of interference. However, tuning the application to deal with interference is challenging because of two fundamental reasons --- the configuration depends on the nature and degree of interference and there are inter-parameter dependencies. We design and implement the IC2 system (Interference-aware Cloud application Configuration) to address the challenges of detection and mitigation of performance interference in clouds. Compared to an interference-agnostic configuration, the proposed solution provides up to 29% and 40% improvement in average response time on EC2 and a private cloud testbed respectively.",https://dl.acm.org/doi/10.1145/2663165.2663330,True,,59.0,,,,['configuration'],"['cloud', 'vm']",Middleware '14: Proceedings of the 15th International Middleware Conference,True,['resource-provisioning'],,,,,,10.1145/2663165.2663330,,,
681,"The effects of scheduling, workload type and consolidation scenarios on virtual machine performance and their prediction through optimized artificial neural networks","['Kousiouris, George', 'Cucinotta, Tommaso', 'Varvarigou, Theodora']",2011,"The aim of this paper is to study and predict the effect of a number of critical parameters on the performance of virtual machines (VMs). These parameters include allocation percentages, real-time scheduling decisions and co-placement of VMs when these are deployed concurrently on the same physical node, as dictated by the server consolidation trend and the recent advances in the Cloud computing systems. Different combinations of VM workload types are investigated in relation to the aforementioned factors in order to find …",https://www.sciencedirect.com/science/article/pii/S0164121211000951?via%3Dihub,True,,103.0,,,,['workload-prediction'],,,True,['resource-provisioning'],,,,,,,,,
682,Matrix: Achieving Predictable Virtual Machine Performance in the Clouds,"['Chiang, Ron C', 'Hwang, Jinho', 'Huang, H Howie', 'Wood, Timothy']",2014,"The success of cloud computing builds largely upon on-demand supply of virtual machines (VMs) that provide the abstraction of a physical machine on shared resources. Unfortunately, despite recent advances in virtualization technology, there still exists an unpredictable performance gap between the real and desired performance. The main contributing factors include contention to the shared physical resources among co-located VMs, limited control of VM allocation, as well as lack of knowledge on the performance of a specific VM out of …",https://www.researchgate.net/publication/273568940_Matrix_Achieving_Predictable_Virtual_Machine_Performance_in_the_Clouds,True,,20.0,,,,,,11th International Conference on Autonomic Computing ($\{$ICAC$\}$ 14),True,['resource-provisioning'],,,,,,,,,
683,Statistical machine learning makes automatic control practical for internet datacenters,"['Peter Bodík', 'Rean Griffith', 'Charles Sutton', 'Armando Fox', 'Michael Jordan', 'David Patterson']",2009,"Horizontally-scalable Internet services on clusters of commodity computers appear to be a great fit for automatic control: there is a target output (service-level agreement), observed output (actual latency), and gain controller (adjusting the number of servers). Yet few datacenters are automated this way in practice, due in part to well-founded skepticism about whether the simple models often used in the research literature can capture complex real-life workload/performance relationships and keep up with changing conditions that might invalidate the models. We argue that these shortcomings can be fixed by importing modeling, control, and analysis techniques from statistics and machine learning. In particular, we apply rich statistical models of the application's performance, simulation-based methods for finding an optimal control policy, and change-point methods to find abrupt changes in performance. Preliminary results running aWeb 2.0 benchmark application driven by real workload traces on Amazon's EC2 cloud show that our method can effectively control the number of servers, even in the face of performance anomalies.",https://dl.acm.org/doi/10.5555/1855533.1855545,True,,236.0,,,,"['resource-consolidation', 'workload-prediction']",,HotCloud'09: Proceedings of the 2009 conference on Hot topics in cloud computing,True,['resource-provisioning'],,,,,,10.5555/1855533.1855545,,,
684,Augmenting Elasticity Controllers for Improved Accuracy,"['Navaneeth Rameshan', 'Ying Liu', 'Leandro Navarro', 'Vladimir Vlassov']",2016,"Elastic resource provisioning is used to guarantee service level objectives (SLO) at reduced cost in a Cloud platform. However, performance interference in the hosting platform introduces uncertainty in the performance guarantees of provisioned services. Existing elasticity controllers are either unaware of this interference or over-provision resources to meet the SLO. In this paper, we show that assuming predictable performance of VMs in a multi-tenant environment to scale, will result in long periods of SLO violations. We augment the elasticity controller to be aware of interference and improve the convergence time of scaling without over provisioning. We perform experiments with Memcached and compare our solution against a baseline elasticity controller that is unaware of performance interference. Our results show that augmentation can reduce SLO violations by 65% or more and also save provisioning costs compared to an interference oblivious controller.",https://ieeexplore.ieee.org/document/7573123,True,,8.0,,,,,,2016 IEEE International Conference on Autonomic Computing (ICAC),True,['resource-provisioning'],,,,,,,,,
685,"Dynamic, behavioral-based estimation of resource provisioning based on high-level application terms in Cloud platforms","['George Kousiouris', 'Andreas Menychtas', 'Dimosthenis Kyriazis', 'Spyridon Gogouvitis', 'Theodora Varvarigou']",2014,"Delivering Internet-scale services and IT-enabled capabilities as computing utilities has been made feasible through the emergence of Cloud environments. While current approaches address a number of challenges such as quality of service, live migration and fault tolerance, which is of increasing importance, refers to the embedding of users' and applications' behaviour in the management processes of Clouds. The latter will allow for accurate estimation of the resource provision (for certain levels of service quality) with respect to the anticipated users' and applications' requirements. In this paper we present a two-level generic black-box approach for behavioral-based management across the Cloud layers (i.e., Software, Platform, Infrastructure): it provides estimates for resource attributes at a low level by analyzing information at a high level related to application terms (Translation level) while it predicts the anticipated user behaviour (Behavioral level). Patterns in high-level information are identified through a time series analysis, and are afterwards translated to low-level resource attributes with the use of Artificial Neural Networks. We demonstrate the added value and effectiveness of the Translation level through different application scenarios: namely FFMPEG encoding, real-time interactive e-Learning and a Wikipedia-type server. For the latter, we also validate the combined level model through a trace-driven simulation for identifying the overall error of the two-level approach. Translation from high level application terms to resource level attributes. Identification of usage patterns and prediction. Combined validation and trace driven simulation of the two-level prediction.",https://dl.acm.org/doi/10.5555/2748143.2748375,True,,68.0,,,,['workload-prediction'],,Future Generation Computer Systems,True,['resource-provisioning'],,,,,,10.5555/2748143.2748375,,,
686,Economical and Robust Provisioning of N-Tier Cloud Workloads: A Multi-level Control Approach,"['Pengcheng Xiong', 'Zhikui Wang', 'Simon Malkowski', 'Qingyang Wang', 'Deepal Jayasinghe', 'Calton Pu']",2011,"Resource provisioning for N-tier web applications in Clouds is non-trivial due to at least two reasons. First, there is an inherent optimization conflict between cost of resources and Service Level Agreement (SLA) compliance. Second, the resource demands of the multiple tiers can be different from each other, and varying along with the time. Resources have to be allocated to multiple (virtual) containers to minimize the total amount of resources while meeting the end-to-end performance requirements for the application. In this paper we address these two challenges through the combination of the resource controllers on both application and container levels. On the application level, a decision maker (i.e., an adaptive feedback controller) determines the total budget of the resources that are required for the application to meet SLA requirements as the workload varies. On the container level, a second controller partitions the total resource budget among the components of the applications to optimize the application performance (i.e., to minimize the round trip time). We evaluated our method with three different workload models -- open, closed, and semi-open - that were implemented in the RUBiS web application benchmark. Our evaluation indicates two major advantages of our method in comparison to previous approaches. First, fewer resources are provisioned to the applications to achieve the same performance. Second, our approach is robust enough to address various types of workloads with time-varying resource demand without reconfiguration.",https://ieeexplore.ieee.org/document/5961734,True,,75.0,,,,['resource-consolidation'],['container'],2011 31st International Conference on Distributed Computing Systems,True,['resource-provisioning'],,,,,,,,,
687,Model-driven optimal resource scaling in cloud,"['Anshul Gandhi', 'Parijat Dube', 'Alexei Karve', 'Andrzej Kochut', 'Li Zhang']",2018,"Cloud computing offers the flexibility to dynamically size the infrastructure in response to changes in workload demand. While both horizontal scaling and vertical scaling of infrastructure are supported by major cloud providers, these scaling options differ significantly in terms of their cost, provisioning time, and their impact on workload performance. Importantly, the efficacy of horizontal and vertical scaling critically depends on the workload characteristics, such as the workload's parallelizability and its core scalability. In today's cloud systems, the scaling decision is left to the users, requiring them to fully understand the trade-offs associated with the different scaling options. In this paper, we present our solution for optimizing the resource scaling of cloud deployments via implementation in OpenStack. The key component of our solution is the modeling engine that characterizes the workload and then quantitatively evaluates different scaling options for that workload. Our modeling engine leverages Amdahl's Law to model service timescaling in scale-up environments and queueing-theoretic concepts to model performance scaling in scale-out environments. We further employ Kalman filtering to account for inaccuracies in the model-based methodology and to dynamically track changes in the workload and cloud environment.",https://dl.acm.org/doi/10.1007/s10270-017-0584-y,True,,11.0,,,,['resource-consolidation'],,Software and Systems Modeling (SoSyM),True,['resource-provisioning'],,,,,,10.1007/s10270-017-0584-y,,,
688,Automated QoS-oriented cloud resource optimization using containers,"['Yu Sun', 'Jules White', 'Bo Li', 'Michael Walker', 'Hamilton Turner']",2017,"Optimizing the deployment of software in a cloud environment is one approach for maximizing system Quality-of-Service (QoS) and minimizing total cost. A traditional challenge to this optimization is the large amount of benchmarking required to optimize even simplistic cloud systems. This paper introduces $$\hbox {C}^2$$C2RAM, an new approach to enable rapid, optimized deployment of software onto a cloud environment by substantially reducing the number of benchmarks required. $$\hbox {C}^2$$C2RAM continues to perform some benchmarking, and therefore its predictions of application QoS metrics, such as throughput and latency, are very accurate. Our results show a maximum difference of 1.06 % between $$\hbox {C}^2$$C2RAM predicted QoS and empirically measured QoS. Moreover, $$\hbox {C}^2$$C2RAM can be provided with QoS requirements for each software in the system, and will ensure that each requirement is met before presenting a deployment plan.",https://dl.acm.org/doi/10.1007/s10515-016-0191-0,True,,10.0,,,,,,Automated Software Engineering,True,['resource-provisioning'],,,,,,10.1007/s10515-016-0191-0,,,
689,Online Response Time Optimization of Apache Web Server,"['Liu, Xue', 'Sha, Lui', 'Diao, Yixin', 'Froehlich, Steven', 'Hellerstein, Joseph L', 'Parekh, Sujay']",2003,"Properly optimizing the setting of configuration parameters can greatly improve performance, especially in the presence of changing workloads. This paper explores approaches to online optimization of the Apache web server, focusing on the MaxClients parameter (which controls the maximum number of workers). Using both empirical and analytic techniques, we show that MaxClients has a concave upward effect on response time and hence hill climbing techniques can be used to find the optimal value of MaxClients …",https://link.springer.com/chapter/10.1007%2F3-540-44884-5_25,True,,149.0,"['fuzzy-logic', 'linear-regression', 'optimization']",,,"['resource-consolidation', 'workload-prediction']",['apache'],International Workshop on Quality of Service,True,['resource-provisioning'],,,,,,,,,
690,SLA-based optimization of power and migration cost in cloud computing,"['Hadi Goudarzi', 'Mohammad Ghasemazar', 'Massoud Pedram']",2012,"Cloud computing systems (or hosting datacenters) have attracted a lot of attention in recent years. Utility computing, reliable data storage, and infrastructure-independent computing are example applications of such systems. Electrical energy cost of a cloud computing system is a strong function of the consolidation and migration techniques used to assign incoming clients to existing servers. Moreover, each client typically has a service level agreement (SLA), which specifies constraints on performance and/or quality of service that it receives from the system. These constraints result in a basic trade-off between the total energy cost and client satisfaction in the system. In this paper, a resource allocation problem is considered that aims to minimize the total energy cost of cloud computing system while meeting the specified client-level SLAs in a probabilistic sense. The cloud computing system pays penalty for the percentage of a client's requests that do not meet a specified upper bound on their service time. An efficient heuristic algorithm based on convex optimization and dynamic programming is presented to solve the aforesaid resource allocation problem. Simulation results demonstrate the effectiveness of the proposed algorithm compared to previous work.",https://ieeexplore.ieee.org/abstract/document/6217419,True,,149.0,['optimization'],"['sla', 'host-metrics']",,"['resource-consolidation', 'power-management']",['vm'],"2012 12th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (ccgrid 2012)",True,['resource-provisioning'],,,,,,,,,
691,Empirical prediction models for adaptive resource provisioning in the cloud,[],2012,"Cloud computing allows dynamic resource scaling for enterprise online transaction systems, one of the key characteristics that differentiates the cloud from the traditional computing paradigm. However, initializing a new virtual instance in a cloud is not instantaneous; cloud hosting platforms introduce several minutes delay in the hardware resource allocation. In this paper, we develop prediction-based resource measurement and provisioning strategies using Neural Network and Linear Regression to satisfy upcoming resource demands. Experimental results demonstrate that the proposed technique offers more adaptive resource management for applications hosted in the cloud environment, an important mechanism to achieve on-demand resource allocation in the cloud.",https://dl.acm.org/doi/10.1016/j.future.2011.05.027,True,,80.0,,,,['resource-consolidation'],,Future Generation Computer Systems,True,['resource-provisioning'],,,,,,10.1016/j.future.2011.05.027,,,
692,Resource prediction based on double exponential smoothing in cloud computing,"['Jinhui Huang', 'Chunlin Li', 'Jie Yu']",2012,"With the development of cloud computing, customers are more and more concerned with cost on the resources which are not free in the cloud. Cloud resource providers can offer users two payment plans, i.e., reservation and on-demand plans for resource provision. In general, cost on resources gained by reservation plan is cheaper than on-demand plan. So the accuracy of resource prediction is of importance. In this paper, we present a resource prediction model based on double exponential smoothing, which considers not only the current state of resources but also the history records. Experiments performed on CloudSim cloud simulator show that the proposed method has a better performance on prediction accuracy.",https://ieeexplore.ieee.org/document/6201461,True,,33.0,,,,['workload-prediction'],,"2012 2nd International Conference on Consumer Electronics, Communications and Networks (CECNet)",True,['resource-provisioning'],,,,,,,,,
693,Pattern matching based forecast of non-periodic repetitive behavior for cloud client,"['Eddy Caron', 'Frédéric Desprez', 'Adrian Muresan']",2011,"The Cloud phenomenon brings along the cost-saving benefit of dynamic scaling. As a result, the question of efficient resource scaling arises. Prediction is necessary as the virtual resources that Cloud computing uses have a setup time that is not negligible. We propose an approach to the problem of workload prediction based on identifying similar past occurrences of the current short-term workload history. We present in detail the Cloud client resource auto-scaling algorithm that uses the above approach to help when scaling decisions are made, as well as experimental results by using real-world Cloud client application traces. We also present an overall evaluation of this approach, its potential and usefulness for enabling efficient auto-scaling of Cloud user resources.",https://dl.acm.org/doi/abs/10.1007/s10723-010-9178-4,True,,60.0,,,,['workload-prediction'],,Journal of Grid Computing,True,['resource-provisioning'],,,,,,10.1007/s10723-010-9178-4,,,
694,Adaptive resource provisioning for read intensive multi-tier applications in the cloud,"['Waheed Iqbal', 'Matthew N. Dailey', 'David Carrera', 'Paul Janecek']",2011,"A Service-Level Agreement (SLA) provides surety for specific quality attributes to the consumers of services. However, current SLAs offered by cloud infrastructure providers do not address response time, which, from the user's point of view, is the most important quality attribute for Web applications. Satisfying a maximum average response time guarantee for Web applications is difficult for two main reasons: first, traffic patterns are highly dynamic and difficult to predict accurately; second, the complex nature of multi-tier Web applications increases the difficulty of identifying bottlenecks and resolving them automatically. This paper proposes a methodology and presents a working prototype system for automatic detection and resolution of bottlenecks in a multi-tier Web application hosted on a cloud in order to satisfy specific maximum response time requirements. It also proposes a method for identifying and retracting over-provisioned resources in multi-tier cloud-hosted Web applications. We demonstrate the feasibility of the approach in an experimental evaluation with a testbed EUCALYPTUS-based cloud and a synthetic workload. Automatic bottleneck detection and resolution under dynamic resource management has the potential to enable cloud infrastructure providers to provide SLAs for Web applications that guarantee specific response time requirements while minimizing resource utilization.",https://dl.acm.org/doi/10.1016/j.future.2010.10.016,True,,50.0,,,,"['failure-detection', 'anomaly-detection', 'resource-consolidation']",,Future Generation Computer Systems,True,"['resource-provisioning', 'failure-management']",,,,,,10.1016/j.future.2010.10.016,,,
695,A Reinforcement Learning Approach to Online Web Systems Auto-configuration,"['Xiangping Bu', 'Jia Rao', 'Cheng-Zhong Xu']",2009,"In a web system, configuration is crucial to the performance and service availability. It is a challenge, not only because of the dynamics of Internet traffic, but also the dynamic virtual machine environment the system tends to be run on. In this paper, we propose a reinforcement learning approach for autonomic configuration and reconfiguration of multi-tier web systems. It is able to adapt performance parameter settings not only to the change of workload, but also to the change of virtual machine configurations. The RL approach is enhanced with an efficient initialization policy to reduce the learning time for online decision. The approach is evaluated using TPC-W benchmark on a three-tier website hosted on a Xen-based virtual machine environment. Experiment results demonstrate that the approach can autoconfigure the web system dynamically in response to the change in both workload and VM resource. It can drive the system into a near-optimal configuration setting in less than 25 trial-and-error iterations.",https://dl.acm.org/doi/10.1109/ICDCS.2009.76,True,,82.0,['reinforcement-learning'],,,['configuration'],,ICDCS '09: Proceedings of the 2009 29th IEEE International Conference on Distributed Computing Systems,True,['resource-provisioning'],,,,,,10.1109/ICDCS.2009.76,,,
696,PERFUME: Power and performance guarantee with fuzzy MIMO control in virtualized servers,"['Palden Lama', 'Xiaobo Zhou']",2011,"It is important but challenging to assure the performance of multi-tier Internet applications with the power consumption cap of virtualized server clusters mainly due to system complexity of shared infrastructure and dynamic and bursty nature of workloads. This paper presents PERFUME, a system that simultaneously guarantees power and performance targets with flexible tradeoffs while assuring control accuracy and system stability. Based on the proposed fuzzy MIMO control technique, it accurately controls both the throughput and percentile-based response time of multi-tier applications due to its novel fuzzy modeling that integrates strengths of fuzzy logic, MIMO control and artificial neural network. It is self-adaptive to highly dynamic and bursty workloads due to online learning of control model parameters using a computationally efficient weighted recursive least-squares method. We implement PERFUME in a testbed of virtualized blade servers hosting two multi-tier RUBiS applications. Experimental results demonstrate its control accuracy, system stability, flexibility in selecting tradeoffs between conflicting targets and robustness against highly dynamic variation and burstiness in workloads. It outperforms a representative utility based approach in providing guarantee of the system throughput, percentile-based response time and power budget in the face of highly dynamic and bursty workloads.",https://ieeexplore.ieee.org/document/5931340,True,,39.0,,,,"['power-management', 'resource-consolidation']",,2011 IEEE Nineteenth IEEE International Workshop on Quality of Service,True,['resource-provisioning'],,,,,,,,,
697,Efficient resource provisioning in compute clouds via VM multiplexing,"['Xiaoqiao Meng', 'Canturk Isci', 'Jeffrey Kephart', 'Li Zhang', 'Eric Bouillet', 'Dimitrios Pendarakis']",2010,"Resource provisioning in compute clouds often require an estimate of the capacity needs of Virtual Machines (VMs). The estimated VM size is the basis for allocating resources commensurate with workload demand. In contrast to the traditional practice of estimating the VM sizes individually, we propose a joint-VM sizing approach in which multiple VMs are consolidated and provisioned, based on an estimate of their aggregate capacity needs. This new approach exploits statistical multiplexing among the workload patterns of multiple VMs, i.e., the peaks and valleys in one workload pattern do not necessarily coincide with the others. Thus, the unused resources of a low utilized VM can be directed to the other co-located VMs with high utilization. Compared to individual VM based provisioning, joint-VM sizing and provisioning may lead to much higher resource utilization. This paper presents three design modules to enable the concept in practice. Specifically, a performance constraint describing the capacity need of a VM for achieving a certain level of application performance; an algorithm for estimating the size of jointly provisioning VMs; a VM selection method that seeks to find good VM combinations for being provisioned together. We showcase that the proposed three modules can be seamlessly plugged into existing applications such as resource provisioning, and providing resource guarantees for VMs. The proposed algorithms and applications are evaluated by monitoring data collected from about 16 thousand VMs in commercial data centers. These evaluations reveal more than 45% improvements in terms of the overall resource utilization.",https://dl.acm.org/doi/10.1145/1809049.1809052,True,,220.0,,['sla'],,['resource-consolidation'],['vm'],ICAC '10: Proceedings of the 7th international conference on Autonomic computing,True,['resource-provisioning'],,,,,,10.1145/1809049.1809052,,,
698,Response time densities in generalised stochastic petri net models,"['Nicholas J. Dingle', 'Peter G. Harrison', 'William J. Knottenbelt']",2002,"Generalised Stochastic Petri nets (GSPNs) have been widely used to analyse the performance of hardware and software systems. This paper presents a novel technique for the numerical determination of response time densities in GSPN models. The technique places no structural restrictions on the models that can be analysed, and allows for the high-level specification of multiple source and destination markings, including any combination of tangible and vanishing markings. The technique is implemented using a scalable parallel Laplace transform inverter that employs a modified Laguerre inversion technique. We present numerical results, including a study of the full distribution of end-to-end response time in a GSPN model of the Courier communication protocol software. The numerical results are validated against simulation.",https://dl.acm.org/doi/abs/10.1145/584369.584377,True,,53.0,,,,,,WOSP '02: Proceedings of the 3rd international workshop on Software and performance,True,['resource-provisioning'],,,,,,10.1145/584369.584377,,,
699,Using magpie for request extraction and workload modelling,"['Paul Barham', 'Austin Donnelly', 'Rebecca Isaacs', 'Richard Mortier']",2004,"Tools to understand complex system behaviour are essential for many performance analysis and debugging tasks, yet there are many open research problems in their development. Magpie is a toolchain for automatically extracting a system's workload under realistic operating conditions. Using low-overhead instrumentation, we monitor the system to record fine-grained events generated by kernel, middleware and application components. The Magpie request extraction tool uses an application-specific event schema to correlate these events, and hence precisely capture the control flow and resource consumption of each and every request. By removing scheduling artefacts, whilst preserving causal dependencies, we obtain canonical request descriptions from which we can construct concise workload models suitable for performance prediction and change detection. In this paper we describe and evaluate the capability of Magpie to accurately extract requests and construct representative models of system behaviour.",https://dl.acm.org/doi/10.5555/1251254.1251272,True,,231.0,,,,['failure-detection'],,OSDI'04: Proceedings of the 6th conference on Symposium on Operating Systems Design & Implementation - Volume 6,True,['failure-management'],,False,['anomaly-detection'],,,10.5555/1251254.1251272,,,
700,An Improved Genetic Algorithm with Limited Iteration for Grid Scheduling,"['Hao Yin', 'Huilin Wu', 'Jiliu Zhou']",2007,"In grid environment the numbers of resources and tasks to be scheduled are usually variable. This kind of characteristics of grid makes the scheduling approach a complex optimization problem. Genetic algorithm (GA) has been widely used to solve these difficult NP-complete problems. However the conventional GA is too slow to be used in a realistic scheduling due to its time-consuming iteration. This paper presents an improved genetic algorithm for scheduling independent tasks in grid environment, which can increase search efficiency with limited number of iteration by improving the evolutionary process while meeting a feasible result.",https://ieeexplore.ieee.org/document/4293783,True,,31.0,,,,,,Sixth International Conference on Grid and Cooperative Computing (GCC 2007),True,['resource-provisioning'],,,,,,,,,
701,Genetic simulated annealing algorithm for task scheduling based on cloud computing environment,"['Guo-ning Gan', 'Ting-lei Huang', 'Shuai Gao']",2010,"Scheduling is a very important part of the cloud computing system. This paper introduces an optimized algorithm for task scheduling based on genetic simulated annealing algorithm in cloud computing and its implementation. Algorithm considers the QOS requirements of different type tasks, the QOS parameters are dealt with dimensionless. The algorithm efficiently completes tasks scheduling in the cloud computing environment computing.",https://ieeexplore.ieee.org/document/5655013?reload=true,True,,116.0,['genetic-programming'],['sla'],,['scheduling'],['cloud'],2010 International Conference on Intelligent Computing and Integrated Systems,True,['resource-provisioning'],,,,,,,,,
702,Independent Tasks Scheduling Based on Genetic Algorithm in Cloud Computing,"['Chenhong Zhao', 'Shanshan Zhang', 'Qingfeng Liu', 'Jian Xie', 'Jicheng Hu']",2009,"Task scheduling algorithm, which is an NP-completeness problem, plays a key role in cloud computing systems. In this paper, we propose an optimized algorithm based on genetic algorithm to schedule independent and divisible tasks adapting to different computation and memory requirements. We prompt the algorithm in heterogeneous systems, where resources (including CPUs) are of computational and communication heterogeneity. Dynamic scheduling is also in consideration. Though GA is designed to solve combinatorial optimization problem, it's inefficient for global optimization. So we conclude with further researches in optimized genetic algorithm.",https://ieeexplore.ieee.org/document/5301850,True,,196.0,['genetic-programming'],,['novel-use'],['scheduling'],,"2009 5th International Conference on Wireless Communications, Networking and Mobile Computing",True,['resource-provisioning'],,,,,,,,,
703,Improved Genetic Algorithms and List Scheduling Techniques for Independent Task Scheduling in Distributed Systems,"['Thanasis Loukopoulos', 'Petros Lampsas', 'Panos Sigalas']",2007,"Given a set of tasks with certain characteristics, e.g., data size, estimated execution time and a set of processing nodes with their own parameters, the goal of task scheduling is to allocate tasks at nodes so that the total makespan is minimized. The problem has been studied under various assumptions concerning task and node parameters with the resulting problem statements usually being NP-complete. List scheduling (LS) heuristics such as MaxMin and MinMin together with genetic algorithms (GAs) were applied in the past to find solutions. In this paper we investigate new heuristics for both the LS and the GA paradigm with the specific aim of improving the performance of the standard algorithms when task computations involve large data transfers. Experimental results under various environment assumptions illustrate the merits of the new algorithms.",https://ieeexplore.ieee.org/document/4420143,True,,4.0,,,,,,,True,['resource-provisioning'],,,,,,,,,
704,Genetic algorithm for grid scheduling using best rank power,"['Wael Abdulal', 'Omar Al Jadaan', 'Ahmad Jabas', 'S. Ramachandram']",2009,"The large computing capacity provided by grid systems is beneficial for solving complex problems by using many nodes of the grid at the same time. The usefulness of a grid system largely depends, among other factors, on the efficiency of the system regarding the allocation of jobs to grid resources. This paper proposes an Roulette Wheel Selection Genetic Algorithm using Best Rank Power(PRRWSGA) for scheduling independent tasks in the grid environment. The modified algorithm speeds up convergence and shortens the search time more than IRRWSGA, at the same time the heuristic initialization of initial population using MCT algorithm allow the algorithm to obtain a high quality feasible scheduling solution. The simulation results, show that PRRWSGA has better search time than both IRRWSGA and standard genetic algorithms. Real-world scheduling problems may utilize this algorithm for better results.",https://ieeexplore.ieee.org/document/5393679,True,,19.0,,,,,,2009 World Congress on Nature \& Biologically Inspired Computing (NaBIC),True,['resource-provisioning'],,,,,,,,,
705,A new algorithm for grid independent task schedule: Genetic simulated annealing,"['Jianqin Wang', 'Qingling Duan', 'Yuxin Jiang', 'Xiuna Zhu']",2010,"Task schedule is a critical issue of distributed computing. Foster et al. (2001) defined ""Grid problem"", which is defined as flexible, secure, coordinated resource sharing among dynamic collections of individuals, institutions, and resources -what they referred to as virtual organizations (VO). Improving the performance of grid computing relies much on the grid task scheduling algorithm. In this paper, a new genetic simulated annealing (GSA) algorithm which combines genetic algorithm (GA) with simulated annealing (SA) algorithm for grid task scheduling is proposed, it could avoid trapping in a local minimum effectively and get the global optimization at last. The algorithm performs better than genetic algorithm and simulated annealing algorithm respectively.",https://ieeexplore.ieee.org/document/5665473,True,,287.0,['genetic-programming'],,['novel-use'],['scheduling'],,2010 World Automation Congress,True,['resource-provisioning'],,,,,,,,,
706,GENETIC ALGORITHM FOR MULTIPROCESSOR TASK SCHEDULING,"['Verma, Ritu', 'Dhingra, Sunita']",2011,"Multiprocessor task scheduling (MPTS) is an important and computationally difficult problem. Multiprocessors have emerged as a powerful computing means for running real-time applications especially due to limitation of uni-processor system for not having sufficient enough capability to execute all the tasks. This paper describes multiprocessor task scheduling in the form of permutation flow shop scheduling, which has an objective function for minimizing the makespan. Here, we will conclude how the performance of genetic …",http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.208.3816,True,,10.0,,,,,,,True,['resource-provisioning'],,,,,,,,,
707,Cloud Computing—Task scheduling based on genetic algorithms,"['Eleonora Maria Mocanu', 'Mihai Florea', 'Mugurel Ionuţ Andreica', 'Nicolae Ţăpuş']",2012,"Cloud Computing is a cutting edge technology for managing and delivering services over the Internet. Map-Reduce is the programming model used in cloud computing for processing large data sets in parallel over huge clusters. In order to increase efficiency, a good task scheduling is needed. Genetic algorithms are very useful and accurate in finding solutions to large scale optimization problems, such as task scheduling. They have gained immense popularity over last few years as a robust and easily adaptable search technique. Hadoop, the open source implementation of Map-Reduce, has several task schedulers available (FIFO, Fair, Capacity Schedulers), but neither one of them is focused on minimizing the global execution time. The goal of this project is to improve Hadoop's functionality by implementing a scheduler based on a genetic algorithm, solving the stated problem.",https://ieeexplore.ieee.org/abstract/document/6189509,True,,38.0,,,,['scheduling'],,2012 IEEE International Systems Conference SysCon 2012,True,['resource-provisioning'],,,,,,,,,
708,Grid Workflow Scheduling based on improved genetic algorithm,"['Xue Zhang', 'Wenhua Zeng']",2010,"Grid Workflow Scheduling represented by DAG(Directed Acyclic Graph) is a typical NP-complete problem, and thus a scheduling algorithm of high efficiency is required. So an improved genetic algorithm is proposed to solve this problem. In the algorithm, chromosomes of poor fitness make secondary preferential hybridization and mutation with the overall best individual. It not only guarantees the population diversity but increases the convergence rate of population. Experiment results based on Gridsim prove it available and better than standard genetic algorithm.",https://ieeexplore.ieee.org/document/5541161,True,,7.0,,,,,,2010 International Conference On Computer Design and Applications,True,['resource-provisioning'],,,,,,,,,
709,iCBS: incremental cost-based scheduling under piecewise linear SLAs,"['Yun Chi', 'Hyun Jin Moon', 'Hakan Hacigümüş']",2011,"In a cloud computing environment, it is beneficial for the cloud service provider to offer differentiated services among different customers, who often have different cost profiles. Therefore, cost-aware scheduling of queries is important. A practical cost-aware scheduling algorithm must be able to handle the highly demanding query volumes in the scheduling queues to make online scheduling decisions very quickly. We develop such a highly efficient cost-aware query scheduling algorithm, called iCBS. iCBS takes the query costs derived from the service level agreements (SLAs) between the service provider and its customers into account to make cost-aware scheduling decisions. iCBS is an incremental variation of an existing scheduling algorithm, CBS. Although CBS exhibits an exceptionally good cost performance, it has a prohibitive time complexity. Our main contributions are (1) to observe how CBS behaves under piecewise linear SLAs, which are very common in cloud computing systems, and (2) to efficiently leverage these observations and to reduce the online time complexity from O(N) for the original version CBS to O(log2 N) for iCBS.",https://dl.acm.org/doi/abs/10.14778/2002938.2002942?download=true,True,,65.0,,,,['scheduling'],,Proceedings of the VLDB Endowment,True,['resource-provisioning'],,,,,,10.14778/2002938.2002942?download=true,,,
710,Characterizing tenant behavior for placement and crisis mitigation in multitenant DBMSs,"['Aaron J. Elmore', 'Sudipto Das', 'Alexander Pucher', 'Divyakant Agrawal', 'Amr El Abbadi', 'Xifeng Yan']",2013,"A multitenant database management system (DBMS) in the cloud must continuously monitor the trade-off between efficient resource sharing among multiple application databases (tenants) and their performance. Considering the scale of \attn{hundreds to} thousands of tenants in such multitenant DBMSs, manual approaches for continuous monitoring are not tenable. A self-managing controller of a multitenant DBMS faces several challenges. For instance, how to characterize a tenant given its variety of workloads, how to reduce the impact of tenant colocation, and how to detect and mitigate a performance crisis where one or more tenants' desired service level objective (SLO) is not achieved. We present Delphi, a self-managing system controller for a multitenant DBMS, and Pythia, a technique to learn behavior through observation and supervision using DBMS-agnostic database level performance measures. Pythia accurately learns tenant behavior even when multiple tenants share a database process, learns good and bad tenant consolidation plans (or packings), and maintains a pertenant history to detect behavior changes. Delphi detects performance crises, and leverages Pythia to suggests remedial actions using a hill-climbing search algorithm to identify a new tenant placement strategy to mitigate violating SLOs. Our evaluation using a variety of tenant types and workloads shows that Pythia can learn a tenant's behavior with more than 92% accuracy and learn the quality of packings with more than 86% accuracy. During a performance crisis, Delphi is able to reduce 99th percentile latencies by 80%, and can consolidate 45% more tenants than a greedy baseline, which balances tenant load without modeling tenant behavior.",https://dl.acm.org/doi/10.1145/2463676.2465308,True,,52.0,,,,['resource-consolidation'],,SIGMOD '13: Proceedings of the 2013 ACM SIGMOD International Conference on Management of Data,True,['resource-provisioning'],,,,,,10.1145/2463676.2465308,,,
711,Entropy: a Consolidation Manager for Clusters,"['Fabien Hermenier', 'Xavier Lorca', 'Jean-Marc Menaud', 'Gilles Muller', 'Julia Lawall']",2009,"Clusters provide powerful computing environments, but in practice much of this power goes to waste, due to the static allocation of tasks to nodes, regardless of their changing computational requirements. Dynamic consolidation is an approach that migrates tasks within a cluster as their computational requirements change, both to reduce the number of nodes that need to be active and to eliminate temporary overload situations. Previous dynamic consolidation strategies have relied on task placement heuristics that use only local optimization and typically do not take migration overhead into account. However, heuristics based on only local optimization may miss the globally optimal solution, resulting in unnecessary resource usage, and the overhead for migration may nullify the benefits of consolidation. In this paper, we propose the Entropy resource manager for homogeneous clusters, which performs dynamic consolidation based on constraint programming and takes migration overhead into account. The use of constraint programming allows Entropy to find mappings of tasks to nodes that are better than those found by heuristics based on local optimizations, and that are frequently globally optimal in the number of nodes. Because migration overhead is taken into account, Entropy chooses migrations that can be implemented efficiently, incurring a low performance overhead.",https://dl.acm.org/doi/10.1145/1508293.1508300,True,,311.0,['constraint-solving'],,,['resource-consolidation'],,VEE '09: Proceedings of the 2009 ACM SIGPLAN/SIGOPS international conference on Virtual execution environments,True,['resource-provisioning'],,,,,,10.1145/1508293.1508300,,,
712,A generic availability model for clustered computing systems,"['H. Sun', 'J.J. Han', 'H. Levendel']",2001,"We study the availability of a clustered computing system with one cluster manager and ""N+M"" processing nodes, where M processing nodes serve as spares for the N active processing nodes. The functionality of an individual processing node is dissected into application software, management software, OS and hardware. The dependency among these entities is considered. Stochastic Petri net models are constructed to investigate the cluster availability. In order to deal with a cluster of a very large size, a solution based on state aggregation and fixed-point iteration is proposed. The existence and uniqueness of the fixed point is proved. The impact of a cluster manager, switchover time and coverage ratio are quantitatively studied. From the numerical results of a simple cluster with ""2+1"" processing nodes, we find that: (1) the availability of the cluster manager does not have a significant impact on the system availability, (2) system availability increases with the coverage ratio and decreases with the switchover time. Mechanisms to improve the system availability are discussed.",https://ieeexplore.ieee.org/document/992704,True,,10.0,,,,,,Proceedings 2001 Pacific Rim International Symposium on Dependable Computing,True,['resource-provisioning'],,,,,,,,,
713,Middleware for Resource-Aware Deployment and Configuration of Fault-Tolerant Real-time Systems,"['Jaiganesh Balasubramanian', 'Aniruddha Gokhale', 'Abhishek Dubey', 'Friedhelm Wolf', 'Chenyang Lu', 'Chris Gill', 'Douglas Schmidt']",2010,"Developing large-scale distributed real-time and embedded (DRE) systems is hard in part due to complex deployment and configuration issues involved in satisfying multiple quality for service (QoS) properties, such as real-timeliness and fault tolerance. This paper makes three contributions to the study of deployment and configuration middleware for DRE systems that satisfy multiple QoS properties. First, it describes a novel task allocation algorithm for passively replicated DRE systems to meet their real-time and fault-tolerance QoS properties while consuming significantly less resources. Second, it presents the design of a strategizable allocation engine that enables application developers to evaluate different allocation algorithms. Third, it presents the design of a middleware agnostic configuration framework that uses allocation decisions to deploy application components/replicas and configure the underlying middleware automatically on the chosen nodes. These contributions are realized in the DeCoRAM (Deployment and Configuration Reasoning and Analysis via Modeling) middleware. Empirical results on a distributed testbed demonstrate DeCoRAM's ability to handle multiple failures and provide efficient and predictable real-time performance.",https://ieeexplore.ieee.org/document/5465967,True,,30.0,,,,,,2010 16th IEEE Real-Time and Embedded Technology and Applications Symposium,True,['resource-provisioning'],,,,,,,,,
714,Autonomic virtual resource management for service hosting platforms,"['Hien Nguyen Van', 'Frederic Dang Tran', 'Jean-Marc Menaud']",2009,"Cloud platforms host several independent applications on a shared resource pool with the ability to allocate computing power to applications on a per-demand basis. The use of server virtualization techniques for such platforms provide great flexibility with the ability to consolidate several virtual machines on the same physical server, to resize a virtual machine capacity and to migrate virtual machine across physical servers. A key challenge for cloud providers is to automate the management of virtual servers while taking into account both high-level QoS requirements of hosted applications and resource management costs. This paper proposes an autonomic resource manager to control the virtualized environment which decouples the provisioning of resources from the dynamic placement of virtual machines. This manager aims to optimize a global utility function which integrates both the degree of SLA fulfillment and the operating costs. We resort to a Constraint Programming approach to formulate and solve the optimization problem. Results obtained through simulations validate our approach.",https://dl.acm.org/doi/10.1109/CLOUD.2009.5071526,True,,32.0,,,,,,CLOUD '09: Proceedings of the 2009 ICSE Workshop on Software Engineering Challenges of Cloud Computing,True,['resource-provisioning'],,,,,,10.1109/CLOUD.2009.5071526,,,
715,Mapping parallelism to multi-cores: a machine learning based approach,"['Zheng Wang', ""Michael F.P. O'Boyle""]",2009,"The efficient mapping of program parallelism to multi-core processors is highly dependent on the underlying architecture. This paper proposes a portable and automatic compiler-based approach to mapping such parallelism using machine learning. It develops two predictors: a data sensitive and a data insensitive predictor to select the best mapping for parallel programs. They predict the number of threads and the scheduling policy for any given program using a model learnt off-line. By using low-cost profiling runs, they predict the mapping for a new unseen program across multiple input data sets. We evaluate our approach by selecting parallelism mapping configurations for OpenMP programs on two representative but different multi-core platforms (the Intel Xeon and the Cell processors). Performance of our technique is stable across programs and architectures. On average, it delivers above 96% performance of the maximum available on both platforms. It achieve, on average, a 37% (up to 17.5 times) performance improvement over the OpenMP runtime default scheme on the Cell platform. Compared to two recent prediction models, our predictors achieve better performance with a significant lower profiling cost.",https://dl.acm.org/doi/10.1145/1594835.1504189,True,,219.0,"['multilayer-perceptron', 'support-vector-machine']",,['novel-use'],"['scheduling', 'workload-prediction']",['multiprocessor'],ACM SIGPLAN Notices,True,['resource-provisioning'],,,,,,10.1145/1594835.1504189,,,
716,A machine learning-based approach for thread mapping on transactional memory applications,"['Márcio Castro', 'Luís Fabrício Wanderley Góes', 'Christiane Pousa Ribeiro', 'Murray Cole', 'Marcelo Cintra', 'Jean-François Méhaut']",2011,"Thread mapping has been extensively used as a technique to efficiently exploit memory hierarchy on modern chip-multiprocessors. It places threads on cores in order to amortize memory latency and/or to reduce memory contention. However, efficient thread mapping relies upon matching application behavior with system characteristics. Particularly, Software Transactional Memory (STM) applications introduce another dimension due to its runtime system support. Existing STM systems implement several conflict detection and resolution mechanisms, which leads STM applications to behave differently for each combination of these mechanisms. In this paper we propose a machine learning-based approach to automatically infer a suitable thread mapping strategy for transactional memory applications. First, we profile several STM applications from the STAMP benchmark suite considering application, STM system and platform features to build a set of input instances. Then, such data feeds a machine learning algorithm, which produces a decision tree able to predict the most suitable thread mapping strategy for new unobserved instances. Results show that our approach improves performance up to 18.46% compared to the worst case and up to 6.37% over the Linux default thread mapping strategy.",https://ieeexplore.ieee.org/abstract/document/6152736,True,,47.0,['decision-tree'],,['novel-use'],['scheduling'],,2011 18th International Conference on High Performance Computing,True,['resource-provisioning'],,,,,,,,,
717,Genetic Algorithm Based Schedulers for Grid Computing Systems,"['Xhafa Xhafa, Fatos', ""Carretero Casado, Javier Sebasti{\\'a}n"", 'Abraham, Ajith']",2007,"In this paper we present Genetic Algorithms (GAs) based schedulers for efficiently allocating jobs to resources in a Grid system. Scheduling is a key problem in emergent computational systems, such as Grid and P2P, in order to benefit from the large computing capacity of such systems. We present an extensive study on the usefulness of GAs for designing efficient Grid schedulers when makespan and flowtime are minimized. Two encoding schemes have been considered and most of GA operators for each of them are implemented and …",https://www.researchgate.net/publication/228624815_Genetic_Algorithm_Based_Schedulers_for_Grid_Computing_Systems,True,,260.0,['genetic-programming'],,['novel-use'],['scheduling'],,,True,['resource-provisioning'],,,,,,,,,
718,Observations on using genetic algorithms for dynamic load-balancing,"['A.Y. Zomaya', 'Yee-Hwei Teh']",2001,"Load-balancing problems arise in many applications, but, most importantly, they play a special role in the operation of parallel and distributed computing systems. Load-balancing deals with partitioning a program into smaller tasks that can be executed concurrently and mapping each of these tasks to a computational resource such as a processor (e.g., in a multiprocessor system) or a computer (e.g., in a computer network). By developing strategies that can map these tasks to processors in a way that balances out the load, the total processing time will be reduced with improved processor utilization. Most of the research on load-balancing focused on static scenarios that, in most of the cases, employ heuristic methods. However, genetic algorithms have gained immense popularity over the last few years as a robust and easily adaptable search technique. The work proposed here investigates how a genetic algorithm can be employed to solve the dynamic load-balancing problem. A dynamic load-balancing algorithm is developed whereby optimal or near-optimal task allocations can ""evolve"" during the operation of the parallel computing system. The algorithm considers other load-balancing issues such as threshold policies, information exchange criteria, and interprocessor communication. The effects of these and other issues on the success of the genetic-based load-balancing algorithm as compared with the first-fit heuristic are outlined.",https://ieeexplore.ieee.org/document/954620,True,,478.0,['genetic-programming'],,"['new-method', 'comparison']",['resource-consolidation'],,,True,['resource-provisioning'],,,,,,,,,
719,"Smart, adaptive mapping of parallelism in the presence of external workload","['Murali Krishna Emani', 'Zheng Wang', ""Michael F. P. O'Boyle""]",2013,"Given the wide scale adoption of multi-cores in main stream computing, parallel programs rarely execute in isolation and have to share the platform with other applications that compete for resources. If the external workload is not considered when mapping a program, it leads to a significant drop in performance. This paper describes an automatic approach that combines compile-time knowledge of the program with dynamic runtime workload information to determine the best adaptive mapping of programs to available resources. This approach delivers increased performance for the target application without penalizing the existing workload. This approach is evaluated on NAS and SpecOMP parallel bench-mark programs across a wide range of workload scenarios. On average, our approach achieves performance gain of 1.5× over a state-of-art scheme on a 12 core machine.",https://ieeexplore.ieee.org/document/6495010,True,,22.0,,,,,,,True,['resource-provisioning'],,,,,,,,,
720,Run-time prediction of parallel applications on shared environments,"['Byoung-Dai Lee', 'Schopf']",2003,"Application run-time is a fundamental component in application and job scheduling. However, accurate predictions of run times are difficult to achieve for parallel applications running in shared environments where resource capacities can change dynamically over time. In this paper, we propose a run-time prediction technique for parallel applications that uses regression methods and filtering techniques to derive the application execution time without using standard performance models. The experimental results show that our use of regression models delivers tolerable prediction accuracy and that we can improve the accuracy dramatically by using appropriate filters.",https://ieeexplore.ieee.org/document/1253355,True,,5.0,,,,,,2003 Proceedings IEEE International Conference on Cluster Computing,True,['resource-provisioning'],,,,,,,,,
721,A Novel Adaptive Support Vector Machine based Task Scheduling,"['Park, Yong-won', 'Baskiyar, S', 'Casey, K']",2010,"In this paper, we propose a novel instance based Support Vector based machine learning approach to task scheduling in heterogeneous computational grids. The system is composed of two components: a Support Vector Machine (SVM) scheduler and a dynamic learner. The scheduler is a multiclass SVM which maps any task to a machine based on the current load on the machines and the queue of tasks waiting to be dispatched. To support dynamic adaptation of the scheduler, we have designed a dynamic learning system which incorporates new knowledge into the scheduler. It allows the scheduler to update the training set and adapt to changing conditions. As newer and better scheduling strategies are discovered, they are incorporated into the training set and thereby used for scheduling tasks in the future. We demonstrate that our SVM scheduler has comparable performance to conventional task scheduling heuristics. Our approach allows for dynamic and adaptive behavior of the system. For grid based systems, where the pattern of tasks is unpredictable, such a technique may be superior to specific heuristic methods and fulfill a much needed instance based learner approach.",https://www.researchgate.net/publication/228574204_A_Novel_Adaptive_Support_Vector_Machine_based_Task_Scheduling,True,,4.0,['support-vector-machine'],,['novel-use'],['scheduling'],,"Proceedings the 9th International Conference on Parallel and Distributed Computing and Networks, Austria",True,['resource-provisioning'],,,,,,,,,
722,Dynamic Thread Mapping Based on Machine Learning for Transactional Memory Applications,"[""Castro, M{\\'a}rcio"", ""G{\\'o}es, Lu{\\'\\i}s Fabr{\\'\\i}cio Wanderley"", 'Fernandes, Luiz Gustavo', ""M{\\'e}haut, Jean-Fran{\\c{c}}ois""]",2012,"Thread mapping is an appealing approach to efficiently exploit the potential of modern chip-multiprocessors. However, efficient thread mapping relies upon matching the behavior of an application with system characteristics. In particular, Software Transactional Memory (STM) introduces another dimension due to its runtime system support. In this work, we propose a dynamic thread mapping approach to automatically infer a suitable thread mapping strategy for transactional memory applications composed of multiple execution phases with …",https://link.springer.com/chapter/10.1007/978-3-642-32820-6_47,True,,13.0,['decision-tree'],,['new-method'],['scheduling'],,European Conference on Parallel Processing,True,['resource-provisioning'],,,,,,,,,
723,Using a multi-agent system and artificial intelligence for monitoring and improving the cloud performance and security,"['Grzonka, Daniel', 'Jakobik, Agnieszka', 'Ko{\\l}odziej, Joanna', 'Pllana, Sabri']",2018,"Cloud Computing is one of the most intensively developed solutions for large-scale distributed processing. Effective use of such environments, management of their high complexity and ensuring appropriate levels of Quality of Service (QoS) require advanced monitoring systems. Such monitoring systems have to support the scalability, adaptability and reliability of Cloud. Most of existing monitoring systems do not incorporate any Artificial Intelligence (AI) algorithms for supporting the change inside the task stream or environment itself. They focus only on monitoring or enabling the control of the system as a part of a separated service. An effective monitoring system for the Cloud environment should gather information about all stages of tasks processing and should actively control the monitored environment.",https://www.sciencedirect.com/science/article/abs/pii/S0167739X17310531?via%3Dihub,True,,33.0,,,,['scheduling'],,,True,['resource-provisioning'],,,,,,,,,
724,A workload-aware mapping approach for data-parallel programs,"['Dominik Grewe', 'Zheng Wang', ""Michael F. P. O'Boyle""]",2011,"Much compiler-orientated work in the area of mapping parallel programs to parallel architectures has ignored the issue of external workload. Given that the majority of platforms will not be dedicated to just one task at a time, the impact of other jobs needs to be addressed. As mapping is highly dependent on the underlying machine, a technique that is easily portable across platforms is also desirable. In this paper we develop an approach for predicting the optimal number of threads for a given data-parallel application in the presence of external workload. We achieve 93.7% of the maximum speedup available which gives an average speedup of 1.66 on 4 cores, a factor 1.24 times better than the OpenMP compiler's default policy. We also develop an alternative cooperative model that minimizes the impact on external workload while still giving an improved average speedup. Finally, we evaluate our approach on a separate 8-core machine giving an average 1.33 times speedup over the default policy showing the portability of our approach.",https://dl.acm.org/doi/abs/10.1145/1944862.1944881,True,,34.0,,,,,,HiPEAC '11: Proceedings of the 6th International Conference on High Performance and Embedded Architectures and Compilers,True,['resource-provisioning'],,,,,,10.1145/1944862.1944881,,,
725,Runtime Empirical Selection of Loop Schedulers on Hyperthreaded SMPs,"['Yun Zhang', 'Michael Voss']",2005,"Hyperthreaded (HT) and simultaneous multithreaded (SMT) processors are now available in commodity workstations and servers. This technology is designed to increase throughput by executing multiple concurrent threads on a single physical processor. These multiple threads share the processor's functional units and on-chip memory hierarchy in an attempt to make better use of idle resources. Most OpenMP applications have been written assuming an Symmetric Multiprocessor (SMP), not an SMT, model. Threads executing on the same physical processor have interactions on data locality and resource sharing that do not occur on traditional SMPs. This work focuses on tuning the behavior of OpenMP applications executing on SMPs with SMT processors. We propose two adaptive loop schedulers that determine effective hierarchical schedulers for individual parallel loops. We compare the performance of our two proposed schedulers against several standard schedulers and the per-region adaptive scheduler proposed by Zhang et al. using the SPEC and NAS OpenMP benchmark suites. We show that both of our proposed schedulers outperform all other schedulers on average, and increase speedup on average by over 25% when all thread contexts are used.",https://dl.acm.org/doi/10.1109/IPDPS.2005.386,True,,36.0,,,,,,IPDPS '05: Proceedings of the 19th IEEE International Parallel and Distributed Processing Symposium (IPDPS'05) - Papers - Volume 01,True,['resource-provisioning'],,,,,,10.1109/IPDPS.2005.386,,,
726,Framework for Task Scheduling in Heterogeneous Distributed Computing Using Genetic Algorithms,"['Page, Andrew J', 'Naughton, Thomas J']",2005,"An algorithm has been developed to dynamically schedule heterogeneous tasks on heterogeneous processors in a distributed system. The scheduler operates in an environment with dynamically changing resources and adapts to variable system resources. It operates in a batch fashion and utilises a genetic algorithm to minimise the total execution time. We have compared our scheduler to six other schedulers, three batch-mode and three immediate-mode schedulers. Experiments show that the algorithm outperforms each of the …",https://link.springer.com/article/10.1007%2Fs10462-005-9002-x,True,,104.0,['genetic-programming'],,['novel-use'],['scheduling'],,,True,['resource-provisioning'],,,,,,,,,
727,MRONLINE: MapReduce online performance tuning,"['Min Li', 'Liangzhao Zeng', 'Shicong Meng', 'Jian Tan', 'Li Zhang', 'Ali R. Butt', 'Nicholas Fuller']",2014,"MapReduce job parameter tuning is a daunting and time consuming task. The parameter configuration space is huge; there are more than 70 parameters that impact job performance. It is also difficult for users to determine suitable values for the parameters without first having a good understanding of the MapReduce application characteristics. Thus, it is a challenge to systematically explore the parameter space and select a near-optimal configuration. Extant offline tuning approaches are slow and inefficient as they entail multiple test runs and significant human effort. To this end, we propose an online performance tuning system, MRONLINE, that monitors a job's execution, tunes associated performance-tuning parameters based on collected statistics, and provides fine-grained control over parameter configuration. MRONLINE allows each task to have a different configuration, instead of having to use the same configuration for all tasks. Moreover, we design a gray-box based smart hill climbing algorithm that can efficiently converge to a near-optimal configuration with high probability. To improve the search quality and increase convergence speed, we also incorporate a set of MapReduce-specific tuning rules in MRONLINE. Our results using a real implementation on a representative 19-node cluster show that dynamic performance tuning can effectively improve MapReduce application performance by up to 30% compared to the default configuration used in YARN.",https://dl.acm.org/doi/10.1145/2600212.2600229,True,,50.0,,,,['configuration'],,HPDC '14: Proceedings of the 23rd international symposium on High-performance parallel and distributed computing,True,['resource-provisioning'],,,,,,10.1145/2600212.2600229,,,
728,Dynamic Task Scheduling with Load Balancing using Genetic Algorithm,"['Chouhan Kumar Rath', 'Prasanti Biswal', 'Shashank Sekhar Suar']",2018,"Parallel computing is better suitable for modeling, simulating complex problems. Simultaneous use of multiple tasks to compute a particular problem defines Parallel Computing. There are various techniques available to build a powerful parallel computing processor. An efficient task scheduling problem improves the performance of the multiprocessing system. So, distribution of tasks in a multiprocessing environment is the most critical issue today. This paper presents an approach to dynamic task scheduling with load balancing. Load distribution is a major issue for parallel processing now a days. In order to minimize the average make-span and solve the load balancing problem, the Genetic algorithm is used. New kind of fitness function is evaluated using standard deviation. The system is modeled by taking a standard task graph with communication cost and computation cost. Task graph consists of nodes (tasks) allocated to the processors with dependency. The experimental results depict the efficiency of the proposed method.",https://ieeexplore.ieee.org/document/8724218,True,,1.0,['genetic-programming'],,,['scheduling'],,2018 International Conference on Information Technology (ICIT),True,['resource-provisioning'],,,,,,,,,
729,Dynamic task scheduling with load balancing using parallel orthogonal particle swarm optimisation,"['S. N. Sivanandam', 'P. Visalakshi']",2009,"This paper presents a novel approach for dynamic task scheduling using particle swarm optimisation. Particle swarm optimisation (PSO) is a population-based meta-heuristic method which can be used to solve np-hard problems. The algorithm has been developed to dynamically schedule heterogeneous tasks on to heterogeneous processors in a distributed setup. Load balancing which is a major issue in task scheduling is also considered. The nature of the tasks are independent and non pre-emptive. Different approaches using PSO has been tried namely PSO with fixed inertia, PSO with variable inertia, PSO with elitism, MPSO, parallel PSO, hybrid PSO, orthogonal PSO and parallel orthogonal PSO. The performance of PSO and its variants is also compared with the genetic algorithm concept. The objective of the algorithms is to minimise the make-span of the entire schedule. Benchmark problems have been taken and validated. The result depicts that the dynamic task scheduling implemented using parallel orthogonal particle swarm optimisation technique is cost-effective in nature when compared to the other algorithms tested.",https://dl.acm.org/doi/10.1504/IJBIC.2009.024726,True,,46.0,['particle-swarm'],,,"['resource-consolidation', 'scheduling']",,International Journal of Bio-Inspired Computation,True,['resource-provisioning'],,,,,,10.1504/IJBIC.2009.024726,,,
730,A dynamic scheduling framework for emerging heterogeneous systems,"['Vignesh T. Ravi', 'Gagan Agrawal']",2011,"A trend that has materialized, and has given rise to much attention, is of the increasingly heterogeneous computing platforms. Recently, it has become very common for a desktop or a notebook computer to be equipped with both a multi-core CPU and a GPU. Application development for exploiting the aggregate computing power of such an environment is a major challenge today. Particularly, we need dynamic work distribution schemes that are adaptable to different computation and communication patterns in applications, and to various heterogeneous configurations. This paper describes a general dynamic scheduling framework for mapping applications with different communication patterns to heterogeneous architectures. We first make key observations about the architectural tradeoffs among heterogeneous resources and the communication pattern of an application, and then infer constraints for the dynamic scheduler. We then present a novel cost model for choosing the optimal chunk size in a heterogeneous configuration. Finally, based on general framework and cost model we provide optimized work distribution schemes to further improve the performance.",https://ieeexplore.ieee.org/document/6152724,True,,42.0,,,,['scheduling'],,2011 18th International Conference on High Performance Computing,True,['resource-provisioning'],,,,,,,,,
731,Harmony: an execution model and runtime for heterogeneous many core systems,"['Gregory F. Diamos', 'Sudhakar Yalamanchili']",2008,"The emergence of heterogeneous many core architectures presents a unique opportunity for delivering order of magnitude performance increases to high performance applications by matching certain classes of algorithms to specifically tailored architectures. Their ubiquitous adoption, however, has been limited by a lack of programming models and management frameworks designed to reduce the high degree of complexity of software development intrinsic to heterogeneous architectures. This paper proposes Harmony, a runtime supported programming and execution model that provides: (1) semantics for simplifying parallelism management, (2) dynamic scheduling of compute intensive kernels to heterogeneous processor resources, and (3) online monitoring driven performance optimization for heterogeneous many core systems. We are particulably concerned with simplifying development and ensuring binary portability and scalability across system configurations and sizes. Initial results from ongoing development demonstrate the binary compatibility with variable number of cores, as well as dynamic adaptation of schedules to data sets. We present preliminary results of key features for some benchmark applications.",https://dl.acm.org/doi/abs/10.1145/1383422.1383447,True,,107.0,,,,"['scheduling', 'anomaly-detection']",['hpc'],HPDC '08: Proceedings of the 17th international symposium on High performance distributed computing,True,['resource-provisioning'],,,,,,10.1145/1383422.1383447,,,
732,PEPPHER: Efficient and Productive Usage of Hybrid Computing Systems,"['Siegfried Benkner', 'Sabri Pllana', 'Jesper Larsson Traff', 'Philippas Tsigas', 'Uwe Dolinsky', 'Cedric Augonnet', 'Beverly Bachmayer', 'Christoph Kessler', 'David Moloney', 'Vitaly Osipov']",2011,"PEPPHER, a three-year European FP7 project, addresses efficient utilization of hybrid (heterogeneous) computer systems consisting of multicore CPUs with GPU-type accelerators. This article outlines the PEPPHER performance-aware component model, performance prediction means, runtime system, and other aspects of the project. A larger example demonstrates performance portability with the PEPPHER approach across hybrid systems with one to four GPUs.",https://ieeexplore.ieee.org/document/5959147,True,,56.0,,,,['resource-consolidation'],,,True,['resource-provisioning'],,,,,,,,,
733,Adaptive Off-Line Tuning for Optimized Composition of Components for Heterogeneous Many-Core Systems,"['Li, Lu', 'Dastgeer, Usman', 'Kessler, Christoph']",2012,"In recent years heterogeneous multi-core systems have been given much attention. However, performance optimization on these platforms remains a big challenge. Optimizations performed by compilers are often limited due to lack of dynamic information and run time environment, which makes applications often not performance portable. One current approach is to provide multiple implementations for the same interface that could be used interchangeably depending on the call context, and expose the composition choices to a compiler, deployment-time composition tool and/or run-time system. Using off-line machine-learning techniques allows to improve the precision and reduce the run-time overhead of run-time composition and leads to an improvement of performance portability. In this work we extend the run-time composition mechanism in the PEPPHER composition tool by off-line composition and present an adaptive machine learning algorithm for generating compact and efficient dispatch data structures with low training time. As dispatch data structure we propose an adaptive decision tree structure, which implies an adaptive training algorithm that allows to control the trade-off between training time, dispatch precision and run-time dispatch overhead. We have evaluated our optimization strategy with simple kernels (matrix-multiplication and sorting) as well as applications from RODINIA benchmark on two GPU-based heterogeneous systems. On average, the precision for composition choices reaches 83.6 percent with approximately 34 minutes off-line training time.",https://link.springer.com/chapter/10.1007/978-3-642-38718-0_32,True,,20.0,,,,,,,True,['resource-provisioning'],,,,,,,,,
734,Optimized composition of performance-aware parallel components,"['C. Kessler', 'W. Löwe']",2012,"We describe the principles of a novel framework for performance-aware composition of sequential and explicitly parallel software components with implementation variants. Automatic composition results in a table-driven implementation that, for each parallel call of a performance-aware component, looks up the expected best implementation variant, processor allocation and schedule given the current problem, and processor group sizes. The dispatch tables are computed off-line at component deployment time by an interleaved dynamic programming algorithm from time-prediction meta-code provided by the component supplier. Copyright © 2011 John Wiley & Sons, Ltd.",https://dl.acm.org/doi/10.1002/cpe.1844,True,,44.0,,,,['service-composition'],,Concurrency and Computation: Practice & Experience,True,['resource-provisioning'],,,,,,10.1002/cpe.1844,,,
735,Runtime Prediction of Service Level Agreement Violations for Composite Services,"['Leitner, Philipp', 'Wetzstein, Branimir', 'Rosenberg, Florian', 'Michlmayr, Anton', 'Dustdar, Schahram', 'Leymann, Frank']",2009,"SLAs are contractually binding agreements between service providers and consumers, mandating concrete numerical target values which the service needs to achieve. For service providers, it is essential to prevent SLA violations as much as possible to enhance customer satisfaction and avoid penalty payments. Therefore, it is desirable for providers to predict possible violations before they happen, while it is still possible to set counteractive measures. We propose an approach for predicting SLA violations at runtime, which uses …",https://link.springer.com/chapter/10.1007/978-3-642-16132-2_17,True,,125.0,,['events'],['new-method'],"['failure-prediction', 'workload-prediction', 'system-failure-prediction']",,Service-oriented computing. ICSOC/ServiceWave 2009 workshops,True,"['resource-provisioning', 'failure-management']",,,,,,,,,
736,Combining QoS-based service selection with performance prediction,"['Zhengdong Gao', 'Gengfeng Wu']",2005,"Currently, much efforts have being focused on dynamic, personalized QoS-based service selection, however, current QoS models are generally composed of static QoS parameters and haven't taken the dynamic nature of service performance into consideration. In our framework, we extend existing QoS model by adding new attributes that reflect performance of services and rely on ANN to provide client dynamic, on demand service performance prediction. Through this way, a client may be more capable of finding the best service based both on his/her preferences and on the service performance estimation",https://ieeexplore.ieee.org/document/1552957,True,,56.0,,,,['workload-prediction'],,IEEE International Conference on e-Business Engineering (ICEBE'05),True,['resource-provisioning'],,,,,,,,,
737,QoS-Based Service Selection and Ranking with Trust and Reputation Management,"['Vu, Le-Hung', 'Hauswirth, Manfred', 'Aberer, Karl']",2005,"QoS-based service selection mechanisms will play an essential role in service-oriented architectures, as e-Business applications want to use services that most accurately meet their requirements. Standard approaches in this field typically are based on the prediction of services' performance from the quality advertised by providers as well as from feedback of users on the actual levels of QoS delivered to them. The key issue in this setting is to detect and deal with false ratings by dishonest providers and users, which has only received …",https://link.springer.com/chapter/10.1007/11575771_30,True,,366.0,,,['new-method'],['service-composition'],,"OTM Confederated International Conferences"" On the Move to Meaningful Internet Systems""",True,['resource-provisioning'],,,,,,,,,
738,Event-Driven Quality of Service Prediction,"['Zeng, Liangzhao', 'Lingenfelder, Christoph', 'Lei, Hui', 'Chang, Henry']",2008,"Quality of Service Management (QoSM) is a new task in IT-enabled enterprises that supports monitoring, collecting and predicting QoS data. QoSM solutions must be able to efficiently process runtime events, compute and pre dict QoS metrics, and provide real-time visibility and prediction of key perform ance indicators (KPI). Currently, most QoSM systems focus on moni tor ing of QoS constraints, ie, they report what has been happened. In a way, this provides the awareness of past developments and sets the basis for decisions …",https://link.springer.com/chapter/10.1007/978-3-540-89652-4_14,True,,62.0,,,,['workload-prediction'],,International Conference on Service-Oriented Computing,True,['resource-provisioning'],,,,,,,,,
739,Self-Learning Prediction System for Optimisation of Workload Management in a Mainframe Operating System,"['Bensch, Michael', 'Brugger, Dominik', 'Rosenstiel, Wolfgang', 'Bogdan, Martin', 'Spruth, Wilhelm G', 'Baeuerle, Peter']",2007,"We present a framework for extraction and prediction of online workload data from a workload manager of a mainframe operating system. To boost overall system performance, the prediction will be incorporated into the workload manager to take preventive action before a bottleneck develops. Model and feature selection automatically create a prediction model based on given training data, thereby keeping the system flexible. We tailor data extraction, preprocessing and training to this specific task, keeping in mind the …",https://www.researchgate.net/publication/220708659_Self-Learning_Prediction_System_for_Optimisation_of_Workload_Management_in_a_Mainframe_Operating_System,True,,6.0,,,,,,ICEIS (2),True,['resource-provisioning'],,,,,,,,,
740,An Approach to Forecasting QoS Attributes of Web Services Based on ARIMA and GARCH Models,"['Ayman Amin', 'Alan Colman', 'Lars Grunske']",2012,"Availability of several web services having a similar functionality has led to using quality of service (QoS) attributes to support services selection and management. To improve these operations and be performed proactively, time series ARIMA models have been used to forecast the future QoS values. However, the problem is that in this extremely dynamic context the observed QoS measures are characterized by a high volatility and time-varying variation to the extent that existing ARIMA models cannot guarantee accurate QoS forecasting where these models are based on a homogeneity (constant variation over time) assumption, which can introduce critical problems such as proactively selecting a wrong service and triggering unrequired adaptations and thus leading to follow-up failures and increased costs. To address this limitation, we propose a forecasting approach that integrates ARIMA and GARCH models to be able to capture the QoS attributes' volatility and provide accurate forecasts. Using QoS datasets of real-world web services we evaluate the accuracy and performance aspects of the proposed approach. Results show that the proposed approach outperforms the popular existing ARIMA models and improves the forecasting accuracy of QoS measures and violations by on average 28.7% and 15.3% respectively.",https://ieeexplore.ieee.org/abstract/document/6257792?section=abstract,True,,101.0,,,,['workload-prediction'],,,True,['resource-provisioning'],,,,,,,,,
741,Revenue Maximization Using Adaptive Resource Provisioning in Cloud Computing Environments,"['Guofu Feng', 'Saurabh Garg', 'Rajkumar Buyya', 'Wenzhong Li']",2012,"Compared with the traditional computing models such as grid computing and cluster computing, a key advantage of Cloud computing is that it provides a practical business model for customers to use remote resources. However, it is challenging for Cloud providers to allocate the pooled computing resources dynamically among the differentiated customers so as to maximize their revenue. It is not an easy task to transform the customer-oriented service metrics into operating level metrics, and control the Cloud resources adaptively based on Service Level Agreement (SLA). This paper addresses the problem of maximizing the provider's revenue through SLA-based dynamic resource allocation as SLA plays a vital role in Cloud computing to bridge service providers and customers. We formalize the resource allocation problem using Queuing Theory and propose optimal solutions for the problem considering various Quality of Service (QoS) parameters such as pricing mechanisms, arrival rates, service rates and available resources. The experimental results, both with the synthetic dataset and with traced dadataset, show that our algorithms outperform related work.",https://ieeexplore.ieee.org/document/6319170,True,,73.0,,,,['resource-consolidation'],,2012 ACM/IEEE 13th International Conference on Grid Computing,True,['resource-provisioning'],,,,,,,,,
742,"An Inter-cloud Outsourcing Model to Scale Performance, Availability and Security","['Emiliano Casalicchio', 'Luca Silvestri']",2012,"This paper presents a model of a horizontal cloud federation and studies the optimal resource selection and allocation policy a service provider should put in place to scale the performance, availability and security guarantees offered to customers. The proposed model considers: (i) resources located in different zones, characterized by different hourly costs and specific performance, availability, and security properties, (ii) service provider customers, dispersed in various zones(characterized by different latencies), and demanding services with different QoS levels defined in Service Level Agreements(SLAs). From this model we define an optimization problem allowing to determine the optimal distribution of the incoming load and the proper allocation of outsourced resources that satisfies the SLAs and that minimizes the outsourcing costs, thus allowing the maximization of the SP revenue. Experiments show how the optimal policy scales the service provider capabilities when the workload grows 10 times and more.",https://ieeexplore.ieee.org/document/6424940,True,,1.0,,,,,,2012 IEEE Fifth International Conference on Utility and Cloud Computing,True,['resource-provisioning'],,,,,,,,,
743,Dual time-scale distributed capacity allocation and load redirect algorithms for cloud systems,"['Danilo Ardagna', 'Sara Casolari', 'Michele Colajanni', 'Barbara Panicucci']",2012,"Resource management remains one of the main issues of cloud computing providers because system resources have to be continuously allocated to handle workload fluctuations while guaranteeing Service Level Agreements (SLA) to the end users. In this paper, we propose novel capacity allocation algorithms able to coordinate multiple distributed resource controllers operating in geographically distributed cloud sites. Capacity allocation solutions are integrated with a load redirection mechanism which, when necessary, distributes incoming requests among different sites. The overall goal is to minimize the costs of allocated resources in terms of virtual machines, while guaranteeing SLA constraints expressed as a threshold on the average response time. We propose a distributed solution which integrates workload prediction and distributed non-linear optimization techniques. Experiments show how the proposed solutions improve other heuristics proposed in literature without penalizing SLAs, and our results are close to the global optimum which can be obtained by an oracle with a perfect knowledge about the future offered load.",https://dl.acm.org/doi/10.1016/j.jpdc.2012.02.014,True,,95.0,,,,['resource-consolidation'],,Journal of Parallel and Distributed Computing,True,['resource-provisioning'],,,,,,10.1016/j.jpdc.2012.02.014,,,
744,Exploring decentralized dynamic scheduling for grids and clouds using the community-aware scheduling algorithm,"['Ye Huang', 'Nik Bessis', 'Peter Norrington', 'Pierre Kuonen', 'Beat Hirsbrunner']",2013,"Job scheduling strategies have been studied for decades in a variety of scenarios. Due to the new characteristics of the emerging computational systems, such as the grid and cloud, metascheduling turns out to be an important scheduling pattern because it is responsible for orchestrating resources managed by independent local schedulers and bridges the gap between participating nodes. Equally, to overcome issues such as bottleneck, single point failure, and impractical unique administrative management, which are normally led by conventional centralized or hierarchical schemes, the decentralized scheduling scheme is emerging as a promising approach because of its capability with regards to scalability and flexibility.In this work, we introduce a decentralized dynamic scheduling approach entitled the community-aware scheduling algorithm (CASA). The CASA is a two-phase scheduling solution comprised of a set of heuristic sub-algorithms to achieve optimized scheduling performance over the scope of overall grid or cloud, instead of individual participating nodes. The extensive experimental evaluation with a real grid workload trace dataset shows that, when compared to the centralized scheduling scheme with BestFit as the metascheduling policy, the use of CASA can lead to a 30%-61% better average job slowdown, and a 68%-86% shorter average job waiting time in a decentralized scheduling manner without requiring detailed real-time processing information from participating nodes. Highlights We introduce a decentralized scheduling algorithm without requiring detailed node information. Our algorithm is able to adapt to the changes in grids through time by rescheduling. Comparisons with the known BestFit algorithm within a centralized scheduling scheme are made. Our algorithm leads to a 30%-61% better average job slowdown. Our algorithm leads to a 68%-86% shorter average job waiting time.",https://dl.acm.org/doi/10.1016/j.future.2011.05.006,True,,89.0,,,,['scheduling'],,Future Generation Computer Systems,True,['resource-provisioning'],,,,,,10.1016/j.future.2011.05.006,,,
745,An adaptive resource management scheme in cloud computing,"['Chenn-Jung Huang', 'Chih-Tai Guan', 'Heng-Ming Chen', 'Yu-Wu Wang', 'Shun-Chih Chang', 'Ching-Yu Li', 'Chuan-Hsiang Weng']",2013,"There are various significant issues in resource allocation, such as maximum computing performance and green computing, which have attracted researchers' attention recently. Therefore, how to accomplish tasks with the lowest cost has become an important issue, especially considering the rate at which the resources on the Earth are being used. The goal of this research is to design a sub-optimal resource allocation system in a cloud computing environment. A prediction mechanism is realized by using support vector regressions (SVRs) to estimate the number of resource utilization according to the SLA of each process, and the resources are redistributed based on the current status of all virtual machines installed in physical machines. Notably, a resource dispatch mechanism using genetic algorithms (GAs) is proposed in this study to determine the reallocation of resources. The experimental results show that the proposed scheme achieves an effective configuration via reaching an agreement between the utilization of resources within physical machines monitored by a physical machine monitor and service level agreements (SLA) between virtual machines operators and a cloud services provider. In addition, our proposed mechanism can fully utilize hardware resources and maintain desirable performance in the cloud environment.",https://www.sciencedirect.com/science/article/abs/pii/S0952197612002655?via%3Dihub,True,,57.0,,,,['resource-consolidation'],,Engineering Applications of Artificial Intelligence,True,['resource-provisioning'],,,,,,10.1016/j.engappai.2012.10.004,,,
746,An Evolutionary Approach for SLA-based Cloud Resource Provisioning,"['Victor Ion Munteanu', 'Teodor-Florin Fortis', 'Viorel Negru']",2013,"As the number of existing cloud vendors rises, resource count and types are ever increasing leading to a need of cloud management solutions which facilitate easy cloud adoption. While providing several services, cloud management's primary role is resource provisioning. In order to meet application needs in terms of resources, cloud developers must carefully choose among the existing offers in order to deploy their applications. The research presented in this paper enables developers to automate the process of resource provisioning by specifying their preference towards resources and resource attributes based on which the system can propose solutions for their requirements.",https://ieeexplore.ieee.org/document/6531797,True,,4.0,,,,,,2013 IEEE 27th international conference on advanced information networking and applications (AINA),True,['resource-provisioning'],,,,,,,,,
747,A statistical based resource allocation scheme in cloud,"['Zhenzhong Zhang', 'Haiyan Wang', 'Limin Xiao', 'Li Ruan']",2011,"Recently, cloud computing has emerged as a new computing paradigm on the Internet. With the development of cloud computing, enterprise data centers shift towards a utility computing model where many critical business applications share a common pool of infrastructure resources offering capacity on demand. The virtual machine with the features of strong isolation and flexible is usually assigned as the basic unit. However, as the demand of each type of VM can fluctuate independently at run time, it becomes a challenging problem to allocate data center resources to each VM to balance the workload in the cloud. In this paper, we introduce an approach (Statistic based Load Balance, SLB) that makes use of the statistical prediction and available resource evaluation mechanism to make online resource allocation decisions. Unlike the methods that balance load based on SLA (Service Level Agreement) of VMs, SLB achieves load balancing by predicting the VM's resource demand. The approach includes two parts:(1) A data analysis of on-line historical performance for forecasting the resource demand of each VM, and (2) An algorithm for choosing a proper host in the resource pool to run the VM. Experiments show that SLB can perform load balance in time, and also perform more balanced use of different resources.",https://ieeexplore.ieee.org/document/6138531,True,,16.0,,,,,,2011 International Conference on Cloud and Service Computing,True,['resource-provisioning'],,,,,,,,,
748,URL: A unified reinforcement learning approach for autonomic cloud management,"['Cheng-Zhong Xu', 'Jia Rao', 'Xiangping Bu']",2012,"Cloud computing is emerging as an increasingly important service-oriented computing paradigm. Management is a key to providing accurate service availability and performance data, as well as enabling real-time provisioning that automatically provides the capacity needed to meet service demands. In this paper, we present a unified reinforcement learning approach, namely URL, to automate the configuration processes of virtualized machines and appliances running in the virtual machines. The approach lends itself to the application of real-time autoconfiguration of clouds. It also makes it possible to adapt the VM resource budget and appliance parameter settings to the cloud dynamics and the changing workload to provide service quality assurance. In particular, the approach has the flexibility to make a good trade-off between system-wide utilization objectives and appliance-specific SLA optimization goals. Experimental results on Xen VMs with various workloads demonstrate the effectiveness of the approach. It can drive the system into an optimal or near-optimal configuration setting in a few trial-and-error iterations.",https://dl.acm.org/doi/10.1016/j.jpdc.2011.10.003,True,,96.0,,,,"['configuration', 'resource-consolidation']",,Journal of Parallel and Distributed Computing,True,['resource-provisioning'],,,,,,10.1016/j.jpdc.2011.10.003,,,
749,An adaptive power management framework for autonomic resource configuration in cloud computing infrastructures,"['Ziming Zhang', 'Qiang Guan', 'Song Fu']",2012,"Power is becoming an increasingly important concern for large-scale cloud computing systems. Meanwhile, cloud service providers leverage virtualization technologies to facilitate service consolidation and enhance resource utilization. However, the introduction of virtualization makes the cloud infrastructure more complex, and thus challenges cloud power management. In a virtualized environment, resource needs to be configured at runtime at the cloud, server and virtual machine levels to achieve high power efficiency. In addition, cloud power management should guarantee high users' SLA (service level agreement) satisfaction. In this paper, we present an adaptive power management framework in the cloud to achieve autonomic resource configuration. We propose a software and lightweight approach to accurately estimate the power usage of virtual machines and cloud servers. It explores hypervisor-observable performance metrics to build the power usage model. To configure cloud resources, we consider both the system power usage and the SLA requirements, and leverage learning techniques to achieve autonomic resource allocation and optimal power efficiency. We implement a prototype of the proposed power management system and test it on a cloud testbed. Experimental results show the high accuracy (over 90%) of our power usage estimation mechanism and our resource configuration approach achieves the lowest energy usage among the compared four approaches.",https://ieeexplore.ieee.org/document/6407738,True,,6.0,,,,,,2012 IEEE 31st International Performance Computing and Communications Conference (IPCCC),True,['resource-provisioning'],,,,,,,,,
750,IdleCached: An Idle Resource Cached Dynamic Scheduling Algorithm in Cloud Computing,"['Hu Song', 'Jing Li', 'Xinchun Liu']",2012,"Cloud computing service provides elastic and infinite computing resources to meet users' QoS requirements such as deadline and service cost. However, some tasks may be delayed because of the long waiting time before their resources are available. So, how to supply dynamic resource scaling and task migration is an important question which determines users' Service-level agreement(SLA) whether can be achieved. In this work, we construct the Torque Cloud management system to support dynamic task scheduling in the Cloud through applying Torque distributed resource management software to Eucalyptus Cloud platform. And also an idle resource cached dynamic scheduling algorithm(Idle Cached) is proposed to dynamically adjust tasks, which will take full advantage of idle resources in the resource pool. The results prove that the algorithm can quickly meet tasks' resource demands, minimize users' service cost and system energy-consumption.",https://ieeexplore.ieee.org/document/6332105,True,,5.0,,,,,,2012 9th International Conference on Ubiquitous Intelligence and Computing and 9th International Conference on Autonomic and Trusted Computing,True,['resource-provisioning'],,,,,,,,,
751,Optimizing Resource allocation while handling SLA violations in Cloud Computing platforms,"['Lionel Eyraud-Dubois', 'Hubert Larchevêque']",2013,"In this paper, we study a resource allocation problem in the context of Cloud Computing, in which a set of Virtual Machines (VM) has to be allocated on a set of Physical Machines (PM). Each VM has a given demand (e.g. CPU demand), and each PM has a capacity. However, VMs only use a fraction of their demand. The aim is to exploit the difference between the demand of the VM and its actual resource usage, to achieve a higher utilization on the PMs. However, the resource consumption of the VMs might change over time (while staying under its original demand), implying sometimes expensive “SLA violations” when the demand of some VMs is not satisfied because of overloaded PMs. Thus, while optimizing the global resource utilization of the PMs, it is necessary to ensure that at any moment a VM's need evolves, a few number of migrations (moving a VM from PM to PM) is sufficient to find a new configuration in which all the VMs' consumptions are satisfied. We model this problem using a fully dynamic bin packing approach and we present an algorithm ensuring a global utilization of the resources of 66%. Moreover, each time a PM is overloaded, at most one migration is sufficient to fall back in a configuration with no overloaded PM, and at most 3 different PMs are concerned by required migrations that may occur to keep the global resource utilization correct. This allows the platform to be highly resilient to a great number of changes.",https://ieeexplore.ieee.org/document/6569802,True,,16.0,,,,,,2013 IEEE 27th International Symposium on Parallel and Distributed Processing,True,['resource-provisioning'],,,,,,,,,
752,Dynamic Power- and Failure-Aware Cloud Resources Allocation for Sets of Independent Tasks,"['Altino M. Sampaio', 'Jorge G. Barbosa']",2013,"Cloud computing is increasingly being adopted in different scenarios, like social networking, business applications, scientific experiments, etc. Relying in virtualization technology, the construction of these computing environments targets improvements in the infrastructure, such as power-efficiency and fulfillment of users' SLA specifications. The methodology usually applied is packing all the virtual machines on the proper physical servers. However, failure occurrences in these networked computing systems can induce substantial negative impact on system performance, deviating the system from ours initial objectives. In this work, we propose adapted algorithms to dynamically map virtual machines to physical hosts, in order to improve cloud infrastructure power-efficiency, with low impact on users' required performance. Our decision making algorithms leverage proactive fault-tolerance techniques to deal with systems failures, allied with virtual machine technology to share nodes resources in an accurately and controlled manner. The results indicate that our algorithms perform better targeting power-efficiency and SLA fulfillment, in face of cloud infrastructure failures.",https://ieeexplore.ieee.org/document/6529262,True,,12.0,,,,,,2013 IEEE International Conference on Cloud Engineering (IC2E),True,['resource-provisioning'],,,,,,,,,
753,Workflow scheduling for SaaS / PaaS cloud providers considering two SLA levels,"['Thiago A. L. Genez', 'Luiz F. Bittencourt', 'Edmundo R. M. Madeira']",2012,"Cloud computing is being used to avoid maintenance costs and upfront investment, while providing elasticity to the available computational power in a pay-per-use basis. Customers can make use of the cloud as a software (SaaS), platform (PaaS), or infrastructure (IaaS) provider. When one customer utilizes an environment provided by a SaaS cloud, she is unaware of any details about the computational infrastructure where her requests are being processed. Therefore, such infrastructure can be composed of computational resources from a datacenter owned by the SaaS or its resources can be leased from a cloud infrastructure provider. In this paper we present an integer linear program (ILP) formulation for the problem of scheduling SaaS customer's workflows into multiple IaaS providers where SLA exists at two levels. In addition, we present heuristics to solve the relaxed version of the presented ILP. Simulation results show that the proposed ILP is able to find low-cost solutions for short deadlines, while the proposed heuristics are effective when deadlines are larger.",https://ieeexplore.ieee.org/document/6212007?reload=true&arnumber=6212007&contentType=Conference%20Publications,True,,48.0,,,,['scheduling'],,2012 IEEE Network Operations and Management Symposium,True,['resource-provisioning'],,,,,,,,,
754,Integrating Resource Consumption and Allocation for Infrastructure Resources on-Demand,"['Ying Zhang', 'Gang Huang', 'Xuanzhe Liu', 'Hong Mei']",2010,"Infrastructure resources on-demand requires resource provision (e.g., CPU and memory) to be both sufficient and necessary, which is the most important issue and a challenge in Cloud Computing. Platform as a service (PaaS) encapsulates a layer of software that includes middleware, and even development environment, and provides them as a service for building and deploying cloud applications. In PaaS, the issue of on-demand infrastructure resource management becomes more challenging due to the thousands of cloud applications that share and compete for resources simultaneously. The fundamental solution is to integrate and coordinate the resource consumption and allocation management of a cloud application. The difficulties of such a solution in PaaS are essentially how to maximize the resource utilization of an application, and how to allocate resources to guarantee adequate resource provision for the system. In this paper, we propose an approach to managing infrastructure resources in PaaS by leveraging two adaptive control loops: the resource consumption optimization loop and the resource allocation loop. The optimization loop improves the resource utilization of a cloud application via management functions provided by the corresponding middleware layers of PaaS. The allocation loop provides or reclaims appropriate amounts of resources to/from the application system while guaranteeing its performance. The two loops are integrated to run consecutively and repeatedly to provide infrastructure resources on-demand by first trying to improve resource utilization, and then allocating more resources when necessary. We implement a framework, SmartRod, to investigate our approach. The experiment on SmartRod proves its effectiveness on infrastructure resource management.",https://ieeexplore.ieee.org/document/5558007,True,,64.0,,,,['resource-consolidation'],,2010 IEEE 3rd International Conference on Cloud Computing,True,['resource-provisioning'],,,,,,,,,
755,A reinforcement learning framework for utility-based scheduling in resource-constrained systems,"['Vengerov, David']",2009,"This paper presents a general methodology for online scheduling of parallel jobs onto multi-processor servers in a soft real-time environment, where the final utility of each job decreases with the job completion time. A solution approach is presented where each server uses Reinforcement Learning for tuning its own value function, which predicts the average future utility per time step obtained from completed jobs based on the dynamically observed state information. The server then selects jobs from its job queue, possibly preempting some …",https://dl.acm.org/doi/book/10.5555/1698177,True,,46.0,,,,['scheduling'],,,True,['resource-provisioning'],,,,,,10.5555/1698177,,,
756,Bayes predictive analysis of a fundamental software reliability model,['A. Csenki'],1990,"The concepts of Bayes prediction analysis are used to obtain predictive distributions of the next time to failure of software when its past failure behavior is known. The technique is applied to the Jelinski-Moranda software-reliability model, which in turn can show an improved predictive performance for some data sets even when compared with some more sophisticated software-reliability models. A Bayes software-reliability model is presented which can be applied to obtain the next time to failure PDF (probability distribution function) and CDF (cumulative distribution function) for all testing protocols. The number of initial faults and the per-fault failure rate are assumed to be s-independent and Poisson and gamma distributed respectively. For certain data sets, the technique yields better predictions than some alternative methods if the frequential likelihood and U-plot criteria are adopted.<
>",https://ieeexplore.ieee.org/document/55879,True,,13.0,,,,['failure-prediction'],,,True,['failure-management'],,,,,,,,,
757,A nonparametric nonstationary procedure for failure prediction,"['J.D. Pfefferman', 'B. Cernuschi-Frias']",2002,"The time between failures is a very useful measurement to analyze reliability models for time-dependent systems. In many cases, the failure-generation process is assumed to be stationary, even though the process changes its statistics as time elapses. This paper presents a new estimation procedure for the probabilities of failures; it is based on estimating time-between-failures. The main characteristics of this procedure are that no probability distribution function is assumed for the failure process, and that the failure process is not assumed to be stationary. The model classifies the failures in Q different types, and estimates the probability of each type of failure s-independently from the others. This method does not use histogram techniques to estimate the probabilities of occurrence of each failure-type; rather it estimates the probabilities directly from the values of the time-instants at which the failures occur. The method assumes quasistationarity only in the interval of time between the last 2 occurrences of the same failure-type. An inherent characteristic of this method is that it assigns different sizes for the time-windows used to estimate the probabilities of each failure-type. For the failure-types with low probability, the estimator uses wide windows, while for those with high probability the estimator uses narrow windows. As an example, the model is applied to software reliability data.",https://ieeexplore.ieee.org/document/1044341,True,,31.0,,,,['failure-prediction'],,,True,['failure-management'],,,,,,,,,
758,Dependability measurement and modeling of a multicomputer system,"['D. Tang', 'R.K. Iyer']",1993,"A measurement-based analysis of error data collected from a DEC VAXcluster multicomputer system is presented. Basic system dependability characteristics such as error/failure distributions and hazard rate are obtained for both the individual machine and the entire VAXcluster. Markov reward models are developed to analyze error/failure behavior and to evaluate performance loss due to errors/failures. Correlation analysis is then performed to quantify relationships of error/failures across machines and across time. It is found that shared resources constitute a major reliability bottleneck. It is shown that for measured system, the homogeneous Markov model, which assumes constant failure rates, overestimates the transient reward rate for the short-term operation, and underestimates it for the long-term operation. Correlation analysis shows that errors are highly correlated across machines and across time. The failure correlation coefficient is low. However, its effect on system unavailability is significant.<
>",https://ieeexplore.ieee.org/document/192214,True,,86.0,,,,['failure-prediction'],,,True,['failure-management'],,,,,,,,,
759,FluxRank: A Widely-Deployable Framework to Automatically Localizing Root Cause Machines for Software Service Failure Mitigation,"['Ping Liu', 'Yu Chen', 'Xiaohui Nie', 'Jing Zhu', 'Shenglin Zhang', 'Kaixin Sui', 'Ming Zhang', 'Dan Pei']",2019,"The failures of software service directly affect user experiences and service revenue. Thus operators monitor both service-level KPIs (e.g., response time) and machine-level KPIs (e.g., CPU usage) on each machine underlying the service. When a service fails, the operators must localize the root cause machines, and mitigate the failure as quickly as possible. Existing approaches have limited application due to the difficulty to obtain the required additional measurement data. As a result, failure localization is largely manual and very time-consuming. This paper presents FluxRank, a widely-deployable framework that can automatically and accurately localize the root cause machines, so that some actions can be triggered to mitigate the service failure. Our evaluation using historical cases from five real services (with tens of thousands of machines) of a top search company shows that the root cause machines are ranked top 1 (top 3) for 55 (66) cases out of 70 cases. Comparing to existing approaches, FluxRank cuts the localization time by more than 80% on average. FluxRank has been deployed online at one Internet service and six banking services for three months, and correctly localized the root cause machines as the top 1 for 55 cases out of 59 cases.",https://ieeexplore.ieee.org/document/8987478,True,,0.0,,,,['failure-prediction'],,2019 IEEE 30th International Symposium on Software Reliability Engineering (ISSRE),True,['failure-management'],,,,,,,,,
760,Generic and Robust Localization of Multi-dimensional Root Causes,"['Zeyan Li', 'Chengyang Luo', 'Yiwei Zhao', 'Yongqian Sun', 'Kaixin Sui', 'Xiping Wang', 'Dapeng Liu', 'Xing Jin', 'Qi Wang', 'Dan Pei']",2019,"Operators of online software services periodically collect various measures with many attributes. When a measure becomes abnormal, indicating service problems such as reliability degrade, operators would like to rapidly and accurately localize the root cause attribute combinations within a huge multi-dimensional search space. Unfortunately, previous approaches are not generic or robust in that they all suffer from impractical root cause assumptions, handling only directly collected measures but not derived ones, handling only anomalies with signicant magnitudes but not those insignicant but important ones, requiring manual parameter ne-tuning, or being too slow. This paper proposes a generic and robust multi-dimensional root cause localization approach, Squeeze, that overcomes all above limitations, the first in the literature. Through our novel bottom-up then top-down searching strategy and the techniques based on our proposed generalized ripple effect and generalized potential score, Squeeze is able to reach a good trade off between search speed and accuracy in a generic and robust manner. Case studies in several banks and an Internet company show that Squeeze can localize root causes much more rapidly and accurately than the traditional manual analysis. Furthermore, our extensive experiments on semi-synthetic datasets show that the F1-score of Squeeze outperforms previous approaches by 0.4 on average, while its localization time is only about 10 seconds",https://ieeexplore.ieee.org/document/8987454,True,,0.0,,,,['root-cause-analysis'],,2019 IEEE 30th International Symposium on Software Reliability Engineering (ISSRE),True,['failure-management'],,,,,,,,,
761,LogAnomaly: Unsupervised Detection of Sequential and Quantitative Anomalies in Unstructured Logs,"['Meng, Weibin', 'Liu, Ying', 'Zhu, Yichen', 'Zhang, Shenglin', 'Pei, Dan', 'Liu, Yuqing', 'Chen, Yihao', 'Zhang, Ruizhi', 'Tao, Shimin', 'Sun, Pei', 'others']",2019,"Recording runtime status via logs is common for almost computer system, and detecting anomalies in logs is crucial for timely identifying malfunctions of systems. However, manually detecting anomalies for logs is time-consuming, error-prone, and infeasible. Existing automatic log anomaly detection approaches, using indexes rather than semantics of log templates, tend to cause false alarms. In this work, we propose LogAnomaly, a framework to model a log stream as a natural language sequence. Empowered by …",https://www.ijcai.org/Proceedings/2019/658,True,,7.0,['rnn'],['logs'],,['failure-detection'],,Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence,True,['failure-management'],True,,"['anomaly-detection', 'log-enhancement']",,,,,,
762,Dynamic TCP Initial Windows and Congestion Control Schemes through Reinforcement Learning,"['Xiaohui Nie', 'Youjian Zhao', 'Zhihan Li', 'Guo Chen', 'Kaixin Sui', 'Jiyang Zhang', 'Zijie Ye', 'Dan Pei']",2019,"Despite many years of improvements to it, TCP still suffers from an unsatisfactory performance. For services dominated by short flows (e.g., web search and e-commerce), TCP suffers from the flow startup problem and cannot fully utilize the available bandwidth in the modern Internet: TCP starts from a conservative and static initial window (IW, 2-4 or 10), while most of the web flows are too short to converge to the best sending rate before the session ends. For services dominated by long flows (e.g., video streaming and file downloading), the congestion control (CC) scheme manually and statically configured might not offer the best performance for the latest network conditions. To address these two challenges, we propose TCP-RL, which uses reinforcement learning (RL) techniques to dynamically configure IW and CC in order to improve the performance of TCP flow transmission. Basing on the latest network conditions observed at the server side of a web service, TCP-RL dynamically configures a suitable IW for short flows through group-based RL, and dynamically configures a suitable CC scheme for long flows through deep RL. Our extensive experiments show that for short flows, TCP-RL can reduce the average transmission time by about 23%; and for long flows, compared with the performance of 14 CC schemes, TCP-RL's performance ranks top 5 for about 85% of the 288 given static network conditions, whereas for about 90% of conditions, its performance drops by less than 12% compared with that of the best-performing CC schemes for the same network conditions.",https://ieeexplore.ieee.org/document/8668690,True,,3.0,"['multilayer-perceptron', 'reinforcement-learning']",,['new-method'],['configuration'],"['web', 'tcp', 'network']",,True,['resource-provisioning'],,,,,,,,,
763,Label-Less: A Semi-automatic Labelling Tool for KPI Anomalies,"['Nengwen Zhao', 'Jing Zhu', 'Rong Liu', 'Dapeng Liu', 'Ming Zhang', 'Dan Pei']",2019,"KPI (Key Performance Indicator) anomaly detection is critical for Internet-based services to ensure the quality and reliability. However, existing algorithms' performance in reality is far from satisfying due to the lack of sufficient KPI anomaly data to help train and evaluate these algorithms. In this paper, we argue that labeling overhead is the main hurdle to obtain such datasets. Thus, we novelly propose a semi-automatic labelling tool called Label-Less, which minimizes the labeling overhead in order to enable an ImageNet-like large-scale KPI anomaly dataset with high-quality ground truth. One novel technique in Label-Less is robust and rapid anomaly similarity search, which saves operators from scanning and checking the long KPIs back and forth for abnormal patterns or label consistency. In our evaluations using 30 real KPIs from a large Internet company, our anomaly similarity search achieves the best F-score of 0.95 on average, and a real-time per-KPI response time (less than 0.5 second). Overall, the feedback from deployment in practice shows that Label-Less can reduce operators' labeling overhead by more than 90%.",https://ieeexplore.ieee.org/document/8737429,True,,0.0,,['kpis'],['new-method'],['failure-detection'],,IEEE INFOCOM 2019-IEEE Conference on Computer Communications,True,['failure-management'],,,['anomaly-detection'],,,,,,
764,Mining Causality Graph For Automatic Web-based Service Diagnosis,"['Xiaohui Nie', 'Youjian Zhao', 'Kaixin Sui', 'Dan Pei', 'Yu Chen', 'Xianping Qu']",2016,"It is crucial for Internet company to provide highly reliable web-based services. The web-based services always have many components running in the large-scale infrastructure with complex interactions. As an indispensable part of high reliability, the diagnosis remains to be a thorny problem. With the growth of system scale and complexity, it becomes even more difficult. In this paper, we propose an automatic diagnosis system based on causality graph to help system operators find the root causes. The causality graph is mainly extracted from the historical data of the monitoring system, and the method consists of four steps. 1) It utilizes a data mining method to extract the initial causality graph. 2) Once a failure happens, it lists top-k suspects with a ranking algorithm based on the causality graph. 3) Then system operators check the suspects and label them either right or wrong. 4) A supervised learning algorithm takes the labels as the input to tune the causality graph, in order to improve the diagnosis accuracy on step 2 iteratively. This method requires neither knowledge about the design and implementation details of the web-based service, nor instrumenting the services' source code. Our controlled experiments show that the root causes can be ranked in top 3 with 100% accuracy after countable learning iterations.",https://ieeexplore.ieee.org/document/7820614,True,,7.0,"['graph-mining', 'rule-mining', 'random-forest']","['kpis', 'events']",,['root-cause-analysis'],,2016 IEEE 35th International Performance Computing and Communications Conference (IPCCC),True,['failure-management'],,,['root-cause-diagnosis'],,,,,,
765,BlueGene/L Failure Analysis and Prediction Models,"['Yinglung Liang', 'Yanyong Zhang', 'A. Sivasubramaniam', 'M. Jette', 'R. Sahoo']",2006,"The growing computational and storage needs of several scientific applications mandate the deployment of extreme-scale parallel machines, such as IBM's BlueGene/L which can accommodate as many as 128 K processors. One of the challenges when designing and deploying these systems in a production setting is the need to take failure occurrences, whether it be in the hardware or in the software, into account. Earlier work has shown that conventional runtime fault-tolerant techniques such as periodic checkpointing are not effective to the emerging systems. Instead, the ability to predict failure occurrences can help develop more effective checkpointing strategies. Failure prediction has long been regarded as a challenging research problem, mainly due to the lack of realistic failure data from actual production systems. In this study, we have collected RAS event logs from BlueGene/L over a period of more than 100 days. We have investigated the characteristics of fatal failure events, as well as the correlation between fatal events and non-fatal events. Based on the observations, we have developed three simple yet effective failure prediction methods, which can predict around 80% of the memory and network failures, and 47% of the application I/O failures",https://ieeexplore.ieee.org/document/1633531,True,,307.0,['correlation'],"['events', 'logs']",,['failure-prediction'],[],International Conference on Dependable Systems and Networks (DSN'06),True,['failure-management'],False,,['system-failure-prediction'],False,,,,,
766,Quantifying Temporal and Spatial Correlation of Failure Events for Proactive Management,"['Song Fu', 'Cheng-Zhong Xu']",2007,"Networked computing systems continue to grow in scale and in the complexity of their components and interactions. Component failures become norms instead of exceptions in these environments. Moreover, failure events exhibit strong correlations in time and space domain. In this paper, we develop a spherical covariance model with an adjustable timescale parameter to quantify the temporal correlation and a stochastic model to characterize spatial correlation. The models are further extended to take into account the information of application allocation to discover more correlations among failure instances. We cluster failure events based on their correlations and predict their future occurrences. Experimental results on a production coalition system, the Wayne State Grid, show the offline and online predictions by our predicting system can forecast 72.7% to 85.3% of the failure occurrences and capture failure correlations in cluster coalition environment.",https://ieeexplore.ieee.org/document/4365694,True,,68.0,,,,['failure-prediction'],,2007 26th IEEE International Symposium on Reliable Distributed Systems (SRDS 2007),True,['failure-management'],,,,,,,,,
767,An approach for estimation of software aging in a Web server,"['Lei Li', 'K. Vaidyanathan', 'K.S. Trivedi']",2002,"A number of recent studies have reported the phenomenon of ""software aging"", characterized by progressive performance degradation or a sudden hang/crash of a software system due to exhaustion of operating system resources, fragmentation and accumulation of errors. To counteract this phenomenon, a proactive technique called ""software rejuvenation"" has been proposed. This essentially involves stopping the running software, cleaning its internal state and then restarting it. Software rejuvenation, being preventive in nature, begs the question as to when to schedule it. Periodic rejuvenation, while straightforward to implement, may not yield the best results. A better approach is based on actual measurement of system resource usage and activity that detects and estimates resource exhaustion times. Estimating the resource exhaustion times makes it possible for software rejuvenation to be initiated or better planned so that the system availability is maximized in the face of time-varying workload and system behavior. We propose a methodology based on time series analysis to detect and estimate resource exhaustion times due to software aging in a Web server while subjecting it to an artificial workload. We first collect and log data on several system resource usage and activity parameters on a Web server. Time-series ARMA models are then constructed from the data to detect aging and estimate resource exhaustion times. The results are then compared with previous measurement-based models and found to be more efficient and computationally less intensive. These models can be used to develop proactive management techniques like software rejuvenation which are triggered by actual measurements.",https://ieeexplore.ieee.org/document/1166929,True,,222.0,['autoregression'],,,['failure-prevention'],['server'],Proceedings International Symposium on Empirical Software Engineering,True,['failure-management'],False,,['rejuvenation'],,,,,,
768,Deterministic Models of Software Aging and Optimal Rejuvenation Schedules,"['Artur Andrzejak', 'Luis Silva']",2007,"Automated modeling of software aging processes is a prerequisite for cost-effective usage of adaptive software rejuvenation as a self-healing technique. We consider the problem of such automated modeling in server-type applications whose performance degrades depending on the ""work"" done since last rejuvenation, for example the number of served requests. This type of performance degradation - caused mostly by resource depletion - is common, as we illustrate in a study of the popular Axis Soap server 1.3. In particular, we propose deterministic models for approximating the leading indicators of aging and an automated procedure for statistical testing of their correctness. We further demonstrate how to use these models for finding optimal rejuvenation schedules under utility functions. Our focus is on the important case that the utility function is the average of a performance metric (such as maximum service rate). We also consider optional SLA constraints under which the performance should never drop below a specified level. Our approach is verified by a study of the aging processes in the Axis Soap 1.3 server. The experiments show that the deterministic modeling technique is appropriate in this case, and that the optimization of rejuvenation schedules can greatly improve the average maximum service rate of an aging application.",https://ieeexplore.ieee.org/document/4258532,True,,55.0,,,,['failure-prediction'],,2007 10th IFIP/IEEE International Symposium on Integrated Network Management,True,['failure-management'],,,,,,,,,
769,Approaches for early fault detection in large scale engineering plants,"['Neville, Stephen William']",1998,"In general, it is difficult to automatically detect faults within large scale engineering plants early during their onset. This is due to a number of factors including the large number of components typically present in such plants and the complex interactions of these components, which are typically poorly understood. Traditionally, fault detection within these plants has been performed through the use of status monitoring systems employing limit checking fault detection. In this approach, upper and lower bounds are placed on what is …",http://hdl.handle.net/1828/8310,True,,8.0,,,,['failure-prediction'],,,True,['failure-management'],,,,,,,,,
770,A Best Practice Guide to Resources Forecasting for the Apache Webserver,"['Gunther A. Hoffmann', 'Kishor S. Trivedi', 'Miroslaw Malek']",2006,"Recently, measurement based studies of software systems prolifirated, reflecting an increasingly empirical focus on system availability, reliability, aging and fault tolerance. However, it is a non-trivial, error-prone, arduous, and time-consuming task even for experienced system administrators and statistical analysts to know what a reasonable set of steps should include to model and successfullypredict performance variables or system failures of a complex software system. Reported results are fragmented and focus on applying statistical regression techniques to captured numerical system data. In this paper, we propose a best practice guide for building empirical models based on our experience with forecasting Apache web sewer performance variables and forecasting call availability of a real world telecommunication system. To substantiate the presented guide and to demonstrate our approach step-by-step we model and predict the response time and the amount of free physical memory of an Apache web sewer system. Additionally, we present concrete results for a) variable selection where we cross benchmark three procedures, b) empirical model building where we cross benchmark four techniques and c) sensitivity analysis. This besr practice guide intends to assist in configuring modeling approaches systematically for best estimation andprediction results.",https://dl.acm.org/doi/10.1109/PRDC.2006.5,True,,36.0,,,,['failure-prediction'],,PRDC '06: Proceedings of the 12th Pacific Rim International Symposium on Dependable Computing,True,['failure-management'],,,,,,10.1109/PRDC.2006.5,,,
771,A Best Practice Guide to Resource Forecasting for Computing Systems,"['Guenther A. Hoffmann', 'Kishor S. Trivedi', 'Miroslaw Malek']",2007,"Recently, measurement-based studies of software systems have proliferated, reflecting an increasingly empirical focus on system availability, reliability, aging, and fault tolerance. However, it is a nontrivial, error-prone, arduous, and time-consuming task even for experienced system administrators, and statistical analysts to know what a reasonable set of steps should include to model, and successfully predict performance variables, or system failures of a complex software system. Reported results are fragmented, and focus on applying statistical regression techniques to monitored numerical system data. In this paper, we propose a best practice guide for building empirical models based on our experience with forecasting Apache web server performance variables, and forecasting call availability of a real-world telecommunication system. To substantiate the presented guide, and to demonstrate our approach in a step by step manner, we model, and predict the response time, and the amount of free physical memory of an Apache web server system, as well as the call availability of an industrial telecommunication system. Additionally, we present concrete results for a) variable selection where we cross benchmark three procedures, b) empirical model building where we cross benchmark four techniques, and c) sensitivity analysis. This best practice guide intends to assist in configuring modeling approaches systematically for best estimation, and prediction results.",https://ieeexplore.ieee.org/document/4378407,True,,78.0,,,,['workload-prediction'],,,True,['resource-provisioning'],,,,,,,,,
772,Predictive application-performance modeling in a computational grid environment,"['N.H. Kapadia', 'J.A.B. Fortes', 'C.E. Brodley']",1999,"This paper describes and evaluates the application of three local learning algorithms-nearest-neighbor, weighted-average, and locally-weighted polynomial regression-for the prediction of run-specific resource-usage on the basis of run-time input parameters supplied to tools. A two-level knowledge base allows the learning algorithms to track short-term fluctuations in the performances of computing systems, and the use of instance editing techniques improves the scalability of the performance-modeling system. The learning algorithms assist PUNCH, a network-computing system at Purdue University, in emulating an ideal user in terms of its resource management and usage policies.",https://ieeexplore.ieee.org/document/805281,True,,174.0,"['similarity-matching', 'autoregression']",,"['comparison', 'novel-use']",['workload-prediction'],,Proceedings. The Eighth International Symposium on High Performance Distributed Computing (Cat. No. 99TH8469),True,['resource-provisioning'],,,,,,,,,
773,Genetic programming approach for fault modeling of electronic hardware,"['A. Abraham', 'C. Grosan']",2005,"This paper presents two variants of genetic programming (GP) approaches for intelligent online performance monitoring of electronic circuits and systems. Reliability modeling of electronic circuits can be best performed by the stressor - susceptibility interaction model. A circuit or a system is deemed to be failed once the stressor has exceeded the susceptibility limits. For on-line prediction, validated stressor vectors may be obtained by direct measurements or sensors, which after preprocessing and standardization are fed into the GP models. Empirical results are compared with artificial neural networks trained using backpropagation algorithm. The performance of the proposed method is evaluated by comparing the experiment results with the actual failure model values. The developed model reveals that GP could play an important role for future fault monitoring systems.",https://ieeexplore.ieee.org/document/1554875,True,,8.0,['genetic-programming'],,"['comparison', 'novel-use']",['failure-prediction'],,2005 IEEE Congress on Evolutionary Computation,True,['failure-management'],,,['hardware-failure-prediction'],,,,,,
774,Bayesian approaches to failure prediction for disk drives,"['Greg Hamerly', 'Charles Elkan']",2001,,https://dl.acm.org/doi/10.5555/645530.655825,True,,209.0,['naive-bayes'],['host-metrics'],['novel-use'],['failure-prediction'],['best'],ICML '01: Proceedings of the Eighteenth International Conference on Machine Learning,True,['failure-management'],True,,['hardware-failure-prediction'],True,26.0,10.5555/645530.655825,,,
775,Hard drive failure prediction using non-parametric statistical methods,"['Murray, Joseph F', 'Hughes, Gordon F', 'Kreutz-Delgado, Kenneth']",2003,"We present a case study of a difficult real-world pattern recognition problem: predicting hard drive failure using attributes monitored internally by individual drives. We compare the performance of support vector machines (SVMs), unsupervised clustering, and non-parametric statistical tests (rank-sum and reverse arrangements). Somewhat surprisingly, the rank-sum method outperformed the other methods, including SVMs. We also show the utility of using non-parametric tests for feature set selection.",https://www.researchgate.net/publication/228972414_Hard_drive_failure_prediction_using_non-parametric_statistical_methods,True,,57.0,,,,['failure-prediction'],,Proc. ICANN/ICONIP 2003,True,['failure-management'],,,,,,,,,
776,Optimal discrimination between transient and permanent faults,"['M. Pizza', 'L. Strigini', 'A. Bondavalli', 'F. Di Giandomenico']",1998,"An important practical problem in fault diagnosis is discriminating between permanent faults and transient faults. In many computer systems, the majority of errors are due to transient faults. Many heuristic methods have been used for discriminating between transient and permanent faults; however, we have found no previous work stating this decision problem in clear probabilistic terms. We present an optimal procedure for discriminating between transient and permanent faults, based on applying Bayesian inference to the observed events (correct and erroneous results). We describe how the assessed probability that a module is permanently faulty must vary with observed symptoms. We describe and demonstrate our proposed method on a simple application problem, building the appropriate equations and showing numerical examples. The method can be implemented as a run-time diagnosis algorithm at little computational cost; it can also be used to evaluate any heuristic diagnostic procedure by comparison.",https://ieeexplore.ieee.org/document/731615,True,,46.0,,,,['failure-prediction'],,Proceedings Third IEEE International High-Assurance Systems Engineering Symposium (Cat. No. 98EX231),True,['failure-management'],,,,,,,,,
777,Failure Prediction in Hardware Systems,"['Turnbull, Doug', 'Alldrin, Neil']",2003,"We analyze hardware sensor data to predict failures in a high-end computer server. Features are extracted using sensor windows and potential failure windows. We then train radial basis function networks on these features and achieve a 0.87 true positive rate and 0.10 false positive rate for predicting failures using a data set where failures and non-failures are equally likely. This shows that sensor data can be used to predict failures in hardware systems. We conclude by discussing the effects of costs, benefits, and prediction accuracy in …",http://cseweb.ucsd.edu/~dturnbul/Papers/ServerPrediction.pdf,True,,24.0,,,,['failure-prediction'],,,True,['failure-management'],,,,,,,,,
778,Hidden Markov Models as a Support for Diagnosis: Formalization of the Problem and Synthesis of the Solution,"['Alessandro Daidone', 'Felicita Di Giandomenico', 'Andrea Bondavalli', 'Silvano Chiaradonna']",2006,"In modern information infrastructures, diagnosis must be able to assess the status or the extent of the damage of individual components. Traditional one-shot diagnosis is not adequate, but streams of data on component behavior need to be collected and filtered over time as done by some existing heuristics. This paper proposes instead a general framework and a formalism to model such over-time diagnosis scenarios, and to find appropriate solutions. As such, it is very beneficial to system designers to support design choices. Taking advantage of the characteristics of the hidden Markov models formalism, widely used in pattern recognition, the paper proposes a formalization of the diagnosis process, addressing the complete chain constituted by monitored component, deviation detection and state diagnosis. Hidden Markov models are well suited to represent problems where the internal state of a certain entity is not known and can only be inferred from external observations of what this entity emits. Such over-time diagnosis is a first class representative of this category of problems. The accuracy of diagnosis carried out through the proposed formalization is then discussed, as well as how to concretely use it to perform state diagnosis and allow direct comparison of alternative solutions",https://ieeexplore.ieee.org/document/4032486,True,,59.0,,,,['root-cause-analysis'],,2006 25th IEEE Symposium on Reliable Distributed Systems (SRDS'06),True,['failure-management'],,,,,,,,,
779,Anomalies as precursors of field failures,"['S. Elbaum', 'S. Kanduri', 'A. Amschler']",2003,"Reproducing and learning from failures in deployed software is costly and difficult. Those activities can be facilitated, however, if the circumstances leading to a failure are properly captured. In this paper, we empirically investigate how various anomaly detection schemes can serve to identify the conditions that precede failures in deployed software. Our results expose the tradeoffs between different detection algorithms applied to several types of events under varying levels of in-house testing.",https://ieeexplore.ieee.org/abstract/document/1251035,True,,15.0,,,,['failure-prediction'],,"14th International Symposium on Software Reliability Engineering, 2003. ISSRE 2003.",True,['failure-management'],,,,,,,,,
780,Combining Visualization and Statistical Analysis to Improve Operator Confidence and Efficiency for Failure Detection and Localization,"['P. Bodic', 'G. Friedman', 'L. Biewald', 'H. Levine', 'G. Candea', 'K. Patel', 'G. Tolle', 'J. Hui', 'A. Fox', 'M.I. Jordan', 'D. Patterson']",2005,"Web applications suffer from software and configuration faults that lower their availability. Recovering from failure is dominated by the time interval between when these faults appear and when they are detected by site operators. We introduce a set of tools that augment the ability of operators to perceive the presence of failure: an automatic anomaly detector scours HTTP access logs to find changes in user behavior that are indicative of site failures, and a visualizer helps operators rapidly detect and diagnose problems. Visualization addresses a key question of autonomic computing of how to win operators' confidence so that new tools will be embraced. Evaluation performed using HTTP logs from Ebates.com demonstrates that these tools can enhance the detection of failure as well as shorten detection time. Our approach is application-generic and can be applied to any Web application without the need for instrumentation",https://ieeexplore.ieee.org/document/1498055,True,,123.0,,['logs'],,"['failure-detection', 'root-cause-analysis']",,Second International Conference on Autonomic Computing (ICAC'05),True,['failure-management'],,,"['anomaly-detection', 'root-cause-diagnosis']",,,,,,
781,Path-based failure and evolution management,"['Mike Y. Chen', 'Anthony Accardi', 'Emre Kiciman', 'Jim Lloyd', 'Dave Patterson', 'Armando Fox', 'Eric Brewer']",2004,"We present a new approach to managing failures and evolution in large, complex distributed systems using runtime paths. We use the paths that requests follow as they move through the system as our core abstraction, and our ""macro"" approach focuses on component interactions rather than the details of the components themselves. Paths record component performance and interactions, are user- and request-centric, and occur in sufficient volume to enable statistical analysis, all in a way that is easily reusable across applications. Automated statistical analysis of multiple paths allows for the detection and diagnosis of complex failures and the assessment of evolution issues. In particular, our approach enables significantly stronger capabilities in failure detection, failure diagnosis, impact analysis, and understanding system evolution. We explore these capabilities with three real implementations, two of which service millions of requests per day. Our contributions include the approach; the maintainable, extensible, and reusable architecture; the various statistical analysis engines; and the discussion of our experience with a high-volume production service over several years.",https://dl.acm.org/doi/10.5555/1251175.1251198,True,,418.0,"['language-modeling', 'decision-tree']",['traces'],['novel-use'],"['root-cause-analysis', 'failure-detection']",,NSDI'04: Proceedings of the 1st conference on Symposium on Networked Systems Design and Implementation - Volume 1,True,['failure-management'],True,,['root-cause-diagnosis'],True,51.0,10.5555/1251175.1251198,,,
782,A methodology for detection and estimation of software aging,"['S. Garg', 'A. van Moorsel', 'K. Vaidyanathan', 'K.S. Trivedi']",1998,"The phenomenon of software aging refers to the accumulation of errors during the execution of the software which eventually results in it's crash/hang failure. A gradual performance degradation may also accompany software aging. Pro-active fault management techniques such as ""software rejuvenation"" (Y. Huang et al., 1995) may be used to counteract aging if it exists. We propose a methodology for detection and estimation of aging in the UNIX operating system. First, we present the design and implementation of an SNMP based, distributed monitoring tool used to collect operating system resource usage and system activity data at regular intervals, from networked UNIX workstations. Statistical trend detection techniques are applied to this data to detect/validate the existence of aging. For quantifying the effect of aging in operating system resources, we propose a metric: ""estimated time to exhaustion"", which is calculated using well known slope estimation techniques. Although the distributed data collection tool is specific to UNIX, the statistical techniques can be used for detection and estimation of aging in other software as well.",https://ieeexplore.ieee.org/document/730892,True,,369.0,['linear-regression'],['host-metrics'],['novel-use'],['failure-prevention'],['software'],Proceedings Ninth International Symposium on Software Reliability Engineering (Cat. No. 98TB100257),True,['failure-management'],True,,['rejuvenation'],True,17.0,,,,
783,Proactive management of software aging,"['V. Castelli', 'R. E. Harper', 'P. Heidelberger', 'S. W. Hunter', 'K. S. Trivedi', 'K. Vaidyanathan', 'W. P. Zeggert']",2001,"Software failures are now known to be a dominant source of system outages. Several studies and much anecdotal evidence point to “software aging” as a common phenomenon, in which the state of a software system degrades with time. Exhaustion of system resources, data corruption, and numerical error accumulation are the primary symptoms of this degradation, which may eventually lead to performance degradation of the software, crash/hang failure, or other undesirable effects. “Software rejuvenation” is a proactive technique intended to reduce the probability of future unplanned outages due to aging. The basic idea is to pause or halt the running software, refresh its internal state, and resume or restart it. Software rejuvenation can be performed by relying on a variety of indicators of aging, or on the time elapsed since the last rejuvenation. In response to the strong desire of customers to be provided with advance notice of unplanned outages, our group has developed techniques that dete ct the occurrence of software aging due to resource exhaustion, estimate the time remaining until the exhaustion reaches a critical level, and automatically perform proactive software rejuvenation of an application, process group, or entire operating system, depending on the pervasiveness of the resource exhaustion and our ability to pinpoint the source. This technology has been incorporated into the IBM Director for xSeries servers. To quantitatively evaluate the impact of different rejuvenation policies on the availability of cluster systems, we have developed analytical models based on stochastic reward nets (SRNs). For time-based rejuvenation policies, we determined the optimal rejuvenation interval based on system availability and cost. We also analyzed a rejuvenation policy based on prediction, and showed that it can further increase system availability and reduce downtime cost. These models are very general and can capture a multitude of cluster system characteristics, failure ...",https://ieeexplore.ieee.org/document/5389090,True,,385.0,"['linear-regression', 'petri-net']",['host-metrics'],,['failure-prevention'],,,True,['failure-management'],True,True,['rejuvenation'],True,20.0,,,,
784,Application Cluster Service Scheme for Near-Zero-Downtime Services,"['Fan-Tien Cheng', 'Shang-Lun Wu', 'Ping-Yen Tsai', 'Yun-Ta Chung', 'Haw-Ching Yang']",2005,"The required reliability in applications of a distributed computer system is continuous service for 24 hours a day, 7 days a week. However, computer failures due to exhaustion of operating system resources, data corruption, numerical error accumulation, and so on, may interrupt services and cause significant losses. Hence, this work proposes an application cluster service (APCS) scheme. The proposed APCS provides both a failover scheme and a state recovery scheme for failure management. The failover scheme is designed mainly to automatically activate the backup application for replacing the failed application whenever it is sick or down. Meanwhile, the state recovery scheme is intended primarily to provide an inheritable design pattern to support applications with state recovery requirements. An application simply needs to inherit and implement this design pattern, and then can accomplish the task of state backup and recovery. Furthermore, a performance evaluator (PEV) that can detect performance degradation and predict time to failure is developed in this study. By using these detection and prediction capabilities, the APCS can perform the failover process before node breakdown. Thus, applying APCS and PEV can enable a distributed computer system to provide services with near-zero-downtime.",https://ieeexplore.ieee.org/document/1570743,True,,14.0,,,,['failure-prediction'],,Proceedings of the 2005 IEEE International Conference on Robotics and Automation,True,['failure-management'],,,,,,,,,
785,Software aging and multifractality of memory resources,"['M. Shereshevsky', 'J. Crowell', 'B. Cukic', 'V. Gandikota', 'Yan Liu']",2003,"We investigate the dynamics of monitored memory resource utilizations in an operating system under stress using quantitative methods of fractal analysis. In the experiments, we recorded the time series representing various memory related parameters of the operating system. We observed that parameters demonstrate clear multifractal behavior. The degree of fractality of these time series tends to increase as the system workload increases. We conjecture that the Hölder exponent that measures the local rate of fractality may be used as …",https://ieeexplore.ieee.org/document/1209987,True,,77.0,,,,['failure-prediction'],,"2003 International Conference on Dependable Systems and Networks, 2003. Proceedings.",True,['failure-management'],,,,,,,,,
786,An approach to predictive detection for service management,"['J.L. Hellerstein', 'Fan Zhang', 'P. Shahabuddin']",1999,"Service providers typically define quality of service problems using threshold tests, such as ""are HTTP operations greater than 12 per second on server XYZ?"" This paper explores the feasibility of predicting violations of threshold tests. Such a capability would allow providers to take corrective actions in advance of service disruptions. Our approach estimates the probability of threshold violations for specific times in the future. We modeled the threshold metric (e.g., HTTP operations per second) at two levels: (1) nonstationary behavior (as is done in workload forecasting for capacity planning) and (2) stationary, time-serial dependencies. Using these models, we compute the probability of threshold violations. We asses our approach using measurements of HTTP operations per second collected from a production Web server. These assessments suggest that our approach works well if: (a) the actual values of predicted metrics are sufficiently distant from their thresholds; and/or (b) the prediction horizon is not too far into the future.",https://ieeexplore.ieee.org/document/770691,True,,96.0,['autoregression'],,['new-method'],"['failure-prediction', 'failure-detection', 'anomaly-detection', 'workload-prediction', 'system-failure-prediction']",,Integrated Network Management VI. Distributed Management for the Networked Millennium. Proceedings of the Sixth IFIP/IEEE International Symposium on Integrated Network Management.(Cat. No. 99EX302),True,"['resource-provisioning', 'failure-management']",,,,,,,,,
787,Critical event prediction for proactive management in large-scale computer clusters,"['R. K. Sahoo', 'A. J. Oliner', 'I. Rish', 'M. Gupta', 'J. E. Moreira', 'S. Ma', 'R. Vilalta', 'A. Sivasubramaniam']",2003,"As the complexity of distributed computing systems increases, systems management tasks require significantly higher levels of automation; examples include diagnosis and prediction based on real-time streams of computer events, setting alarms, and performing continuous monitoring. The core of autonomic computing, a recently proposed initiative towards next-generation IT-systems capable of 'self-healing', is the ability to analyze data in real-time and to predict potential problems. The goal is to avoid catastrophic failures through prompt execution of remedial actions.This paper describes an attempt to build a proactive prediction and control system for large clusters. We collected event logs containing various system reliability, availability and serviceability (RAS) events, and system activity reports (SARs) from a 350-node cluster system for a period of one year. The 'raw' system health measurements contain a great deal of redundant event data, which is either repetitive in nature or misaligned with respect to time. We applied a filtering technique and modeled the data into a set of primary and derived variables. These variables used probabilistic networks for establishing event correlations through prediction algorithms. We also evaluated the role of time-series methods, rule-based classification algorithms and Bayesian network models in event prediction.Based on historical data, our results suggest that it is feasible to predict system performance parameters (SARs) with a high degree of accuracy using time-series models. Rule-based classification techniques can be used to extract machine-event signatures to predict critical events with up to 70% accuracy.",https://dl.acm.org/doi/10.1145/956750.956799,True,,167.0,['bayesian-network'],['logs'],,['failure-prediction'],,KDD '03: Proceedings of the ninth ACM SIGKDD international conference on Knowledge discovery and data mining,True,['failure-management'],True,,['system-failure-prediction'],,,10.1145/956750.956799,,,
788,Software Aging Prediction Model Based on Fuzzy Wavelet Network with Adaptive Genetic Algorithm,"['Meng Hai Ning', 'Qi Yong', 'Hou Di', 'Chen Ying', 'Zhao Ji Zhong']",2006,"According to the characteristics of the operational behavior and runtime state of application sever, the resource consumption time series are observed and modeled by fuzzy wavelet network (FWN) with fuzzy logic inference and learning capability. The objective is to model the extracted data series of systematic performance parameters to predict software aging in application server. The dimensionality of input variables of FWN is reduced by principal components analysis (PCA), and the structure and parameters of FWN are optimized with adaptive genetic algorithm (GA). Judging by the model, we can get the aging threshold before application server failed and preventively maintenance the application server before systematic parameter value reaches the threshold. The experiments are carried out to validate the efficiency of the proposed model and show that the aging prediction model based on FWN with adaptive genetic algorithm is superior to the neural network (NN) model and wavelet network (WN) model in the aspects of convergence rate and prediction precision",https://ieeexplore.ieee.org/abstract/document/4031957,True,,17.0,,,,['failure-prediction'],,2006 18th IEEE International Conference on Tools with Artificial Intelligence (ICTAI'06),True,['failure-management'],,,,,,,,,
789,TASA: Telecommunication Alarm Sequence Analyzer or how to enjoy faults in your network,"['K. Hatonen', 'M. Klemettinen', 'H. Mannila', 'P. Ronkainen', 'H. Toivonen']",1996,"Today's large and complex telecommunication networks produce large amounts of alarms daily. The sequence of alarms contains valuable knowledge about the behavior of the network, but much of the knowledge is fragmented and hidden in the vast amount of data. Regularities in the alarms can be used in fault management applications, e.g., for filtering redundant alarms, locating problems in the network, and possibly in predicting severe faults. In this paper we describe TASA (Telecommunication Alarm Sequence Analyzer), a novel system for discovering interesting regularities in the alarms. In the core of the system are algorithms for locating frequent alarm episodes from the alarm stream and presenting them as rules. Discovered rules can then be explored with flexible information retrieval tools that support iteration. The user interface is hypertext, based on HTML, and can be used with a standard WWW browser. TASA is in experimental use and has already discovered rules that have been integrated into the alarm handling software of an operator.",https://ieeexplore.ieee.org/abstract/document/539622,True,,71.0,,,,['failure-prediction'],,Proceedings of NOMS'96-IEEE Network Operations and Management Symposium,True,['failure-management'],,,,,,,,,
790,Industry: predicting telecommunication equipment failures from sequences of network alarms,['Gary M. Weiss'],2002,"The computer and telecommunication industries rely heavily on knowledge-based expert systems to manage the performance of their networks. These expert systems are developed by knowledge engineers, who must first interview domain experts to extract the pertinent knowledge. This knowledge acquisition process is laborious and costly, and typically is better at capturing qualitative knowledge than quantitative knowledge. This is a liability, especially for domains like the telecommunication domain, where enormous amounts of data are readily available for analysis. Data mining holds tremendous promise for the development of expert systems for monitoring network performance since it provides a way of automatically identifying subtle, yet important, patterns in data. This case study describes a project in which a temporal data mining system called Timeweaver is used to identify faulty telecommunication equipment from logs of network alarm messages.",https://dl.acm.org/doi/10.5555/778212.778336,True,,35.0,,,,['failure-prediction'],,Handbook of data mining and knowledge discovery,True,['failure-management'],,,,,,10.5555/778212.778336,,,
791,Predicting rare events in temporal domains,"['R. Vilalta', 'Sheng Ma']",2002,"Temporal data mining aims at finding patterns in historical data. Our work proposes an approach to extract temporal patterns from data to predict the occurrence of target events, such as computer attacks on host networks, or fraudulent transactions in financial institutions. Our problem formulation exhibits two major challenges: 1) we assume events being characterized by categorical features and displaying uneven inter-arrival times; such an assumption falls outside the scope of classical time-series analysis, 2) we assume target events are highly infrequent; predictive techniques must deal with the class-imbalance problem. We propose an efficient algorithm that tackles the challenges above by transforming the event prediction problem into a search for all frequent eventsets preceding target events. The class imbalance problem is overcome by a search for patterns on the minority class exclusively; the discrimination power of patterns is then validated against other classes. Patterns are then combined into a rule-based model for prediction. Our experimental analysis indicates the types of event sequences where target events can be accurately predicted.",https://ieeexplore.ieee.org/document/1183991,True,,70.0,,,,['failure-prediction'],,"2002 IEEE International Conference on Data Mining, 2002. Proceedings.",True,['failure-management'],,,,,,,,,
792,Software failure prediction based on a Markov Bayesian network model,"['C.G. Bai', 'Q.P. Hu', 'M. Xie', 'S.H. Ng']",2005,"Due to the complexity of software products and development processes, software reliability models need to possess the ability of dealing with multiple parameters. Also in order to adapt to the continually refreshed data, they should provide flexibility in model construction in terms of information updating. Existing software reliability models are not flexible in this context. The main reason for this is that there are many static assumptions associated with the models. Bayesian network is a powerful tool for solving this problem, as it exhibits strong ability to adapt in problems involving complex variant factors. In this paper, a software prediction model based on Markov Bayesian networks is developed, and a method to solve the network model is proposed. The use of our model is illustrated with an example.",https://www.sciencedirect.com/science/article/pii/S0164121204000081?via%3Dihub,True,,60.0,,,,['failure-prediction'],,Journal of Systems and Software,True,['failure-management'],,,,,,10.1016/j.jss.2004.02.028,,,
793,A Classification Approach for Prediction of Target Events in Temporal Sequences,"['Carlotta Domeniconi', 'Chang-Shing Perng', 'Ricardo Vilalta', 'Sheng Ma']",2002,"Learning to predict significant events from sequences of data with categorical features is an important problem in many application areas. We focus on events for system management, and formulate the problem of prediction as a classification problem. We perform co-occurrence analysis of events by means of Singular Value Decomposition (SVD) of the examples constructed from the data. This process is combined with Support Vector Machine (SVM) classification, to obtain efficient and accurate predictions. We conduct an analysis of statistical properties of event data, which explains why SVM classification is suitable for such data, and perform an empirical study using real data.",https://dl.acm.org/doi/10.5555/645806.670309,True,,28.0,,,,['failure-prediction'],,PKDD '02: Proceedings of the 6th European Conference on Principles of Data Mining and Knowledge Discovery,True,['failure-management'],,,,,,10.5555/645806.670309,,,
794,Structured Comparative Analysis of Systems Logs to Diagnose Performance Problems,"['Karthik Nagaraj', 'Charles Killian', 'Jennifer Neville']",2012,"Diagnosis and correction of performance issues in modern, large-scale distributed systems can be a daunting task, since a single developer is unlikely to be familiar with the entire system and it is hard to characterize the behavior of a software system without completely understanding its internal components. This paper describes DISTALYZER, an automated tool to support developer investigation of performance issues in distributed systems. We aim to leverage the vast log data available from large scale systems, while reducing the level of knowledge required for a developer to use our tool. Specifically, given two sets of logs, one with good and one with bad performance, DISTALYZER uses machine learning techniques to compare system behaviors extracted from the logs and automatically infer the strongest associations between system components and performance. The tool outputs a set of inter-related event occurrences and variable values that exhibit the largest divergence across the logs sets and most directly affect the overall performance of the system. These patterns are presented to the developer for inspection, to help them understand which system component(s) likely contain the root cause of the observed performance issue, thus alleviating the need for many human hours of manual inspection. We demonstrate the generality and effectiveness of DISTALYZER on three real distributed systems by showing how it discovers and highlights the root cause of six performance issues across the systems. DISTALYZER has broad applicability to other systems since it is dependent only on the logs for input, and not on the source code.",https://dl.acm.org/doi/10.5555/2228298.2228334,True,,56.0,['bayesian-network'],['logs'],['new-method'],['root-cause-analysis'],,NSDI'12: Proceedings of the 9th USENIX conference on Networked Systems Design and Implementation,True,['failure-management'],True,,['root-cause-diagnosis'],,,10.5555/2228298.2228334,,,
795,Diagnosing Latency in Multi-Tier Black-Box Services,"['Ostrowski, Krzysztof', 'Mann, Gideon', 'Sandler, Mark']",2011,"As multi-tier cloud applications become pervasive, we need better tools for understanding their performance. This paper presents a system that analyzes observed or desired changes to end-to-end latency pro le in a large distributed application, and identi fies their underlying causes. It recognizes changes to system con guration, workload, or performance of individual services that lead to the observed or desired outcome. Experiments on an industrial datacenter demonstrate the utility of the system.",https://research.google/pubs/pub37477/,True,,17.0,"['similarity-matching', 'probabilistic-modeling']","['kpis', 'traces', 'host-metrics']",,"['failure-prediction', 'root-cause-analysis']",,,True,['failure-management'],,,['anomaly-detection'],,,,,,
796,Adaptive on-line software aging prediction based on machine learning,"['Javier Alonso', 'Jordi Torres', 'Josep Ll. Berral', 'Ricard Gavaldà']",2010,"The growing complexity of software systems is resulting in an increasing number of software faults. According to the literature, software faults are becoming one of the main sources of unplanned system outages, and have an important impact on company benefits and image. For this reason, a lot of techniques (such as clustering, fail-over techniques, or server redundancy) have been proposed to avoid software failures, and yet they still happen. Many software failures are those due to the software aging phenomena. In this work, we present a detailed evaluation of our chosen machine learning prediction algorithm (M5P) in front of dynamic and non-deterministic software aging. We have tested our prediction model on a three-tier web J2EE application achieving acceptable prediction accuracy against complex scenarios with small training data sets. Furthermore, we have found an interesting approach to help to determine the root cause failure: The model generated by machine learning algorithms.",https://ieeexplore.ieee.org/document/5544275,True,,77.0,['regression-tree'],['host-metrics'],['novel-use'],['failure-prevention'],['software'],2010 IEEE/IFIP International Conference on Dependable Systems \& Networks (DSN),True,['failure-management'],True,,['rejuvenation'],True,19.0,,,,
797,Using machine learning for non-intrusive modeling and prediction of software aging,"['Artur Andrzejak', 'Luis Silva']",2008,"The wide-spread phenomenon of software (running image) aging is known to cause performance degradation, transient failures or even crashes of applications. In this work we describe first a method for monitoring and modeling of performance degradation in SOA applications, particularly application servers. This method works for a large class of the aging processes caused by resource depletion (e.g. memory leaks). It can be deployed non-intrusively in a production environment, under arbitrary service request distributions. Based on this schema we investigate in the second part of the paper how machine learning (classification) algorithms can be used for proactive detection of performance degradation or sudden drops caused by aging. We leverage the predictive power of these algorithms with several techniques to make the measurement-based aging models more adaptive and more robust against transient failures. We evaluate several state-of-the-art classification methods for their accuracy and computational efficiency in this scenario. The studies are performed on a data set generated by a TPC-W benchmark instrumented with a memory leak injector. The results show that the probing method yields accurate aging models with low overhead and the machine learning approach gives statistically significant short-term predictions of degrading application performance. Both approaches can be used directly to fight aging via adaptive software rejuvenation (restart of the application), for operator alerting, or for short-term capacity planning.",https://ieeexplore.ieee.org/document/4575113,True,,37.0,,,,['failure-prediction'],,NOMS 2008-2008 IEEE Network Operations and Management Symposium,True,['failure-management'],,,,,,,,,
798,Accurate proactive adaptation of service-oriented systems,"['Metzger, Andreas', 'Sammodi, Osama', 'Pohl, Klaus']",2013,"As service-oriented systems are increasingly composed of third-party services accessible over the Internet, self-adaptation capabilities promise to make these systems become robust and resilient against third-party service failures that may negatively impact on system quality. In such a setting, proactive adaptation capabilities will provide significant benefits by predicting pending service failures and mitigating their negative impact on system quality. Proactive adaptation requires accurate quality prediction techniques; firstly, because …",https://link.springer.com/chapter/10.1007/978-3-642-36249-1_9,True,,30.0,,,,['failure-prediction'],,Assurances for Self-Adaptive Systems,True,['failure-management'],,,,,,,,,
799,Finding soon-to-fail disks in a haystack,['Moises Goldszmidt'],2012,"This paper presents a detector of soon-to-fail disks based on a combination of statistical models. During operation the detector takes as input a performance signal from each disk and sends and alarm when there is enough evidence (according to the models) that the disk is not healthy. The parameters of these models are automatically trained using signals from healthy and failed disks. In an evaluation on a population of 1190 production disks from a popular customer-facing internet service, the detector was able to predict 15 out of the 17 failed disks (88:2% detection) with 30 false alarms (2:56% false positive rate).",https://dl.acm.org/doi/10.5555/2342806.2342814,True,,28.0,,,,['failure-prediction'],,HotStorage'12: Proceedings of the 4th USENIX conference on Hot Topics in Storage and File Systems,True,['failure-management'],,,,,,10.5555/2342806.2342814,,,
800,Improving storage system reliability with proactive error prediction,"['Farzaneh Mahdisoltani', 'Ioan Stefanovici', 'Bianca Schroeder']",2017,"This paper proposes the use of machine learning techniques to make storage systems more reliable in the face of sector errors. Sector errors are partial drive failures, where individual sectors on a drive become unavailable, and occur at a high rate in both hard disk drives and solid state drives. The data in the affected sectors can only be recovered through redundancy in the system (e.g. another drive in the same RAID) and is lost if the error is encountered while the system operates in degraded mode, e.g. during RAID reconstruction. In this paper, we explore a range of different machine learning techniques and show that sector errors can be predicted ahead of time with high accuracy. Prediction is robust, even when only little training data or only training data for a different drive model is available. We also discuss a number of possible use cases for improving storage system reliability through the use of sector error predictors. We evaluate one such use case in detail: We show that the mean time to detecting errors (and hence the window of vulnerability to data loss) can be greatly reduced by adapting the speed of a scrubber based on error predictions.",https://dl.acm.org/doi/10.5555/3154690.3154728,True,,7.0,,,,['failure-prediction'],,USENIX ATC '17: Proceedings of the 2017 USENIX Conference on Usenix Annual Technical Conference,True,['failure-management'],,,,,,10.5555/3154690.3154728,,,
801,Online Defect Prediction for Imbalanced Data,"['Ming Tan', 'Lin Tan', 'Sashank Dara', 'Caleb Mayeux']",2015,"Many defect prediction techniques are proposed to improve software reliability. Change classification predicts defects at the change level, where a change is the modifications to one file in a commit. In this paper, we conduct the first study of applying change classification in practice. We identify two issues in the prediction process, both of which contribute to the low prediction performance. First, the data are imbalanced -- there are much fewer buggy changes than clean changes. Second, the commonly used cross-validation approach is inappropriate for evaluating the performance of change classification. To address these challenges, we apply and adapt online change classification, resampling, and updatable classification techniques to improve the classification performance. We perform the improved change classification techniques on one proprietary and six open source projects. Our results show that these techniques improve the precision of change classification by 12.2-89.5% or 6.4 -- 34.8 percentage points (pp.) on the seven projects. In addition, we integrate change classification in the development process of the proprietary project. We have learned the following lessons: 1) new solutions are needed to convince developers to use and believe prediction results, and prediction results need to be actionable, 2) new and improved classification algorithms are needed to explain the prediction results, and insensible and unactionable explanations need to be filtered or refined, and 3) new techniques are needed to improve the relatively low precision.",https://ieeexplore.ieee.org/document/7202954,True,,121.0,"['naive-bayes', 'logistic-regression', 'support-vector-machine', 'linear-regression', 'multilayer-perceptron', 'entropy-selection', 'decision-tree', 'clustering']","['source-code', 'code-metrics']",['novel-use'],['failure-prevention'],['software'],2015 IEEE/ACM 37th IEEE International Conference on Software Engineering,True,['failure-management'],False,,['software-defect-prediction'],,,,,,"['online', 'imbalance']"
802,Online detection of utility cloud anomalies using metric distributions,"['Chengwei Wang', 'Vanish Talwar', 'Karsten Schwan', 'Parthasarathy Ranganathan']",2010,"The online detection of anomalies is a vital element of operations in data centers and in utility clouds like Amazon EC2. Given ever-increasing data center sizes coupled with the complexities of systems software, applications, and workload patterns, such anomaly detection must operate automatically, at runtime, and without the need for prior knowledge about normal or anomalous behaviors. Further, detection should function for different levels of abstraction like hardware and software, and for the multiple metrics used in cloud computing systems. This paper proposes EbAT - Entropy-based Anomaly Testing - offering novel methods that detect anomalies by analyzing for arbitrary metrics their distributions rather than individual metric thresholds. Entropy is used as a measurement that captures the degree of dispersal or concentration of such distributions, aggregating raw metric data across the cloud stack to form entropy time series. For scalability, such time series can then be combined hierarchically and across multiple cloud subsystems. Experimental results on utility cloud scenarios demonstrate the viability of the approach. EbAT outperforms threshold-based methods with on average 57.4% improvement in accuracy of anomaly detection and also does better by 59.3% on average in false alarm rate with a `near-optimum' threshold-based method.",https://ieeexplore.ieee.org/document/5488443,True,,129.0,['entropy-selection'],['host-metrics'],"['new-method', 'comparison']",['failure-detection'],"['vm', 'cloud']",2010 IEEE Network Operations and Management Symposium-NOMS 2010,True,['failure-management'],,,['anomaly-detection'],,,,,,['online']
803,Online Anomaly Detection for Hard Disk Drives Based on Mahalanobis Distance,"['Yu Wang', 'Qiang Miao', 'Eden W. M. Ma', 'Kwok-Leung Tsui', 'Michael G. Pecht']",2013,"A hard disk drive (HDD) failure may cause serious data loss and catastrophic consequences. Online health monitoring provides information about the degradation trend of the HDD, and hence the early warning of failures, which gives us a chance to save the data. This paper developed an approach for HDD anomaly detection using Mahalanobis distance (MD). Critical parameters were selected using failure modes, mechanisms, and effects analysis (FMMEA), and the minimum redundancy maximum relevance (mRMR) method. A self-monitoring, analysis, and reporting technology (SMART) data set is used to evaluate the performance of the developed approach. The result shows that about 67% of the anomalies of failed drives can be detected with zero false alarm rate, and most of them can provide users with at least 20 hours during which to backup the data.",https://ieeexplore.ieee.org/document/6423861,True,,68.0,"['dimensionality-reduction', 'similarity-matching', 'entropy-selection']",['host-metrics'],"['novel-use', 'comparison']",['failure-prediction'],['hard-drive'],,True,['failure-management'],True,,['hardware-failure-prediction'],True,29.0,,,,
804,The mystery machine: end-to-end performance analysis of large-scale internet services,"['Michael Chow', 'David Meisner', 'Jason Flinn', 'Daniel Peek', 'Thomas F. Wenisch']",2014,"Current debugging and optimization methods scale poorly to deal with the complexity of modern Internet services, in which a single request triggers parallel execution of numerous heterogeneous software components over a distributed set of computers. The Achilles' heel of current methods is the need for a complete and accurate model of the system under observation: producing such a model is challenging because it requires either assimilating the collective knowledge of hundreds of programmers responsible for the individual components or restricting the ways in which components interact. Fortunately, the scale of modern Internet services offers a compensating benefit: the sheer volume of requests serviced means that, even at low sampling rates, one can gather a tremendous amount of empirical performance observations and apply ""big data"" techniques to analyze those observations. In this paper, we show how one can automatically construct a model of request execution from pre-existing component logs by generating a large number of potential hypotheses about program behavior and rejecting hypotheses contradicted by the empirical observations. We also show how one can validate potential performance improvements without costly implementation effort by leveraging the variation in component behavior that arises naturally over large numbers of requests to measure the impact of optimizing individual components or changing scheduling behavior. We validate our methodology by analyzing performance traces of over 1.3 million requests to Facebook servers. We present a detailed study of the factors that affect the end-to-end latency of such requests. We also use our methodology to suggest and validate a scheduling optimization for improving Facebook request latency.",https://dl.acm.org/doi/10.5555/2685048.2685066,True,,142.0,['graph-mining'],"['logs', 'traces']",,"['failure-detection', 'root-cause-analysis']",,OSDI'14: Proceedings of the 11th USENIX conference on Operating Systems Design and Implementation,True,['failure-management'],True,,['anomaly-detection'],True,60.0,10.5555/2685048.2685066,,,
805,SherLog: error diagnosis by connecting clues from run-time logs,"['Ding Yuan', 'Haohui Mai', 'Weiwei Xiong', 'Lin Tan', 'Yuanyuan Zhou', 'Shankar Pasupathy']",2010,"Computer systems often fail due to many factors such as software bugs or administrator errors. Diagnosing such production run failures is an important but challenging task since it is difficult to reproduce them in house due to various reasons: (1) unavailability of users' inputs and file content due to privacy concerns; (2) difficulty in building the exact same execution environment; and (3) non-determinism of concurrent executions on multi-processors.Therefore, programmers often have to diagnose a production run failure based on logs collected back from customers and the corresponding source code. Such diagnosis requires expert knowledge and is also too time-consuming, tedious to narrow down root causes. To address this problem, we propose a tool, called SherLog, that analyzes source code by leveraging information provided by run-time logs to infer what must or may have happened during the failed production run. It requires neither re-execution of the program nor knowledge on the log's semantics. It infers both control and data value information regarding to the failed execution.We evaluate SherLog with 8 representative real world software failures (6 software bugs and 2 configuration errors) from 7 applications including 3 servers. Information inferred by SherLog are very useful for programmers to diagnose these evaluated failures. Our results also show that SherLog can analyze large server applications such as Apache with thousands of logging messages within only 40 minutes.",https://dl.acm.org/doi/10.1145/1735970.1736038,True,,255.0,,"['logs', 'source-code']",['new-method'],['root-cause-analysis'],"['source-code', 'apache', 'software']",ACM SIGARCH Computer Architecture News,True,['failure-management'],True,,['fault-localization'],True,80.0,10.1145/1735970.1736038,,,
806,"Toward Fine-Grained, Unsupervised, Scalable Performance Diagnosis for Production Cloud Computing Systems","['Haibo Mi', 'Huaimin Wang', 'Yangfan Zhou', 'Michael Rung-Tsong Lyu', 'Hua Cai']",2013,"Performance diagnosis is labor intensive in production cloud computing systems. Such systems typically face many real-world challenges, which the existing diagnosis techniques for such distributed systems cannot effectively solve. An efficient, unsupervised diagnosis tool for locating fine-grained performance anomalies is still lacking in production cloud computing systems. This paper proposes CloudDiag to bridge this gap. Combining a statistical technique and a fast matrix recovery algorithm, CloudDiag can efficiently pinpoint fine-grained causes of the performance problems, which does not require any domain-specific knowledge to the target system. CloudDiag has been applied in a practical production cloud computing systems to diagnose performance problems. We demonstrate the effectiveness of CloudDiag in three real-world case studies.",https://ieeexplore.ieee.org/document/6410318,True,,76.0,['dimensionality-reduction'],['logs'],['novel-use'],['failure-detection'],,,True,['failure-management'],,,['anomaly-detection'],,,,,,
807,A System Architecture for Real-time Anomaly Detection in Large-scale NFV Systems,"['Gulenko, Anton', 'Wallschl{\\""a}ger, Marcel', 'Schmidt, Florian', 'Kao, Odej', 'Liu, Feng']",2016,"Virtualization as a key IT technology has developed to a predominant model in data centers in recent years. The flexibility regarding scaling-out and migration of virtual machines for seamless maintenance has enabled a new level of continuous operation and changed service provisioning significantly. Meanwhile, services from domains striving for highest possible availability–eg from the telecommunications domain–are adopting this approach as well and are investing significant efforts into the development of Network Function …",https://www.sciencedirect.com/science/article/pii/S1877050916318269?via%3Dihub,True,,10.0,,,,['failure-detection'],,,True,['failure-management'],,,,,,,,,
808,Failure prediction using machine learning in a virtualised HPC system and application,"['Bashir Mohammed', 'Irfan Awan', 'Hassan Ugail', 'Muhammad Younas']",2019,"Failure is an increasingly important issue in high performance computing and cloud systems. As large-scale systems continue to grow in scale and complexity, mitigating the impact of failure and providing accurate predictions with sufficient lead time remains a challenging research problem. Traditional existing fault-tolerance strategies such as regular check-pointing and replication are not adequate because of the emerging complexities of high performance computing systems. This necessitates the importance of having an effective as well as proactive failure management approach in place aimed at minimizing the effect of failure within the system. With the advent of machine learning techniques, the ability to learn from past information to predict future pattern of behaviours makes it possible to predict potential system failure more accurately. Thus, in this paper, we explore the predictive abilities of machine learning by applying a number of algorithms to improve the accuracy of failure prediction. We have developed a failure prediction model using time series and machine learning, and performed comparison based tests on the prediction accuracy. The primary algorithms we considered are the support vector machine (SVM), random forest (RF), k-nearest neighbors (KNN), classification and regression trees (CART) and linear discriminant analysis (LDA). Experimental results indicates that the average prediction accuracy of our model using SVM when predicting failure is 90% accurate and effective compared to other algorithms. This finding implies that our method can effectively predict all possible future system and application failures within the system.",https://dl.acm.org/doi/10.1007/s10586-019-02917-1,True,,5.0,"['random-forest', 'logistic-regression', 'support-vector-machine', 'regression-tree', 'clustering']",,['comparison'],['failure-prediction'],,Cluster Computing,True,['failure-management'],,,['system-failure-prediction'],,,10.1007/s10586-019-02917-1,,,
809,Long Short-Term Memory Network for Remaining Useful Life estimation,"['Shuai Zheng', 'Kosta Ristovski', 'Ahmed Farahat', 'Chetan Gupta']",2017,"Remaining Useful Life (RUL) of a component or a system is defined as the length from the current time to the end of the useful life. Accurate RUL estimation plays a critical role in Prognostics and Health Management(PHM). Data driven approaches for RUL estimation use sensor data and operational data to estimate RUL. Traditional regression based approaches and recent Convolutional Neural Network (CNN) approach use features created from sliding windows to build models. However, sequence information is not fully considered in these approaches. Sequence learning models such as Hidden Markov Models (HMMs) and Recurrent Neural Networks (RNNs) have flaws when modeling sequence information. HMMs are limited to discrete hidden states and are known to have issues when modeling long-term dependencies in the data. RNNs also have issues with long-term dependencies. In this work, we propose a Long Short-Term Memory (LSTM) approach for RUL estimation, which can make full use of the sensor sequence information and expose hidden patterns within sensor data with multiple operating conditions, fault and degradation models. Extensive experiments using three widely adopted Prognostics and Health Management data sets show that LSTM for RUL estimation significantly outperforms traditional approaches for RUL estimation as well as Convolutional Neural Network (CNN).",https://ieeexplore.ieee.org/document/7998311,True,,121.0,['rnn'],,"['comparison', 'novel-use']",['failure-prediction'],['hardware'],,True,['failure-management'],True,,['hardware-failure-prediction'],True,41.0,,,,
810,Alarm correlation and fault identification in communication networks,"['A.T. Bouloutas', 'S. Calo', 'A. Finkel']",1994,"Presents an approach for modeling and solving the problem of fault identification and alarm correlation in large communication networks. A single fault in a large network may result in a large number of alarms, and it is often very difficult to isolate the true cause of the fault. This appears to be one of the most important difficulties in managing faults in today's networks. The problem may become worse in the case of multiple faults. The authors present a general methodology for solving the alarm correlation and fault identification problem. They propose a new alarm structure, propose a general model for representing the network, and give two algorithms which can solve the alarm correlation and fault identification problem in the presence of multiple faults. These algorithms differ in the degree of accuracy achieved in identifying the fault, and in the degree of complexity required for implementation.<
>",https://ieeexplore.ieee.org/document/577079,True,,273.0,"['automaton', 'search', 'bayesian-network']",,,"['root-cause-analysis', 'failure-detection']",['network'],,True,['failure-management'],False,,['fault-localization'],,,,,,
811,Data Mining Static Code Attributes to Learn Defect Predictors,"['Tim Menzies', 'Jeremy Greenwald', 'Art Frank']",2006,"The value of using static code attributes to learn defect predictors has been widely debated. Prior work has explored issues like the merits of ""McCabes versus Halstead versus lines of code counts"" for generating defect predictors. We show here that such debates are irrelevant since how the attributes are used to build predictors is much more important than which particular attributes are used. Also, contrary to prior pessimism, we show that such defect predictors are demonstrably useful and, on the data studied here, yield predictors with a mean probability of detection of 71 percent and mean false alarms rates of 25 percent. These predictors would be useful for prioritizing a resource-bound exploration of code that has yet to be inspected",https://ieeexplore.ieee.org/document/4027145,True,,1246.0,"['naive-bayes', 'decision-tree']",['code-metrics'],"['new-method', 'discussion']",['failure-prevention'],['source-code'],,True,['failure-management'],True,,['software-defect-prediction'],True,2.0,,,,
812,Predicting component failures at design time,"['Adrian Schröter', 'Thomas Zimmermann', 'Andreas Zeller']",2006,"How do design decisions impact the quality of the resulting software? In an empirical study of 52 ECLIPSE plug-ins, we found that the software design as well as past failure history, can be used to build models which accurately predict failure-prone components in new programs. Our prediction only requires usage relationships between components, which are typically defined in the design phase; thus, designers can easily explore and assess design alternatives in terms of predicted quality. In the ECLIPSE study, 90% of the 5% most failure-prone components, as predicted by our model from design data, turned out to actually produce failures later; a random guess would have predicted only 33%.",https://dl.acm.org/doi/10.1145/1159733.1159739,True,,191.0,['naive-bayes'],['code-metrics'],['novel-use'],['failure-prevention'],['source-code'],ISESE '06: Proceedings of the 2006 ACM/IEEE international symposium on Empirical software engineering,True,['failure-management'],,,['software-defect-prediction'],,,10.1145/1159733.1159739,,,
813,Resource Allocation for Antivirus Cloud Appliances,"['Hamzah, SK Ali Abdullah', 'Khattab, Sherif', 'El-Gamal, Salwa S']",2013,"Malware detection or antivirus software has been recently provided as a service in the cloud. A cloud antivirus provider hosts a number of virtual machines each running the same or different antivirus engines on potentially different sets of workloads (files). From the provider's perspective, the problem of optimally allocating physical resources to these virtual machines is crucial to the efficiency of the infrastructure. This paper proposes a search-based optimization approach for solving the resource allocation problem in cloud-based …",https://www.researchgate.net/publication/314400266_Resource_Allocation_for_Antivirus_Cloud_Appliances,True,,1.0,,,,,,,True,['resource-provisioning'],,,,,,,,,
814,Predicting fault incidence using software change history,"['T.L. Graves', 'A.F. Karr', 'J.S. Marron', 'H. Siy']",2000,"This paper is an attempt to understand the processes by which software ages. We define code to be aged or decayed if its structure makes it unnecessarily difficult to understand or change and we measure the extent of decay by counting the number of faults in code in a period of time. Using change management data from a very large, long-lived software system, we explore the extent to which measurements from the change history are successful in predicting the distribution over modules of these incidences of faults. In general, process measures based on the change history are more useful in predicting fault rates than product metrics of the code: For instance, the number of times code has been changed is a better indication of how many faults it will contain than is its length. We also compare the fault rates of code of various ages, finding that if a module is, on the average, a year older than an otherwise similar module, the older module will have roughly a third fewer faults. Our most successful model measures the fault potential of a module as the sum of contributions from all of the times the module has been changed, with large, recent changes receiving the most weight.",https://ieeexplore.ieee.org/document/859533,True,,774.0,['linear-regression'],"['code-metrics', 'code-history']",['discussion'],['failure-prevention'],['source-code'],,True,['failure-management'],True,,['software-defect-prediction'],True,6.0,,,,
815,A bio-inspired approach to provisioning of virtual resources in federated clouds,"['Lucio Agostinho', 'Guilherme Feliciano', 'Leonardo Olivi', 'Eleri Cardozo', 'Eliane Guimaraes']",2011,"In cloud computing the allocation and scheduling of multiple virtual resources, such as virtual machines (VMs), are still a challenge. The optimization of these processes bring the advantage of improving the energy savings and load balancing in large data centers. Resource allocation and scheduling also impact in federated clouds where resources can be leased from partner domains. This paper proposes a bio-inspired VM allocation method based on Genetic Algorithms to optimize the VM distribution across federated cloud domains. The main contribution of this work is an inter-domain allocation algorithm that takes into account the capacity of the links connecting the domains in order to avoid quality of service degradation for VMs allocated on partner domains. An architecture to simulate federated clouds is also a contribution of this paper.",https://ieeexplore.ieee.org/document/6119059,True,,43.0,,,,"['resource-consolidation', 'scheduling']",,"2011 IEEE Ninth International Conference on Dependable, Autonomic and Secure Computing",True,['resource-provisioning'],,,,,,,,,
816,Dynamic voltage frequency scaling for multi-tasking systems using online learning,"['Dhiman, Gaurav', 'Rosing, Tajana Simunic']",2007,This paper presents an extremely lightweight dynamic voltage and frequency scaling technique targeted towards modern multi-tasking systems. The technique utilizes processors runtime statistics and an online learning algorithm to estimate the best suited voltage and frequency setting at any given point in time. We implemented the proposed technique in Linux 2.6. 9 running on an Intel PXA27x platform and performed experiments in both single and multi-task environments. Our measurements show that we can achieve the maximum …,https://ieeexplore.ieee.org/document/5514319,True,,171.0,,,,['power-management'],,Proceedings of the 2007 international symposium on Low power electronics and design (ISLPED'07),True,['resource-provisioning'],,,,,,,,,
817,Workload prediction for adaptive power scaling using deep learning,"['Stephen J. Tarsa', 'Amit P. Kumar', 'H. T. Kung']",2014,"We apply hierarchical sparse coding, a form of deep learning, to model user-driven workloads based on on-chip hardware performance counters. We then predict periods of low instruction throughput, during which frequency and voltage can be scaled to reclaim power. Using a multi-layer coding structure, our method progressively codes counter values in terms of a few prominent features learned from data, and passes them to a Support Vector Machine (SVM) classifier where they act as signatures for predicting future workload states. We show that prediction accuracy and look-ahead range improve significantly over linear regression modeling, giving more time to adjust power management settings. Our method relies on learning and feature extraction algorithms that can discover and exploit hidden statistical invariances specific to workloads. We argue that, in addition to achieving superior prediction performance, our method is fast enough for practical use. To our knowledge, we are the first to use deep learning at the instruction level for workload prediction and on-chip power adaptation.",https://ieeexplore.ieee.org/document/6838580,True,,5.0,,,,,,2014 IEEE International Conference on IC Design \& Technology,True,['resource-provisioning'],,,,,,,,,
818,Software defect prediction using bayesian networks,"['Okutan, Ahmet', 'Y{\\i}ld{\\i}z, Olcay Taner']",2014,"There are lots of different software metrics discovered and used for defect prediction in the literature. Instead of dealing with so many metrics, it would be practical and easy if we could determine the set of metrics that are most important and focus on them more to predict defectiveness. We use Bayesian networks to determine the probabilistic influential relationships among software metrics and defect proneness. In addition to the metrics used in Promise data repository, we define two more metrics, ie NOD for the number of …",https://link.springer.com/article/10.1007/s10664-012-9218-8,True,,221.0,['bayesian-network'],['code-metrics'],['novel-use'],['failure-prevention'],['source-code'],,True,['failure-management'],True,,['software-defect-prediction'],True,5.0,,,,
819,A defect prediction model for open source software,"['Malhotra, Ruchika', 'Singh, Y']",2012,"Defect prediction models are significantly beneficial for software systems, where testing experts need to focus their attention and resources on problematic areas in the software under development. In this paper we find the relation between object oriented metrics and fault proneness using logistic regression method. The results are analyzed using open source software. The performance of the predicted models is evaluated using Receiver Operating Characteristic (ROC) analysis. The results show that Area under Curve …",https://www.researchgate.net/publication/265289928_A_Defect_Prediction_Model_for_Open_Source_Software,True,,8.0,,,,['failure-prediction'],,Proceedings of the World Congress on Engineering,True,['failure-management'],,,,,,,,,
820,Correlation Modeling and Resource Optimization for Cloud Service With Fault Recovery,"['Xiwei Qiu', 'Yuanshun Dai', 'Yanping Xiang', 'Liudong Xing']",2017,"Energy-efficient cloud computing has recently attracted much attention, where not only performance but also energy consumption are important metrics to be considered for designing rational resource scheduling strategies. Most of existing approaches for achieving energy efficient computing focus on connecting these two metrics and balancing the tradeoff between them, which however is inadequate because another important factor reliability is not considered. In fact, both virtual machine (VM) failures and server failures inevitably interrupt execution of a cloud service, and eventually result in spending more time and consuming more energy on completing the cloud service. Therefore, reliability significantly affects service performance and energy consumption, and thus they should not be handled separately. Connecting these correlated metrics is essential for making more precise evaluation and further for developing rational cloud resource scheduling strategies. In this paper, we present a correlated modeling approach applying Semi-Markov models, the Laplace-Stieltjes transform (LST), a Bayesian approach to analyze reliability-performance (R-P) and reliability-energy (R-E) correlations for cloud services using a retrying fault recovery mechanism. A recursive method is also proposed for modeling the correlations for cloud services using a check-pointing fault recovery mechanism. The proposed correlation models can be used to calculate the expected service time and energy consumption for completing a cloud service. Moreover, the models can contribute to analyzing the expected performance-energy tradeoff. We formulate the expected performance-energy optimization problem by describing performance and energy consumption metrics as functions of assigned CPU frequencies. Finally, we use a derivation approach to determine Pareto optimal solutions for the formulated optimization problem. Illustrative examples are provided.",https://ieeexplore.ieee.org/document/7892880,True,,10.0,,,,,,,True,['resource-provisioning'],,,,,,,,,
821,RVLBPNN: A Workload Forecasting Model for Smart Cloud Computing,"['Lu, Yao', 'Panneerselvam, John', 'Liu, Lu', 'Wu, Yan']",2016,"Given the increasing deployments of Cloud datacentres and the excessive usage of server resources, their associated energy and environmental implications are also increasing at an alarming rate. Cloud service providers are under immense pressure to significantly reduce both such implications for promoting green computing. Maintaining the desired level of Quality of Service (QoS) without violating the Service Level Agreement (SLA), whilst attempting to reduce the usage of the datacentre resources is an obvious challenge for the Cloud service providers. Scaling the level of active server resources in accordance with the predicted incoming workloads is one possible way of reducing the undesirable energy consumption of the active resources without affecting the performance quality. To this end, this paper analyzes the dynamic characteristics of the Cloud workloads and defines a hierarchy for the latency sensitivity levels of the Cloud workloads. Further, a novel workload prediction model for energy efficient Cloud Computing is proposed, named RVLBPNN (Rand Variable Learning Rate Backpropagation Neural Network) based on BPNN (Backpropagation Neural Network) algorithm. Experiments evaluating the prediction accuracy of the proposed prediction model demonstrate that RVLBPNN achieves an improved prediction accuracy compared to the HMM and Naïve Bayes Classifier models by a considerable margin",https://www.hindawi.com/journals/sp/2016/5635673/,True,,19.0,['rnn'],,"['comparison', 'novel-use']",['workload-prediction'],,,True,['resource-provisioning'],,,,,,,,,
822,Projecting Disk Usage Based on Historical Trends in a Cloud Environment,"['Stokely, Murray', 'Mehrabian, Amaan', 'Albrecht, Christoph', 'Labelle, Francois', 'Merchant, Arif']",2012,"Provisioning scarce resources among competing users and jobs remains one of the primary challenges of operating large-scale, distributed computing environments. Distributed storage systems, in particular, typically rely on hard operator-set quotas to control disk allocation and enforce isolation for space and I/O bandwidth among disparate users. However, users and operators are very poor at predicting future requirements and, as a result, tend to over-provision grossly.",https://research.google/pubs/pub37747/,True,,28.0,,,,['workload-prediction'],,,True,['resource-provisioning'],,,,,,,,,
823,Resource Central: Understanding and Predicting Workloads for Improved Resource Management in Large Cloud Platforms,"['Eli Cortez', 'Anand Bonde', 'Alexandre Muzio', 'Mark Russinovich', 'Marcus Fontoura', 'Ricardo Bianchini']",2017,"Cloud research to date has lacked data on the characteristics of the production virtual machine (VM) workloads of large cloud providers. A thorough understanding of these characteristics can inform the providers' resource management systems, e.g. VM scheduler, power manager, server health manager. In this paper, we first introduce an extensive characterization of Microsoft Azure's VM workload, including distributions of the VMs' lifetime, deployment size, and resource consumption. We then show that certain VM behaviors are fairly consistent over multiple lifetimes, i.e. history is an accurate predictor of future behavior. Based on this observation, we next introduce Resource Central (RC), a system that collects VM telemetry, learns these behaviors offline, and provides predictions online to various resource managers via a general client-side library. As an example of RC's online use, we modify Azure's VM scheduler to leverage predictions in oversubscribing servers (with oversubscribable VM types), while retaining high VM performance. Using real VM traces, we then show that the prediction-informed schedules increase utilization and prevent physical resource exhaustion. We conclude that providers can exploit their workloads' characteristics and machine learning to improve resource management substantially.",https://dl.acm.org/doi/10.1145/3132747.3132772,True,,69.0,,,,['workload-prediction'],,SOSP '17: Proceedings of the 26th Symposium on Operating Systems Principles,True,['resource-provisioning'],,,,,,10.1145/3132747.3132772,,,
824,Workload prediction for cloud computing elasticity mechanism,"['Yazhou Hu', 'Bo Deng', 'Fuyang Peng', 'Dongxia Wang']",2016,"Elasticity is the key feature of cloud computing technology, which can automatically reduce and add resources to meet users' need. In order to achieve elasticity, we should find how and when to trigger the elasticity automatic scaling mechanism. Workload analyzing is a popular method to solve this problem. In this paper, we propose three models to predict the workload based on analyzing monitoring data. Firstly, we use time series approach to analyze monitoring data. Then, we propose a Kalman filter model to predict the cloud workload. Next, we put forward a novel pattern matching model to analyze and predict the workload. Based on these predicting, we propose a new trigger strategy for cloud computing elasticity automatic scaling mechanism. Finally, experimental results show that our models not only improve the prediction accuracy, but also reduce the automatic scaling delay.",https://ieeexplore.ieee.org/document/7529565,True,,19.0,,,,['workload-prediction'],,2016 IEEE International Conference on Cloud Computing and Big Data Analysis (ICCCBDA),True,['resource-provisioning'],,,,,,,,,
825,Ensemble of Bayesian Predictors and Decision Trees for Proactive Failure Management in Cloud Computing Systems,"['Guan, Qiang', 'Zhang, Ziming', 'Fu, Song']",2012,"In modern cloud computing systems, hundreds and even thousands of cloud servers are interconnected by multi-layer networks. In such large-scale and complex systems, failures are common. Proactive failure management is a crucial technology to characterize system behaviors and forecast failure dynamics in the cloud. To make failure predictions, we need to monitor the system execution and collect health-related runtime performance data. However, in newly deployed or managed cloud systems, these data are usually unlabeled …",https://www.researchgate.net/publication/220520379_Ensemble_of_Bayesian_Predictors_and_Decision_Trees_for_Proactive_Failure_Management_in_Cloud_Computing_Systems,True,,70.0,"['decision-tree', 'entropy-selection', 'dimensionality-reduction', 'naive-bayes']","['host-metrics', 'network-metrics']",['new-method'],['failure-prediction'],['cloud'],,True,['failure-management'],True,,['system-failure-prediction'],,,10.4304/jcm.7.1.52-61,,,
826,Extended forecast of CPU and network load on computational Grid,"['S. Akioka', 'Y. Muraoka']",2004,"To achieve effective load balancing and a robust Grid environment, extended load forecast for computational resources is increasingly required. Thus, this paper proposes a method of predicting network and CPU load variance within a wide range, from several minutes to over a week. This is the widest range of prediction of the existing algorithms in the load of computational resources for the Grid environment. The distinctiveness of our algorithm is in using seasonal load variation for both load variance and one-step-ahead prediction. We apply seasonal fluctuation in CPU load to network load variation especially for network load variance prediction. Furthermore, the Markov model-based meta-predictor is used for one-step-ahead prediction, which is sensitive to late trends. The results of the experiments demonstrate that our algorithm gives a good curve for expected 8-day-long load variance, and makes accurate one-step-ahead predictions. The mean error rate for one-step-ahead predictions is 9.4% in the case of network load, and 6.2% in the case of CPU load. Moreover, the least mean error rate for wider range forecasts is 5.5% for network load variation, and 3.6% for CPU load variation.",https://ieeexplore.ieee.org/document/1336711,True,,66.0,,,,['workload-prediction'],,,True,['resource-provisioning'],,,,,,,,,
827,MJSA: Markov job scheduler based on availability in desktop grid computing environment,"['EunJoung Byun', 'SungJin Choi', 'MaengSoon Baik', 'JoonMin Gil', 'ChanYeol Park', 'ChongSun Hwang']",2007,"In a desktop grid computing environment, voluntary desktops (i.e., resource providers) are free to leave and join independently in the middle of execution. To develop a reliable desktop grid computing system, a scheduling scheme must consider the dynamic nature (i.e., volatility) of volunteers. Existing desktop grid computing systems, however, do not consider volatility in their scheduling procedures. As a result, job execution is often suspended, resulting in delayed completion time and degraded performance and reliability. To solve these limitations, we propose the Markov Job Scheduler based on Availability (MJSA) supporting three advanced scheduling schemes: OPTIMIST, PESSIMIST, and REALIST. These scheduling schemes are based on stochastic modeling of desktop availability. In the OPTIMIST scheme, in which time constraints are relaxed, the MJSA provides reliable resource selection at low cost. In the PESSIMIST scheme, where time constraints are rigid, the MJSA enables stable makespan in strictly time. Finally, in the REALIST scheme, where time constraints are only partially relaxed, the MJSA provides enhanced cost efficiency. In conclusion, the MJSA improves performance and reliability by adapting the appropriate scheduling scheme when selecting volunteers according to the needs of applications.",https://www.sciencedirect.com/science/article/abs/pii/S0167739X06001865?via%3Dihub,True,,31.0,,,,,,Future Generation Computer Systems,True,['resource-provisioning'],,,,,,10.1016/j.future.2006.09.004,,,
828,Deep Recurrent Model for Server Load and Performance Prediction in Data Center,"['Huang, Zheng', 'Peng, Jiajun', 'Lian, Huijuan', 'Guo, Jie', 'Qiu, Weidong']",2017,"Recurrent neural network (RNN) has been widely applied to many sequential tagging tasks such as natural language process (NLP) and time series analysis, and it has been proved that RNN works well in those areas. In this paper, we propose using RNN with long short-term memory (LSTM) units for server load and performance prediction. Classical methods for performance prediction focus on building relation between performance and time domain, which makes a lot of unrealistic hypotheses. Our model is built based on events (user …",https://www.hindawi.com/journals/complexity/2017/8584252/,True,,13.0,['rnn'],,['novel-use'],['workload-prediction'],,,True,['resource-provisioning'],,,,,,10.1155/2017/8584252,,,
829,Host load prediction with long short-term memory in cloud computing,"['Song, Binbin', 'Yu, Yao', 'Zhou, Yu', 'Wang, Ziqiang', 'Du, Sidan']",2018,"Host load prediction is significant for improving resource allocation and utilization in cloud computing. Due to the higher variance than that in a grid, accurate prediction remains a challenge in the cloud system. In this paper, we apply a concise yet adaptive and powerful model called long short-term memory to predict the mean load over consecutive future time intervals and actual load multi-step-ahead. Two real-world load traces were used to evaluate the performance. One is the load trace in the Google data center, and the other is that in a traditional distributed system. The experiment results show that our proposed method achieves state-of-the-art performance with higher accuracy in both datasets.",https://www.semanticscholar.org/paper/Host-load-prediction-with-long-short-term-memory-in-Song-Yu/f5905aa0e7eadc7f531ed504c9985b3de94efede,True,,38.0,['rnn'],,['novel-use'],['workload-prediction'],,,True,['resource-provisioning'],True,,,,,,,,
830,Cross-Project and Within-Project Semisupervised Software Defect Prediction: A Unified Approach,"['Fei Wu', 'Xiao-Yuan Jing', 'Ying Sun', 'Jing Sun', 'Lin Huang', 'Fangyi Cui', 'Yanfei Sun']",2018,"When there exist not enough historical defect data for building an accurate prediction model, semisupervised defect prediction (SSDP) and cross-project defect prediction (CPDP) are two feasible solutions. Existing CPDP methods assume that the available source data are …",https://ieeexplore.ieee.org/document/8320968,True,,31.0,"['optimization', 'dimensionality-reduction']",['code-metrics'],['new-method'],['failure-prevention'],['source-code'],,True,['failure-management'],,,['software-defect-prediction'],,,,,,
831,An approach for software defect prediction by combined soft computing,['T P Pushphavathi'],2017,"Nowadays, software project success is the key challenge. Prediction of software defects is main focus for the engineering community. A recent study in literature shows that data mining techniques are wildly used to predict software projects success. Many software development companies maintain their own software repositories, it is helpful for the prediction of software defects. There is dire need to reduce the gap between the software engineering and data mining community to increase the rate of software projects success. Although software defect prediction using classification/clustering algorithms has been encouraged by many researchers. However the lack of performance due to single classifier/clustering algorithms used for defect prediction. This paper provides software defect prediction using integrated approaches are advocated, instead of single classifier/clustering. The experimental results obtained shows better prediction performance could be achieved using the soft computing techniques (genetic algorithm, fuzzy c-means clustering and random forest classifier).",https://ieeexplore.ieee.org/document/8390007,True,,3.0,"['random-forest', 'genetic-programming', 'clustering']",,,['failure-prevention'],['source-code'],"2017 International Conference on Energy, Communication, Data Analytics and Soft Computing (ICECDS)",True,['failure-management'],,,['software-defect-prediction'],,,,,,
832,A load balancing scheme based on deep-learning in IoT,"['Kim, Hye-Young', 'Kim, Jong-Min']",2017,Extending the current Internet with interconnected objects and devices and their virtual representation has been a growing trend in recent years. The Internet of Things (IoT) contribution is in the increased value of information created by the number of interconnections among things and the transformation of the processed information into knowledge for the benefit of society. Benefit due to the service controlled by communication between objects is now being increased by people who use these services in real life. The …,https://link.springer.com/article/10.1007/s10586-016-0667-5,True,,18.0,,,,['failure-prediction'],,,True,['failure-management'],,,,,,,,,
833,Approach of measuring and predicting software system state based on hidden Markov model,"['Wu, Jia', 'Zeng, Wei-Ru', 'Chen, Han-Lin', 'TANG, Xue-Fei']",2016,"With increased improvement on the capability and performance of software systems, enterprises have improved the management efficiency and enhanced the business model. Meanwhile, as software systems become more and more complex, severe challenges for the management of software systems are encountered. This paper presents a method for measuring and predicting software system state based on hidden Markov model. It establishes the linkage between the exterior state (the observation state) and the interior …",https://www.researchgate.net/publication/312060680_Approach_of_measuring_and_predicting_software_system_state_based_on_hidden_Markov_model,True,,6.0,"['clustering', 'markov-model']",,['new-method'],['failure-prediction'],['software'],,True,['failure-management'],,,['system-failure-prediction'],,,10.13328/j.cnki.jos.005014,,,
834,Joint Optimization of Server and Network Resource Utilization in Cloud Data Centers,"['Biyu Zhou', 'Jie Wu', 'Lin Wang', 'Fa Zhang', 'Zhiyong Liu']",2017,"Virtual machine placement is a key component of cloud resource management, which may affect network bandwidth allocation. In this paper, we revisit the virtual machine placement problem in cloud data centers and aim to maximize the overall resource utilization in multiple dimensions, while ensuring that the resource constraints on both the server such as CPU capacity and the network such as bandwidth are not violated. We model the bandwidth-guaranteed virtual machine placement problem and prove its NP-hardness, and design offline and online algorithms to solve the problem. We first consider the offline version and develop approximation algorithms with bounded performance ratios for both the homogeneous and the heterogeneous cases. Then, for the online version, we propose simple and efficient heuristics based on the insights from the offline algorithm design. Comprehensive experimental results verify that the overall resource utilization can be significantly improved by applying our proposals.",https://ieeexplore.ieee.org/document/8254037,True,,2.0,,,,,,GLOBECOM 2017-2017 IEEE Global Communications Conference,True,['resource-provisioning'],,,,,,,,,
835,Dynamic VM Allocation Algorithm using Clustering in Cloud Computing,"['Panchal, Bhupendra', 'Kapoor, RK']",2013,,https://www.semanticscholar.org/paper/Dynamic-VM-Allocation-Algorithm-using-Clustering-in-Panchal-Kapoor/e343ef29a3c9a32c91a29e6b7cfee8638624e137,True,,27.0,,,,['resource-consolidation'],,,True,['resource-provisioning'],,,,,,,,,
836,Power Cost Reduction in Distributed Data Centers: A Two-Time-Scale Approach for Delay Tolerant Workloads,"['Yuan Yao', 'Longbo Huang', 'Abhishek B. Sharma', 'Leana Golubchik', 'Michael J. Neely']",2012,"This paper considers a stochastic optimization approach for job scheduling and server management in large-scale, geographically distributed data centers. Randomly arriving jobs are routed to a choice of servers. The number of active servers depends on server activation decisions that are updated at a slow time scale, and the service rates of the servers are controlled by power scaling decisions that are made at a faster time scale. We develop a two-time-scale decision strategy that offers provable power cost and delay guarantees. The performance and robustness of the approach is illustrated through simulations.",https://ieeexplore.ieee.org/document/6392828,True,,97.0,,,,['power-management'],,,True,['resource-provisioning'],,,,,,,,,
837,Power-Aware Linear Programming based Scheduling for heterogeneous computer clusters,"['Hadil Al-Daoud', 'Issam Al-Azzoni', 'Douglas G. Down']",2010,"Several power-aware scheduling policies have been proposed for homogeneous clusters. In this work, we propose a new policy for heterogeneous clusters. Our simulation experiments show that using our proposed policy results in significant reduction in energy consumption while performing very competitively in heterogeneous clusters.",https://ieeexplore.ieee.org/document/5598298?reload=true&arnumber=5598298,True,,37.0,,,,"['scheduling', 'power-management']",,,True,['resource-provisioning'],,,,,,,,,
838,A New Genetic Algorithm for Scheduling for Large Communication Delays,"['Johnatan E. Pecero', 'Denis Trystram', 'Albert Y. Zomaya']",2009,"In modern parallel and distributed systems, the time for exchanging data is usually larger than that for computing elementary operations. Consequently, these communications slow down the execution of the application scheduled on such systems. Accounting for these communications is essential for attaining efficient hardware and software utilization. Therefore, we provide in this paper a new combined approach for scheduling parallel applications with large communication delays on an arbitrary number of processors. In this approach, a genetic algorithm is improved with the introduction of some extra knowledge about the scheduling problem. This knowledge is represented by a class of clustering algorithms introduced recently, namely, convex clusters which are based on structural properties of the parallel applications. The developed algorithm is assessed by simulations run on some families of synthetic task graphs and randomly generated applications. The comparison with related approaches emphasizes its interest.",https://dl.acm.org/doi/10.1007/978-3-642-03869-3_25,True,,15.0,,,,,,Euro-Par '09: Proceedings of the 15th International Euro-Par Conference on Parallel Processing,True,['resource-provisioning'],,,,,,10.1007/978-3-642-03869-3_25,,,
839,A genetic algorithm for multiprocessor scheduling,"['E.S.H. Hou', 'N. Ansari', 'Hong Ren']",1994,"The problem of multiprocessor scheduling can be stated as finding a schedule for a general task graph to be executed on a multiprocessor system so that the schedule length can be minimized. This scheduling problem is known to be NP-hard, and methods based on heuristic search have been proposed to obtain optimal and suboptimal solutions. Genetic algorithms have recently received much attention as a class of robust stochastic search algorithms for various optimization problems. In this paper, an efficient method based on genetic algorithms is developed to solve the multiprocessor scheduling problem. The representation of the search node is based on the order of the tasks being executed in each individual processor. The genetic operator proposed is based on the precedence relations between the tasks in the task graph. Simulation results comparing the proposed genetic algorithm, the list scheduling algorithm, and the optimal schedule using random task graphs, and a robot inverse dynamics computational task graph are presented.<
>",https://ieeexplore.ieee.org/document/265940,True,,940.0,['genetic-programming'],['tasks'],,['scheduling'],['scheduler'],,True,['resource-provisioning'],,,,,,,,,['multiprocessor']
840,Heuristics and augmented neural networks for task scheduling with non-identical machines,"['Agarwal, Anurag', 'Colak, Selcuk', 'Jacob, Varghese S', 'Pirkul, Hasan']",2006,"We propose new heuristics along with an augmented-neural-network (AugNN) formulation for solving the makespan minimization task-scheduling problem for the non-identical machine environment. We explore four task and three machine-priority rules, resulting in 12 combinations of single-pass heuristics. The task-priority rules are Highest-Level-First (HLF), Highest-Total-Remaining-Processing-Time-First (HTRPTF), Smallest-Latest-Finish-Time-First (SLFTF) and Minimum-Slack-First (MSF). For machine priority, we propose a greedy …",https://www.sciencedirect.com/science/article/abs/pii/S0377221705004303?via%3Dihub,True,,31.0,,,,,,,True,['resource-provisioning'],,,,,,,,,
841,"Fault-aware, utility-based job scheduling on Blue, Gene/P systems","['Wei Tang', 'Zhiling Lan', 'Narayan Desai', 'Daniel Buettner']",2009,"Job scheduling on large-scale systems is an increasingly complicated affair, with numerous factors influencing scheduling policy. Addressing these concerns results in sophisticated scheduling policies that can be difficult to reason about. In this paper, we present a general utility-based scheduling framework to balance various scheduling requirements and priorities. It enables system owners to customize scheduling policies under different circumstances without changing the scheduling code. We also develop a fault-aware job allocation strategy for Blue Gene/P systems to address the increasing concern of system failures. We demonstrate the effectiveness of these facilities by means of event-driven simulations with real job traces collected from the production Blue Gene/P system at Argonne National Laboratory.",https://ieeexplore.ieee.org/document/5289206,True,,62.0,,,,['scheduling'],,2009 IEEE International Conference on Cluster Computing and Workshops,True,['resource-provisioning'],,,,,,,,,
842,Novel selection policies for Container-based Cloud Deployment Models,"['Walid A. Hanafy', 'Amr E. Mohamed', 'Sameh A. Salem']",2017,"Application deployment models have been introduced over the last decades. These models are characterized by low resource utilization efficiencies which are improved by the appearance of virtualization technologies. Currently, a more efficient utilization model called containers type virtualization has appeared. This allows different applications to share the same operating system kernel resulting in a superior application density. However, the developed infrastructures and frameworks contribute to the energy consumption and violation rates of the service level agreement. In this paper, a novel container and host selection policies for cloud deployment models are proposed. Experimental results show that the proposed policies achieve better energy and services level agreement commitment compared to the other competitors. This paper, demonstrate the existence of an association between container and host selection policies.",https://ieeexplore.ieee.org/document/8289794,True,,4.0,,,,,,2017 13th International Computer Engineering Conference (ICENCO),True,['resource-provisioning'],,,,,,,,,
843,Power-Aware Consolidation of Scientific Workflows in Virtualized Environments,"['Qian Zhu', 'Jiedan Zhu', 'Gagan Agrawal']",2010,"The recent emergence of clouds with large, virtualized pools of compute and storage resources raises the possibility of a new compute paradigm for scientific research. With virtualization technologies, consolidation of scientific workflows presents a promising opportunity for energy and resource cost optimization, while achieving high performance. We have developed pSciMapper, a power-aware consolidation framework for scientific workflows. We view consolidation as a hierarchical clustering problem, and introduce a distance metric that is based on interference between resource requirements. A dimensionality reduction method (KCCA) is used to relate the resource requirements to performance and power consumption. We have evaluated pSciMapper with both real-world and synthetic scientific workflows, and demonstrated that it is able to reduce power consumption by up to 56%, with less than 15% slowdown. Our experiments also show that scheduling overheads of pSciMapper are low, and the algorithm can scale well for workflows with hundreds of tasks.",https://ieeexplore.ieee.org/abstract/document/5644899,True,,118.0,"['dimensionality-reduction', 'clustering']",['tasks'],['novel-use'],"['scheduling', 'power-management']","['vm', 'cloud']","SC'10: Proceedings of the 2010 ACM/IEEE International Conference for High Performance Computing, Networking, Storage and Analysis",True,['resource-provisioning'],,,,,,,,,
844,Classifying server behavior and predicting impact of modernization actions,"['Jasmina Bogojeska', 'David Lanyi', 'Ioana Giurgiu', 'George Stark', 'Dorothea Wiesmann']",2013,"Today the decision of when to modernize which elements of the server HW/SW stack is often done manually based on simple business rules. In this paper we alleviate this problem by supporting the decision process with an automated approach based on incident tickets and server attributes data. As a first step we identify and rank servers with problematic behavior as candidates for modernization using a random forest classifier. Second, this predictive model is used to evaluate the impact of different modernization actions and suggest the most effective ones. We show that our chosen model yields high quality predictions and outperforms traditional linear regression models on a large set of real data.",https://ieeexplore.ieee.org/document/6727810,True,,28.0,,,,['failure-detection'],,Proceedings of the 9th International Conference on Network and Service Management (CNSM 2013),True,['failure-management'],,,,,,,,,
845,Black box scheduling for resource intensive virtual machine workloads with interference models,"['Sam Verboven', 'Kurt Vanmechelen', 'Jan Broeckhove']",2013,"Modern datacenters consist of increasingly powerful hardware. Achieving high levels of utilization on this hardware often requires the execution of multiple concurrent workloads. Virtualization has emerged as an efficient means to isolate workloads by partitioning large physical resources using self-contained virtual machine images. Despite the many advantages, some challenges regarding performance isolation still need to be addressed. Unmanaged multiplexing of resource intensive workloads has the potential to cause unexpected variances in workload performance. In this paper, we address this issue using performance models based on the runtime characteristics of virtualized workloads. A set of resource intensive workloads is benchmarked with increasing degrees of multiplexing. Resource usage profiles are constructed using the metrics made available by the Xen hypervisor. Based on these profiles, performance degradation is predicted using several existing modeling techniques. In addition, we propose a novel approach using both the classification and regression capabilities of support vector machines. Application clustering is used to identify several application types with distinct performance profiles. Finally, we evaluate the developed performance models by introducing several new scheduling techniques. We demonstrate that the integration of these models in the scheduling logic can significantly improve the overall performance of multiplexed workloads.",https://dl.acm.org/doi/10.1016/j.future.2013.04.027,True,,23.0,,,,,,Future Generation Computer Systems,True,['resource-provisioning'],,,,,,10.1016/j.future.2013.04.027,,,
846,Deep convolutional neural networks for detecting noisy neighbours in cloud infrastructure,"['Ordozgoiti, Bruno', 'Mozo, Alberto', ""Canaval, Sandra G{\\'o}mez"", 'Margolin, Udi', 'Rosensweig, Elisha', 'Segall, Itai']",2017,"Cloud infrastructure in data centers is expected to be one of the main technologies supporting Internet communications in the coming years. Virtualization is employed to achieve the flexibility and dynamicity required by the wide variety of applications used today. Therefore, optimal allocation of virtual machines is key to ensuring performance and efficiency. Noisy neighbor is a term used to describe virtual machines competing for physical resources and thus disturbing each other, a phenomenon that can dramatically degrade …",https://www.semanticscholar.org/paper/Deep-convolutional-neural-networks-for-detecting-in-Ordozgoiti-Mozo/72f900b8d89312a392255c4dc9544376e6ad9145,True,,3.0,,,,['failure-detection'],,,True,['failure-management'],,,,,,,,,
847,An unsupervised approach to online noisy-neighbor detection in cloud data centers,"['Lorido-Botran, Tania', 'Huerta, Sergio', ""Tom{\\'a}s, Luis"", 'Tordsson, Johan', 'Sanz, Borja']",2017,"Resource sharing is an inherent characteristic of cloud data centers. Virtual Machines (VMs) and/or Containers that are co-located in the same physical server often compete for resources leading to interference. The noisy neighbor's effect refers to an anomaly caused by a VM/container limiting resources accessed by another one. Our main contribution is an online, lightweight and application-agnostic solution for anomaly detection, that follows an unsupervised approach. It is based on comparing models for different lags: Dirichlet Process …",https://www.sciencedirect.com/science/article/abs/pii/S0957417417305158?via%3Dihub,True,,6.0,,,,['failure-detection'],,,True,['failure-management'],,,,,,,,,
848,Multimedia QoE optimized management using prediction and statistical learning,"['Muslim Elkotob', 'Daniel Grandlund', 'Karl Andersson', 'Christer Ahlund']",2010,"We present a scheme for flow management with heterogeneous access technologies available indoors and in a campus network such as GPRS, 3G and Wi-Fi. Statistical learning is used as a key for optimizing a target variable namely video quality of experience (QoE). First we analyze the data using passive measurements to determine relationships between parameters and their impact on the main performance indicator, video Quality of Experience (QoE). The derived weights are used for performing prediction in every discrete time interval of our designed autonomic control loop to know approximately the QoE in the next time interval and perform a switch to another access technology if it yields a better QoE level. This user-perspective performance optimization is in line with operator and service provider goals. QoE performance models for slow vehicular and pedestrian speeds for Wi-Fi and 3G are derived and compared.",https://ieeexplore.ieee.org/document/5735733,True,,24.0,,,,,,IEEE Local Computer Network Conference,True,['resource-provisioning'],,,,,,,,,
849,Predicting quality of experience in multimedia streaming,"['Vlado Menkovski', 'Adetola Oredope', 'Antonio Liotta', 'Antonio Cuadra Sánchez']",2009,"Measuring and predicting the user's Quality of Experience (QoE) of a multimedia stream is the first step towards improving and optimizing the provision of mobile streaming services. This enables us to better understand how Quality of Service (QoS) parameters affect service quality, as it is actually perceived by the end user. Over the last years this goal has been pursued by means of subjective tests and through the analysis of the user's feedback. Existing statistical techniques have lead to poor accuracy (order of 70%) and inability to evolve prediction models with the system's dynamics. In this paper, we propose a novel approach for building accurate and adaptive QoE prediction models using Machine Learning classification algorithms, trained on subjective test data. These models can be used for real-time prediction of QoE and can be efficiently integrated into online learning systems that can adapt the models according to changes in the environment. Providing high accuracy of above 90%, the classification algorithms become an indispensible component of a mobile multimedia QoE management system.",https://dl.acm.org/doi/10.1145/1821748.1821766,True,,52.0,,,,['workload-prediction'],,MoMM '09: Proceedings of the 7th International Conference on Advances in Mobile Computing and Multimedia,True,['resource-provisioning'],,,,,,10.1145/1821748.1821766,,,
850,Research on relationship between QoE and QoS based on BP Neural Network,"['Haiqing Du', 'Chang Guo', 'Yixi Liu', 'Yong Liu']",2009,"In this paper, a method to connect QoE directly to QoS according to the corresponding level of QoE under different network circumstance is proposed. The algorithm used to analysis a number of data is the BP (back-propagation) neural network toolbox in Matlab software. With the help of video quality evaluation system and network emulator, all kinds of parameters of different network circumstance and the targets based on our dividing standard were used to train the neural network to get the right weights in BP neural network. Based on the trained BP neural network, the input network parameters can be adjusted to get the ideal output to satisfy the users' need.",https://ieeexplore.ieee.org/document/5360833,True,,45.0,['multilayer-perceptron'],,,['workload-prediction'],,2009 IEEE International Conference on Network Infrastructure and Digital Content,True,['resource-provisioning'],,,,,,,,,
851,A large-scale web QoS prediction scheme for the Industrial Internet of Things based on a kernel machine learning algorithm,"['Xiong Luo', 'Ji Liu', 'Dandan Zhang', 'Xiaohui Chang']",2016,"Cloud computing plays an essential role in enabling practical applications based on the Industrial Internet of Things (IIoT). Hence, the quality of these services directly impacts the usability of IIoT applications. To select or recommend the best web and cloud based services, one method is to mine the vast data that are pertinent to the quality of service (QoS) of such services. To enable dynamic discovery and composition of web services, one can use a set of well-defined QoS criteria to describe and distinguish functionally similar web services. In general, QoS is a nonfunctional performance index of web services, and it might be user-dependent. Hence, to fully assess the QoS of all available web services, a user normally would have to invoke every one of them. This implies that the QoS values for services that the user has not invoked would be missing. If the number of web services available is large, it is virtually inevitable for this to happen because invoking every single service would be prohibitively expensive. This issue is typically resolved by employing some predication algorithms to estimate the missing QoS values. In this paper, a data-driven scheme of predicting the missing QoS values for the IIoT based on a kernel least mean square algorithm (KLMS) is proposed. During the data prediction process, the Pearson correlation coefficient (PCC) is initially introduced to find the relevant QoS values from similar service users and web service items for each known QoS entry. Next, KLMS is used to analyze the hidden relationships between all the known QoS data and corresponding QoS data with the highest similarities. We therefore can apply the derived coefficients for the prediction of missing web service QoS values. An extensive performance study based on a public data set is conducted to verify the prediction accuracy of our proposed scheme. This data set includes 200 distributed service users on 500 web service items with a total of 1,858,260 intermediate data values. The experiment results show that our proposed KLMS-based prediction scheme has better prediction accuracy than traditional approaches.",https://dl.acm.org/doi/10.1016/j.comnet.2016.01.004,True,,44.0,,,,['workload-prediction'],,Computer Networks: The International Journal of Computer and Telecommunications Networking,True,['resource-provisioning'],,,,,,10.1016/j.comnet.2016.01.004,,,
852,Job Scheduling for Cloud Computing Using Neural Networks,"['Maqableh, Mahmoud', 'Karajeh, Huda', 'others']",2014,"Cloud computing aims to maximize the benefit of distributed resources and aggregate them to achieve higher throughput to solve large scale computation problems. In this technology, the customers rent the resources and only pay per use. Job scheduling is one of the biggest issues in cloud computing. Scheduling of users' requests means how to allocate resources to these requests to finish the tasks in minimum time. The main task of job scheduling system is to find the best resources for user's jobs, taking into consideration some statistics and …",https://www.scirp.org/journal/paperinformation.aspx?paperid=49261,True,,59.0,,,,['scheduling'],,,True,['resource-provisioning'],,,,,,,,,
853,Thoth: Automatic Resource Management with Machine Learning for Container-based Cloud Platform,"['Akkarit Sangpetch', 'Orathai Sangpetch', 'Nut Juangmarisakul', 'Supakorn Warodom']",2017,"Platform-as-a-Service (PaaS) providers often encounter fluctuation in computing resource usage due to workload changes, resulting in performance degradation. To maintain acceptable service quality, providers may need to manually adjust resource allocation according to workload dynamics. Unfortunately, this approach will not scale well as the number of applications grows. We thus propose Thoth, a dynamic resource management system for PaaS using Docker container technology. Thoth automatically monitors resource usage and dynamically adjusts appropriate amount of resources for each application. To implement the automatic-scaling algorithm, we select three algorithms, namely Neural Network, Q-Learning and our rule-based algorithm, to study and evaluate. The experimental results suggest that Q-Learning can the best adapt to the load changes, followed by a rule-based algorithm and NN. With Q-Learning, Thoth can save computing resources by 28.95% and 21.92%, compared to Neural Network and the rule-based algorithm respectively, without compromising service quality.",https://dl.acm.org/doi/10.5220/0006254601030111,True,,7.0,,,,,,CLOSER 2017: Proceedings of the 7th International Conference on Cloud Computing and Services Science,True,['resource-provisioning'],,,,,,10.5220/0006254601030111,,,
854,Trace-Based Analysis and Prediction of Cloud Computing User Behavior Using the Fractal Modeling Technique,"['Shuang Chen', 'Mahboobeh Ghorbani', 'Yanzhi Wang', 'Paul Bogdan', 'Massoud Pedram']",2014,"The problem of big data analytics is gaining increasing research interest because of the rapid growth in the volume of data to be analyzed in various areas of science and technology. In this paper, we investigate the characteristics of the cloud computing requests received by the cloud infrastructure operators. The cluster usage dataset released by Google is thoroughly studied. To address the self-similarity and non-stationarity characteristics of the workload profile in a cloud computing system, fractal modeling techniques similar to some cyber-physical system (CPS) applications are exploited. A trace-based prediction of the job inter-arrival time and aggregated resource request sent to server cluster in the near future is effectively performed by solving fractional-order differential equations. The distributions of important parameters including job/task duration time and resource request per task in terms of CPU, memory, and storage are extracted from the cluster dataset are fitted using the alpha-stable distribution.",https://ieeexplore.ieee.org/document/6906851,True,,12.0,,,,,,2014 IEEE International Congress on Big Data,True,['resource-provisioning'],,,,,,,,,
855,Predicting host CPU utilization in the cloud using evolutionary neural networks,"['Mason, Karl', 'Duggan, Martin', 'Barrett, Enda', 'Duggan, Jim', 'Howley, Enda']",2018,"The Infrastructure as a Service (IaaS) platform in cloud computing provides resources as a service from a pool of compute, network, and storage resources. One of the major challenges facing cloud computing is to predict the usage of these resources in real time. By knowing future demands, cloud data centres can dynamically scale resources to decrease energy consumption while maintaining a high quality of service. However cloud resource consumption is ever changing, making it difficult for accurate predictions to be produced. This motivates the research presented in this paper which aims to predict in advance the level of CPU consumption of a host. This research implements evolutionary Neural Networks (NN), a powerful machine learning method, to make these predictions. A number of state of the art swarm and evolutionary optimization algorithms are implemented to train the neural networks to predict host utilization: Particle Swarm Optimization (PSO), Differential Evolution (DE) and Covariance Matrix Adaptation Evolutionary Strategy (CMA-ES). The results of this research demonstrate that CMA-ES converges faster to a better solution on the training data. However when evaluated on the test data, DE performs statistically equal to CMA-ES. The results also demonstrate that the trained networks are still accurate when applied to CPU utilization data from different hosts with no further training needed. When evaluated to predict multiple steps into the future, the accuracy of the network understandably decreases but still performs well on average.",https://www.sciencedirect.com/science/article/abs/pii/S0167739X17322793,True,,23.0,"['multilayer-perceptron', 'particle-swarm']",,['novel-use'],['workload-prediction'],,,True,['resource-provisioning'],,,,,,,,,
856,Workload prediction in cloud using artificial neural network and adaptive differential evolution,"['Jitendra Kumar', 'Ashutosh Kumar Singh']",2018,"Cloud computing has drastically transformed the means of computing in recent years. In spite of numerous benefits, it suffers from some challenges too. Major challenges of cloud computing include dynamic resource scaling and power consumption. These factors lead a cloud system to become inefficient and costly. The workload prediction is one of the variables by which the efficiency and operational cost of a cloud can be improved. Accuracy is the key component in workload prediction and the existing approaches lag in producing 100% accurate results. The researchers are also putting their consistent efforts for its improvement. In this paper, we present a workload prediction model using neural network and self adaptive differential evolution algorithm. The model is capable of learning the best suitable mutation strategy along with optimal crossover rate. The experiments were performed on the benchmark data sets of NASA and Saskatchewan servers HTTP traces for different prediction intervals. We compared the results with prediction model based on well known back propagation learning algorithm and received significant improvement. The proposed model attained a shift up to 168 times in the error reduction and prediction error is reduced up to 0.001. The paper presents a workload prediction approach for cloud datacenters using neural network and self adaptive differential evolution.The proposed approach outperforms well known back propagation network approach in accuracy.The root mean squared error is used as accuracy measurement metric and proposed approach is able to achieve significant reduction in the prediction error.",https://dl.acm.org/doi/10.1016/j.future.2017.10.047,True,,58.0,,,,['workload-prediction'],,Future Generation Computer Systems,True,['resource-provisioning'],,,,,,10.1016/j.future.2017.10.047,,,
857,A sequential pattern mining model for application workload prediction in cloud environment,"['Amiri, Maryam', 'Mohammad-Khanli, Leyli', 'Mirandola, Raffaela']",2018,"The resource provisioning is one of the challenging problems in the cloud environment. The resources should be allocated dynamically according to the demand changes of the applications. Over-provisioning increases energy wasting and costs. On the other hand, under-provisioning causes Service Level Agreements (SLA) violation and Quality of Service (QoS) dropping. Therefore the allocated resources should be close to the current demand of applications as much as possible. For this purpose, the future demand of applications should be determined. Thus, the prediction of the future workload of applications is an essential step before the resource provisioning. To the best of our knowledge, for the first time, this paper proposes a novel Prediction mOdel based on SequentIal paTtern mINinG (POSITING) that considers correlation between different resources and extracts behavioural patterns of applications independently of the fixed pattern length explicitly. Based on the extracted patterns and the recent behaviour of the application, the future demand of resources is predicted. The main goal of this paper is to show that models based on pattern mining could offer novel and useful points of view for tackling some of the issues involved in predicting the application workloads. The performance of the proposed model is evaluated based on both real and synthetic workloads. The experimental results show that the proposed model could improve the prediction accuracy in comparison to the other state-of-the-art methods such as moving average, linear regression, neural networks and hybrid prediction approaches.",https://www.sciencedirect.com/science/article/pii/S1084804517304198?via%3Dihub,True,,4.0,,,,,,,True,['resource-provisioning'],,,,,,,,,
858,An online learning model based on episode mining for workload prediction in cloud,"['Amiri, Maryam', 'Mohammad-Khanli, Leyli', 'Mirandola, Raffaela']",2018,"The resource provisioning is one of the challenging problems in the cloud environment. The resources should be allocated dynamically according to the demand changes of the applications. Over-provisioning increases energy wasting and costs. On the other hand, under-provisioning causes Service Level Agreements (SLA) violation and Quality of Service (QoS) dropping. Therefore the allocated resources should be close to the current demand of applications as much as possible. Thus, the prediction of the future workload of applications …",https://www.sciencedirect.com/science/article/abs/pii/S0167739X18300712?via%3Dihub,True,,9.0,,,,['workload-prediction'],,,True,['resource-provisioning'],,,,,,,,,
859,Sample-based software defect prediction with active and semi-supervised learning,"['Li, Ming', 'Zhang, Hongyu', 'Wu, Rongxin', 'Zhou, Zhi-Hua']",2012,"Software defect prediction can help us better understand and control software quality. Current defect prediction techniques are mainly based on a sufficient amount of historical project data. However, historical data is often not available for new projects and for many organizations. In this case, effective defect prediction is difficult to achieve. To address this problem, we propose sample-based methods for software defect prediction. For a large software system, we can select and test a small percentage of modules, and then build a …",https://link.springer.com/article/10.1007/s10515-011-0092-1,True,,141.0,"['random-forest', 'decision-tree', 'logistic-regression', 'naive-bayes']",['code-metrics'],['novel-use'],['failure-prevention'],['source-code'],,True,['failure-management'],False,,['software-defect-prediction'],,,,,,
860,Predicting application run times using historical information,"['Smith, Warren', 'Foster, Ian', 'Taylor, Valerie']",1998,"We present a technique for deriving predictions for the run times of parallel applications from the run times of “similar” applications that have executed in the past. The novel aspect of our work is the use of search techniques to determine those application characteristics that yield the best definition of similarity for the purpose of making predictions. We use four workloads recorded from parallel computers at Argonne National Laboratory, the Cornell Theory Center, and the San Diego Supercomputer Center to evaluate the effectiveness of our …",https://link.springer.com/chapter/10.1007/BFb0053984,True,,431.0,,,,['workload-prediction'],,Workshop on Job Scheduling Strategies for Parallel Processing,True,['resource-provisioning'],,,,,,,,,
861,PQR: Predicting Query Execution Times for Autonomous Workload Management,"['Chetan Gupta', 'Abhay Mehta', 'Umeshwar Dayal']",2008,"Modern enterprise data warehouses have complex workloads that are notoriously difficult to manage. One of the key pieces to managing workloads is an estimate of how long a query will take to execute. An accurate estimate of this query execution time is critical to self managing Enterprise Class Data Warehouses. In this paper we study the problem of predicting the execution time of a query on a loaded data warehouse with a dynamically changing workload. We use a machine learning approach that takes the query plan, combines it with the observed load vector of the system and uses the new vector to predict the execution time of the query. The predictions are made as time ranges. We validate our solution using real databases and real workloads. We show experimentally that our machine learning approach works well. This technology is slated for incorporation into a commercial, enterprise class DBMS.",https://ieeexplore.ieee.org/document/4550823?section=abstract,True,,27.0,,,,,,2008 International Conference on Autonomic Computing,True,['resource-provisioning'],,,,,,,,,
862,The Grid Backfilling: a Multi-Site Scheduling Architecture with Data Mining Prediction Techniques,"['Rodero, Ivan', 'Guim, Francesc', 'Corbalan, Julita', 'Goyeneche, Ariel']",2008,"In large Grids, like the National Grid Service (NGS), or large distributed architecture different scheduling entities are involved. Despite a global scheduling approach would archive higher performance and could increment the utilization of global system in these scenarios usually independent schedulers carry out its own scheduling decisions. In this paper we present howa coordinated scheduling among all the different centers using data mining prediction techniques can substantially improve the performance of the global distributed …",https://www.researchgate.net/publication/226189417_The_Grid_Backfilling_a_Multi-Site_Scheduling_Architecture_with_Data_Mining_Prediction_Techniques,True,,23.0,"['optimization', 'autoregression']",,['new-method'],"['scheduling', 'workload-prediction']",,Grid Middleware and Services,True,['resource-provisioning'],,,,,,,,,
863,A Hybrid Intelligent Method for Performance Modeling and Prediction of Workflow Activities in Grids,"['Rubing Duan', 'Farrukh Nadeem', 'Jie Wang', 'Yun Zhang', 'Radu Prodan', 'Thomas Fahringer']",2009,"Grid schedulers require individual activity performance predictions to map workflow activities on different Grid sites. The effectiveness of the scheduling systems is hampered by inaccurate predictions due to the inability of existing predictors to effectively model the dynamic and heterogeneous nature of Grid resources, or the wide range of problem sizes and runtime arguments. To address this deficiency, we propose a hybrid Bayesian-neural network approach to dynamically model and predict the execution time of activities in real workflow applications. Bayesian network is used for a high-level representation of activities performance probability distribution against different factors affecting the performance. The important attributes are dynamically selected by the Bayesian network and fed into a radial basis function neural network to make further predictions. Our approach is generic to any type of scientific applications, and flexible to import expert knowledge to further improve accuracies. Experimental results for activities from three realworld workflow applications are presented to show effectiveness of our approach.",https://ieeexplore.ieee.org/document/5071890,True,,33.0,,,,,,2009 9th IEEE/ACM International Symposium on Cluster Computing and the Grid,True,['resource-provisioning'],,,,,,,,,
864,DynaQoS: Model-free self-tuning fuzzy control of virtualized resources for QoS provisioning,"['Jia Rao', 'Yudi Wei', 'Jiayu Gong', 'Cheng-Zhong Xu']",2011,"Cloud elasticity allows dynamic resource provisioning in concert with actual application demands. Feedback control approaches have been applied with success to resource allocation in physical servers. However, cloud dynamics make the design of an accurate and stable resource controller more challenging, especially when response time is considered as the measured output. Response time is highly dependent on the characteristics of workload and sensitive to cloud dynamics. To address the challenges, we extend a self-tuning fuzzy control (STFC) approach, originally developed for response time assurance in web servers to resource allocation in virtualized environments. We introduce mechanisms for adaptive output amplification and flexible rule selection in the STFC approach for better adaptability and stability. Based on the STFC, we further design a two-layer QoS provisioning framework, DynaQoS, that supports adaptive multi-objective resource allocation and service differentiation. We implement a prototype of DynaQoS on a Xen-based cloud testbed. Experimental results on an E-Commerce benchmark show that STFC outperforms popular controllers such as Kalman filter, ARMA and adaptive PI by at least 16% and 37% under both static and dynamic workloads, respectively. Further results with multiple control objectives and service classes demonstrate the effectiveness of DynaQoS in performance-power control and service differentiation.",https://ieeexplore.ieee.org/document/5931341,True,,62.0,,,,['resource-consolidation'],,2011 IEEE Nineteenth IEEE International Workshop on Quality of Service,True,['resource-provisioning'],,,,,,,,,
865,IntroPerf: transparent context-sensitive multi-layer performance inference using system stack traces,"['Chung Hwan Kim', 'Junghwan Rhee', 'Hui Zhang', 'Nipun Arora', 'Guofei Jiang', 'Xiangyu Zhang', 'Dongyan Xu']",2014,"Performance bugs are frequently observed in commodity software. While profilers or source code-based tools can be used at development stage where a program is diagnosed in a well-defined environment, many performance bugs survive such a stage and affect production runs. OS kernel-level tracers are commonly used in post-development diagnosis due to their independence from programs and libraries; however, they lack detailed program-specific metrics to reason about performance problems such as function latencies and program contexts. In this paper, we propose a novel performance inference system, called IntroPerf, that generates fine-grained performance information -- like that from application profiling tools -- transparently by leveraging OS tracers that are widely available in most commodity operating systems. With system stack traces as input, IntroPerf enables transparent context-sensitive performance inference, and diagnoses application performance in a multi-layered scope ranging from user functions to the kernel. Evaluated with various performance bugs in multiple open source software projects, IntroPerf automatically ranks potential internal and external root causes of performance bugs with high accuracy without any prior knowledge about or instrumentation on the subject software. Our results show IntroPerf's effectiveness as a lightweight performance introspection tool for post-development diagnosis.",https://dl.acm.org/doi/10.1145/2591971.2592008,True,,20.0,,,,['root-cause-analysis'],,SIGMETRICS '14: The 2014 ACM international conference on Measurement and modeling of computer systems,True,['failure-management'],,,,,,10.1145/2591971.2592008,,,
866,"RAIDShield: Characterizing, Monitoring, and Proactively Protecting Against Disk Failures","['Ma, Ao', 'Traylor, Rachel', 'Douglis, Fred', 'Chamness, Mark', 'Lu, Guanlin', 'Sawyer, Darren', 'Chandra, Surendar', 'Hsu, Windsor']",2015,"Modern storage systems orchestrate a group of disks to achieve their performance and reliability goals. Even though such systems are designed to withstand the failure of individual disks, failure of multiple disks poses a unique set of challenges. We empirically investigate disk failure data from a large number of production systems, specifically focusing on the impact of disk failures on RAID storage systems. Our data covers about one million SATA disks from six disk models for periods up to 5 years. We show how observed disk …",https://www.researchgate.net/publication/272564146_RAIDShield_Characterizing_Monitoring_and_Proactively_Protecting_Against_Disk_Failures,True,,90.0,['naive-bayes'],['host-metrics'],,['failure-prediction'],['raid'],,True,['failure-management'],True,,['hardware-failure-prediction'],True,37.0,,,,
867,An integrated framework on mining logs files for computing system management,"['Tao Li', 'Feng Liang', 'Sheng Ma', 'Wei Peng']",2005,"Traditional approaches to system management have been largely based on domain experts through a knowledge acquisition process that translates domain knowledge into operating rules and policies. This has been well known and experienced as a cumbersome, labor intensive, and error prone process. In addition, this process is difficult to keep up with the rapidly changing environments. In this paper, we will describe our research efforts on establishing an integrated framework for mining system log files for automatic management. In particular, we apply text mining techniques to categorize messages in log files into common situations, improve categorization accuracy by considering the temporal characteristics of log messages, develop temporal mining techniques to discover the relationships between different events, and utilize visualization tools to evaluate and validate the interesting temporal patterns for system management.",https://dl.acm.org/doi/10.1145/1081870.1081972,True,,45.0,"['naive-bayes', 'markov-model']",['logs'],,['root-cause-analysis'],,KDD '05: Proceedings of the eleventh ACM SIGKDD international conference on Knowledge discovery in data mining,True,['failure-management'],,,,,,10.1145/1081870.1081972,,,
868,A Particle Swarm Optimization-Based Heuristic for Scheduling Workflow Applications in Cloud Computing Environments,"['Suraj Pandey', 'Linlin Wu', 'Siddeswara Mayura Guru', 'Rajkumar Buyya']",2010,"Cloud computing environments facilitate applications by providing virtualized resources that can be provisioned dynamically. However, users are charged on a pay-per-use basis. User applications may incur large data retrieval and execution costs when they are scheduled taking into account only the `execution time'. In addition to optimizing execution time, the cost arising from data transfers between resources as well as execution costs must also be taken into account. In this paper, we present a particle swarm optimization (PSO) based heuristic to schedule applications to cloud resources that takes into account both computation cost and data transmission cost. We experiment with a workflow application by varying its computation and communication costs. We compare the cost savings when using PSO and existing `Best Resource Selection' (BRS) algorithm. Our results show that PSO can achieve: (a) as much as 3 times cost savings as compared to BRS, and (b) good distribution of workload onto resources.",https://ieeexplore.ieee.org/document/5474725,True,,362.0,['particle-swarm'],,['novel-use'],['scheduling'],,2010 24th IEEE international conference on advanced information networking and applications,True,['resource-provisioning'],,,,,,,,,
869,Tiresias: Black-Box Failure Prediction in Distributed Systems,"['Andrew W. Williams', 'Soila M. Pertet', 'Priya Narasimhan']",2007,"Faults in distributed systems can result in errors that manifest in several ways, potentially even in parts of the system that are not collocated with the root cause. These manifestations often appear as deviations (or ""errors"") in performance metrics. By transparently gathering, and then identifying escalating anomalous behavior in, various node-level and system-level performance metrics, the Tiresias system makes black-box failure-prediction possible. Through the trend analysis of performance metrics, Tiresias provides a window of opportunity (look-ahead time) for system recovery prior to impending crash failures. We empirically validate the heuristic rules of the Tiresias system by analyzing fault-free and faulty performance data from a replicated middleware-based system.",https://ieeexplore.ieee.org/document/4228073,True,,73.0,,,,['failure-prediction'],,2007 IEEE international parallel and distributed processing symposium,True,['failure-management'],,,,,,,,,
870,NETradamus: A forecasting system for system event messages,"['Alexander Clemm', 'Malte Hartwig']",2010,"Proactive service assurance depends on the ability to anticipate problems in a network. This paper presents a system to forecast high-severity system event messages before they occur. Forecasts are based on a stream of syslog messages from devices, which is compared and scored against a set of codes that are representative of known problem event patterns. The system includes a component to mine relevant event patterns from past problem occurrences. In addition to describing the techniques and algorithms that the system is based on, the paper also provides an assessment of the effectiveness of those techniques.",https://ieeexplore.ieee.org/document/5488430,True,,11.0,,,,['failure-prediction'],,2010 IEEE Network Operations and Management Symposium-NOMS 2010,True,['failure-management'],,,,,,,,,
871,Resolution Recommendation for Event Tickets in Service Management,"['Wubai Zhou', 'Liang Tang', 'Chunqiu Zeng', 'Tao Li', 'Larisa Shwartz', 'Genady Ya. Grabarnik']",2016,"In recent years, IT service providers have rapidly achieved an automated service delivery model. Software monitoring systems are designed to actively collect and signal event occurrences and, when necessary, automatically generate incident tickets. Repeating events generate similar tickets, which in turn have a vast number of repeated problem resolutions likely to be found in earlier tickets. In this paper, we develop techniques to recommend appropriate resolution for incoming events by making use of similarities between the events and historical resolutions of similar events. Built on the traditional k-nearest neighbor algorithm (KNN), our proposed algorithms take into account false positives often generated by monitoring systems. An additional penalty is incorporated into the algorithms to control the number of misleading resolutions in the recommendation results. Moreover, as the effectiveness of the KNN heavily relies on the underlying similarity measurement, we proposed two other approaches to significantly improve our recommendation with respect to resolution relevance. One approach uses topic-level features to incorporate resolution information into the similarity measurement; the other uses metric learning to learn a more effective similarity measure. Extensive empirical evaluations on three ticket data sets demonstrate the effectiveness and efficiency of our proposed methods.",https://ieeexplore.ieee.org/document/7505624,True,,40.0,['similarity-matching'],['tickets'],['novel-use'],['remediation'],,,True,['failure-management'],True,True,['solution-recommendation'],True,89.0,,,,
872,Knowledge Guided Hierarchical Multi-Label Classification Over Ticket Data,"['Chunqiu Zeng', 'Wubai Zhou', 'Tao Li', 'Larisa Shwartz', 'Genady Ya Grabarnik']",2017,"Maximal automation of routine IT maintenance procedures is an ultimate goal of IT service management. System monitoring, an effective and reliable means for IT problem detection, generates monitoring ticket. In light of the ticket description, the underlying categories of the IT problem are determined, and the ticket is assigned to the corresponding processing teams for problem resolving. Automatic IT problem category determination acts as a critical part during the routine IT maintenance procedures. In practice, IT problem categories are naturally organized in a hierarchy by specialization. Utilizing the category hierarchy, this paper comes up with a hierarchical multi-label classification method to classify the monitoring tickets. In order to find the most effective classification, a novel contextual hierarchy (CH) loss is introduced in accordance with the problem hierarchy. Consequently, an arising optimization problem is solved by a new greedy algorithm named GLabel. Furthermore, as well as the ticket instance itself, the knowledge from the domain experts, which partially indicates some categories the given ticket may or may not belong to, can also be leveraged to guide the hierarchical multi-label classification. Accordingly, a multi-label inference with the domain expert knowledge is conducted on the basis of the given label hierarchy. The experiment demonstrates the great performance improvement by incorporating the domain knowledge during the hierarchical multi-label classification over the ticket data.",https://ieeexplore.ieee.org/document/7852503,True,,24.0,"['probabilistic-modeling', 'optimization']",['tickets'],,['remediation'],,,True,['failure-management'],True,,['triage'],True,88.0,,,,
873,Mining console logs for large-scale system problem detection,"['Wei Xu', 'Ling Huang', 'Armando Fox', 'David Patterson', 'Michael Jordan']",2008,"The console logs generated by an application contain messages that the application developers believed would be useful in debugging or monitoring the application. Despite the ubiquity and large size of these logs, they are rarely exploited in a systematic way for monitoring and debugging because they are not readily machine-parsable. In this paper, we propose a novel method for mining this rich source of information. First, we combine log parsing and text mining with source code analysis to extract structure from the console logs. Second, we extract features from the structured information in order to detect anomalous patterns in the logs using Principal Component Analysis (PCA). Finally, we use a decision tree to distill the results of PCA-based anomaly detection to a format readily understandable by domain experts (e.g. system operators) who need not be familiar with the anomaly detection algorithms. As a case study, we distill over one million lines of console logs from the Hadoop file system to a simple decision tree that a domain expert can readily understand; the process requires no operator intervention and we detect a large portion of runtime anomalies that are commonly overlooked.",https://dl.acm.org/doi/10.5555/1855895.1855899,True,,116.0,,['logs'],,['failure-detection'],,SysML'08: Proceedings of the Third conference on Tackling computer systems problems with machine learning techniques,True,['failure-management'],,,['anomaly-detection'],,,10.5555/1855895.1855899,,,
874,Constructing the knowledge base for cognitive IT service management,"['Qing Wang', 'Wubai Zhou', 'Chunqiu Zeng', 'Tao Li', 'Larisa Shwartz', 'Genady Ya. Grabarnik']",2017,"The increasing complexity of IT environments dictates the usage of intelligent automation driven by cognitive technologies, aiming at providing higher quality and more complex services. Inspired by cognitive computing, an integrated framework is proposed for a problem resolution. In order to improve the efficiency of the problem resolution process, it is crucial to formalize problem records and discover relationships between elements of the records, records overall and other technical information. In the proposed framework, the domain knowledge is modeled using ontology. The key contribution of the framework is a novel domain specific approach for extracting useful phrases, that enables an automation improvement through resolution recommendation utilizing the ontology modeling technique. The effectiveness and efficiency of our framework are evaluated by an extensive empirical study of a large scale real ticket data.",https://ieeexplore.ieee.org/document/8035012,True,,15.0,['rule-mining'],['tickets'],['new-method'],['remediation'],,2017 IEEE International Conference on Services Computing (SCC),True,['failure-management'],True,,['solution-recommendation'],True,90.0,,,,
875,Recommending resolutions for problems identified by monitoring,"['Liang Tang', 'Tao Li', 'Larisa Shwartz', 'Genady Grabarnik']",2013,"Service Providers are facing an increasingly intense competitive landscape and growing industry requirements. Modern service infrastructure management focuses on the development of methodologies and tools for improving the efficiency and quality of service. It is desirable to run a service in a fully automated operation environment. Automated problem resolution, however, is difficult. It is particularly difficult for the weakly-coupled service composition, since the coupling is not defined at design time. Monitoring software systems are designed to actively capture events and automatically generate incident tickets or event tickets. Repeating events generate similar event tickets, which in turn have a vast number of repeated problem resolutions likely to be found in earlier tickets. We apply a recommendation systems approach to resolution of event tickets. In addition, we extend the recommendation methodology to take into account possible falsity of some of the tickets. The paper presents an analysis of the historical event tickets from a large service provider and proposes two resolution-recommendation algorithms for event tickets utilizing historical tickets. The recommendation algorithms take into account false positives often generated by monitoring systems. An additional penalty is incorporated in the algorithms to control the number of misleading resolutions in the recommendation results. An extensive empirical evaluation on three ticket data sets demonstrates that our proposed algorithms achieve a high accuracy with a small percentage of misleading results.",https://ieeexplore.ieee.org/document/6572979,True,,25.0,['similarity-matching'],"['host-metrics', 'events', 'tickets']","['new-method', 'comparison']",['remediation'],,2013 IFIP/IEEE International Symposium on Integrated Network Management (IM 2013),True,['failure-management'],True,,['solution-recommendation'],,,,,,
876,LogSig: generating system events from raw textual logs,"['Liang Tang', 'Tao Li', 'Chang-Shing Perng']",2011,"Modern computing systems generate large amounts of log data. System administrators or domain experts utilize the log data to understand and optimize system behaviors. Most system logs are raw textual and unstructured. One main fundamental challenge in automated log analysis is the generation of system events from raw textual logs. Log messages are relatively short text messages but may have a large vocabulary, which often result in poor performance when applying traditional text clustering techniques to the log data. Other related methods have various limitations and only work well for some particular system logs. In this paper, we propose a message signature based algorithm logSig to generate system events from textual log messages. By searching the most representative message signatures, logSig categorizes log messages into a set of event types. logSig can handle various types of log data, and is able to incorporate human's domain knowledge to achieve a high performance. We conduct experiments on five real system log data. Experiments show that logSig outperforms other alternative algorithms in terms of the overall performance.",https://dl.acm.org/doi/10.1145/2063576.2063690,True,,37.0,,,,['root-cause-analysis'],,CIKM '11: Proceedings of the 20th ACM international conference on Information and knowledge management,True,['failure-management'],,,,,,10.1145/2063576.2063690,,,
877,Problem identification by mining trouble tickets,"['Vikrant Shimpi', 'Maitreya Natu', 'Vaishali Sadaphal', 'Vaishali Kulkarni']",2014,"IT systems of today's enterprises are continuously monitored and managed by a team of resolvers. Any problem in the system is reported in the form of trouble-tickets. A ticket contains various details of the observed problem. However, the knowledge of the actual problem is hidden in the ticket description along with other information. Knowledge of issues helps the service providers to better plan to improve cost and quality of operations. In this paper, we address the problem of extracting the issues from ticket descriptions. We discuss various challenges in issue extraction and present algorithms to handle different scenarios. We demonstrate the effectiveness of the proposed algorithms through two real-world case-studies.",https://dl.acm.org/doi/10.5555/2726970.2726983,True,,6.0,,,,['root-cause-analysis'],,COMAD '14: Proceedings of the 20th International Conference on Management of Data,True,['failure-management'],,,,,,10.5555/2726970.2726983,,,
878,Juggling the jigsaw: Towards automated problem inference from network trouble tickets,"['Rahul Potharaju', 'Navendu Jain', 'Cristina Nita-Rotaru']",2013,"This paper presents NetSieve, a system that aims to do automated problem inference from network trouble tickets. Network trouble tickets are diaries comprising fixed fields and free-form text written by operators to document the steps while troubleshooting a problem. Unfortunately, while tickets carry valuable information for network management, analyzing them to do problem inference is extremely difficult--fixed fields are often inaccurate or incomplete, and the free-form text is mostly written in natural language. This paper takes a practical step towards automatically analyzing natural language text in network tickets to infer the problem symptoms, troubleshooting activities and resolution actions. Our system, NetSieve, combines statistical natural language processing (NLP), knowledge representation, and ontology modeling to achieve these goals. To cope with ambiguity in free-form text, NetSieve leverages learning from human guidance to improve its inference accuracy. We evaluate NetSieve on 10K+ tickets from a large cloud provider, and compare its accuracy using (a) an expert review, (b) a study with operators, and (c) vendor data that tracks device replacement and repairs. Our results show that NetSieve achieves 89%-100% accuracy and its inference output is useful to learn global problem trends. We have used NetSieve in several key network operations: analyzing device failure trends, understanding why network redundancy fails, and identifying device problem symptoms.",https://dl.acm.org/doi/10.5555/2482626.2482640,True,,66.0,,['tickets'],,['root-cause-analysis'],['network'],nsdi'13: Proceedings of the 10th USENIX conference on Networked Systems Design and Implementation,True,['failure-management'],,,['root-cause-diagnosis'],,,10.5555/2482626.2482640,,,
879,ConfSeer: leveraging customer support knowledge bases for automated misconfiguration detection,"['Rahul Potharaju', 'Joseph Chan', 'Luhui Hu', 'Cristina Nita-Rotaru', 'Mingshi Wang', 'Liyuan Zhang', 'Navendu Jain']",2015,"We introduce ConfSeer, an automated system that detects potential configuration issues or deviations from identified best practices by leveraging a knowledge base (KB) of technical solutions. The intuition is that these KB articles describe the configuration problems and their fixes so if the system can accurately understand them, it can automatically pinpoint both the errors and their resolution. Unfortunately, finding an accurate match is difficult because (a) the KB articles are written in natural language text, and (b) configuration files typically contain a large number of parameters with a high value range. Thus, expert-driven manual troubleshooting is not scalable.While there are several state-of-the-art techniques proposed for individual tasks such as keyword matching, concept determination and entity resolution, none offer a practical end-to-end solution to detect problems in machine configurations. In this paper, we describe our experiences building ConfSeer using a novel combinations of ideas from natural language processing, information retrieval and interactive learning. ConfSeer powers the recommendation engine behind Microsoft Operations Management Suite that proposes fixes for software configuration errors. The system has been running in production for about a year to proactively find misconfigurations on tens of thousands of servers. Our evaluation of ConfSeer against an expert-defined rule-based commercial system, an expert survey and web search engines shows that it achieves 80%-97.5% accuracy and incurs low runtime overheads.",https://dl.acm.org/doi/abs/10.14778/2824032.2824079,True,,12.0,,,,['root-cause-analysis'],,Proceedings of the VLDB Endowment,True,['failure-management'],,,,,,10.14778/2824032.2824079,,,
880,Alert detection in system logs,"['Adam J. Oliner', 'Alex Aiken', 'Jon Stearley']",2008,"We present Nodeinfo, an unsupervised algorithm for anomaly detection in system logs. We demonstrate Nodeinfo's effectiveness on data from four of the world's most powerful supercomputers: using logs representing over 746 million processor-hours, in which anomalous events called alerts were manually tagged for scoring, we aim to automatically identify the regions of the log containing those alerts. We formalize the alert detection task in these terms, describe how Nodeinfo uses the information entropy of message terms to identify alerts, and present an online version of this algorithm, which is now in production use. This is the first work to investigate alert detection on (several) publicly-available supercomputer system logs, thereby providing a reproducible performance baseline.",https://ieeexplore.ieee.org/document/4781208,True,,105.0,,['logs'],,['failure-detection'],,2008 Eighth IEEE International Conference on Data Mining,True,['failure-management'],,,['anomaly-detection'],,,,,,
881,Correlating events with time series for incident diagnosis,"['Chen Luo', 'Jian-Guang Lou', 'Qingwei Lin', 'Qiang Fu', 'Rui Ding', 'Dongmei Zhang', 'Zhe Wang']",2014,"As online services have more and more popular, incident diagnosis has emerged as a critical task in minimizing the service downtime and ensuring high quality of the services provided. For most online services, incident diagnosis is mainly conducted by analyzing a large amount of telemetry data collected from the services at runtime. Time series data and event sequence data are two major types of telemetry data. Techniques of correlation analysis are important tools that are widely used by engineers for data-driven incident diagnosis. Despite their importance, there has been little previous work addressing the correlation between two types of heterogeneous data for incident diagnosis: continuous time series data and temporal event data. In this paper, we propose an approach to evaluate the correlation between time series data and event data. Our approach is capable of discovering three important aspects of event-timeseries correlation in the context of incident diagnosis: existence of correlation, temporal order, and monotonic effect. Our experimental results on simulation data sets and two real data sets demonstrate the effectiveness of the algorithm.",https://dl.acm.org/doi/10.1145/2623330.2623374,True,,52.0,,,,['root-cause-analysis'],,KDD '14: Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining,True,['failure-management'],,,['root-cause-diagnosis'],,,10.1145/2623330.2623374,,,
882,A clustering model based on matrix approximation with applications to cluster system log files,"['Tao Li', 'Wei Peng']",2005,"In system management applications, to perform automated analysis of the historical data across multiple components when problems occur, we need to cluster the log messages with disparate formats to automatically infer the common set of semantic situations and obtain a brief description for each situation. In this paper, we propose a clustering model where the problem of clustering is formulated as matrix approximations and the clustering objective is minimizing the approximation error between the original data matrix and the reconstructed matrix based on the cluster structures. The model explicitly characterizes the data and feature memberships and thus enables the descriptions of each cluster. We present a two-side spectral relaxation optimization procedure for the clustering model. We also establish the connections between our clustering model with existing approaches. Experimental results show the effectiveness of the proposed approach.",https://dl.acm.org/doi/abs/10.1007/11564096_62,True,,10.0,,,,['failure-detection'],,ECML'05: Proceedings of the 16th European conference on Machine Learning,True,['failure-management'],,,,,,10.1007/11564096_62,,,
883,Automatic classification of change requests for improved it service quality,"['Cristina Kadar', 'Dorothea Wiesmann', 'Jose Iria', 'Dirk Husemann', 'Mario Lucic']",2011,"Faulty changes to the IT infrastructure can lead to critical system and application outages, and therefore cause serious economical losses. In this paper, we describe a change planning support tool that aims at assisting the change requesters in leveraging aggregated information associated with the change, like past failure reasons or best implementation practices. The thus gained knowledge can be used in the subsequent planning and implementation steps of the change. Optimal matching of change requests with the aggregated information is achieved through the classification of the change request into about 200 fine-grained activities. We propose to automatically classify the incoming change requests using various information retrieval and machine learning techniques. The cost of building the classifiers is reduced by employing active learning techniques or by leveraging labeled features. Historical tickets from two customers were used to empirically assess and compare the accuracy of the different classification approaches (Lucene index, multinomial logistic regression, and generalized expectation criteria).",https://ieeexplore.ieee.org/document/5958118,True,,20.0,,,,['remediation'],,2011 Annual SRII Global Conference,True,['failure-management'],,,,,,,,,
884,Modeling Probabilistic Measurement Correlations for Problem Determination in Large-Scale Distributed Systems,"['Jing Gao', 'Guofei Jiang', 'Haifeng Chen', 'Jiawei Han']",2009,"With the growing complexity in computer systems, it has been a real challenge to detect and diagnose problems in today's large-scale distributed systems. Usually, the correlations between measurements collected across the distributed system contain rich information about the system behaviors, and thus a reasonable model to describe such correlations is crucially important in detecting and locating system problems. In this paper, we propose a transition probability model based on Markov properties to characterize pair-wise measurement correlations. The proposed method can discover both the spatial (across system measurements) and temporal (across observation time) correlations, and thus such a model can successfully represent the system normal profiles. Problem determination and localization under this framework is fast and convenient. The framework is general enough to discover any types of correlations (e.g. linear or non-linear). Also, model updating, system problem detection and diagnosis can be conducted effectively and efficiently. Experimental results show that, the proposed method can detect the anomalous events and locate the problematic sources by analyzing the real monitoring data collected from three companies' infrastructures.",https://ieeexplore.ieee.org/document/5158476,True,,44.0,,,,['root-cause-analysis'],,2009 29th IEEE International Conference on Distributed Computing Systems,True,['failure-management'],,,,,,,,,
885,A solution for identifying the root cause of problems in IT change management,"['Ricardo L. dos Santos', 'Juliano A. Wickboldt', 'Roben C. Lunardi', 'Bruno L. Dalmazo', 'Lisandro Z. Granville', 'Luciano P. Gaspary', 'Claudio Bartolini', 'Marianne Hickey']",2011,"The reuse of knowledge acquired by operators to diagnose failures in Information Technology (IT) infrastructures has potential to decrease the recurrence of failures and, consequently, reduce possible losses and maintenance costs. Nevertheless, existing solutions to support failure diagnosis lack of flexibility to adapt to a constantly changing IT environment. As a result, diagnostic is performed in an ad hoc and static fashion, which hampers the reuse of knowledge to solve similar failures affecting different elements of an IT infrastructure. To bridge this gap, in this paper we propose an extension of Common Information Model (CIM), supported by a conceptual solution for the identification of the root causes of problems, adaptable to changes in the target infrastructure and applicable to similar failures. Experiments carried out considering typical failures during the deployment of IT changes provide evidence about the efficacy of the proposed solution.",https://ieeexplore.ieee.org/document/5990563,True,,14.0,,,,['root-cause-analysis'],,12th IFIP/IEEE International Symposium on Integrated Network Management (IM 2011) and Workshops,True,['failure-management'],,,,,,,,,
886,Grid Resources Prediction with Support Vector Regression and Particle Swarm Optimization,"['Guosheng Hu', 'Liang Hu', 'Hongwei Li', 'Kun Li', 'Wei Liu']",2010,"Accurate grid resources prediction is crucial for a grid scheduler. In this study, support vector regression (SVR), which is an effective regression algorithm, is applied to grid resource prediction. In order to obtain better prediction performance, SVR's parameters must be selected carefully. Therefore, a particle swarm optimization-based SVR (PSO-SVR) model, in which PSO is used to determine free parameters of SVR, is presented in this study. The hybrid model (PSO-SVR) can automatically determine the parameters of SVR with higher predictive accuracy and generalization ability simultaneously. The performance of PSO-SVR, the back-propagation neural network (BPNN) and the traditional SVR model whose parameters are obtained by trail-and-error procedure (T-SVR) have been compared with benchmark data set. Experimental results indicate that the PSO-SVR model can achieve higher predictive accuracy than the other two models.",https://ieeexplore.ieee.org/document/5533166,True,,13.0,,,,,,2010 Third International Joint Conference on Computational Science and Optimization,True,['resource-provisioning'],,,,,,,,,
887,Software defect prediction system using multilayer perceptron neural network with data mining,"['Gayathri, M', 'Sudha, A']",2014,"Fault prediction in software systems is crucial for any software organization to produce quality and reliable software. Faults (defects) or fault-proneness of software modules are to be predicted in the early stages of software life cycle, so that more testing efforts can be put on faulty modules. Various metrics in software like Cyclomatic complexity, Lines of Code have been calculated and effectively used for predicting faults. Techniques like statistical methods, data mining, machine learning, and mixed algorithms, which were based on …",https://www.semanticscholar.org/paper/Software-Defect-Prediction-System-using-Multilayer-Gayathri-Sudha/b0dfe906644e9541cf2b79ddf65ad9828c4dcf1d,True,,23.0,,,,['failure-prediction'],,,True,['failure-management'],,,,,,,,,
888,Forecasting for Grid and Cloud Computing On-Demand Resources Based on Pattern Matching,"['Eddy Caron', 'Frederic Desprez', 'Adrian Muresan']",2010,"The Cloud phenomenon brings along the cost-saving benefit of dynamic scaling. As a result, the question of efficient resource scaling arises. Prediction is necessary as the virtual resources that Cloud computing uses have a setup time that is not negligible. We propose an approach to the problem of workload prediction based on identifying similar past occurrences of the current short-term workload history. We present in detail the Cloud client resource auto-scaling algorithm that uses the above approach to help when scaling decisions are made, as well as experimental results by using real-world traces from Cloud and Grid platforms. We also present an overall evaluation of this approach, its potential and usefulness for enabling efficient auto-scaling of Cloud user resources.",https://ieeexplore.ieee.org/document/5708485,True,,134.0,,,,"['resource-consolidation', 'workload-prediction']",,2010 IEEE Second International Conference on Cloud Computing Technology and Science,True,['resource-provisioning'],,,,,,,,,
889,Towards Machine Learning-Based Auto-tuning of MapReduce,"['Nezih Yigitbasi', 'Theodore L. Willke', 'Guangdeng Liao', 'Dick Epema']",2013,"MapReduce, which is the de facto programming model for large-scale distributed data processing, and its most popular implementation Hadoop have enjoyed widespread adoption in industry during the past few years. Unfortunately, from a performance point of view getting the most out of Hadoop is still a big challenge due to the large number of configuration parameters. Currently these parameters are tuned manually by trial and error, which is ineffective due to the large parameter space and the complex interactions among the parameters. Even worse, the parameters have to be re-tuned for different MapReduce applications and clusters. To make the parameter tuning process more effective, in this paper we explore machine learning-based performance models that we use to auto-tune the configuration parameters. To this end, we first evaluate several machine learning models with diverse MapReduce applications and cluster configurations, and we show that support vector regression model (SVR) has good accuracy and is also computationally efficient. We further assess our auto-tuning approach, which uses the SVR performance model, against the Starfish auto tuner, which uses a cost-based performance model. Our findings reveal that our auto-tuning approach can provide comparable or in some cases better performance improvements than Starfish with a smaller number of parameters. Finally, we propose and discuss a complete and practical end-to-end auto-tuning flow that combines our machine learning-based performance models with smart search algorithms for the effective training of the models and the effective exploration of the parameter space.",https://ieeexplore.ieee.org/document/6730744,True,,73.0,,,,['configuration'],,"2013 IEEE 21st International Symposium on Modelling, Analysis and Simulation of Computer and Telecommunication Systems",True,['resource-provisioning'],,,,,,,,,
890,Model ensemble tools for self-management in data centers,"['Jin Chen', 'Gokul Soundararajan', 'Saeed Ghanbari', 'Cristiana Amza']",2013,"We introduce Ensemble, a runtime framework and associated tools for building query latency models on-the-fly. These dynamic performance models can be used to support complex, highly dimensional resource allocation, and/or what-if performance inquiry in modern database environments, such as data centers and Clouds. Ensemble combines simple, partially specified, lower-dimensionality models to provide good initial approximations for higher dimensionality, end-to-end query latency models. We perform an experimental evaluation on industry-standard applications running on a multi-tier dynamic content server. We show that the Ensemble on-the-fly modeling framework provides accurate, fast and flexible performance modelling by using partial, lower dimensionality models to approximate end-to-end query latency models.",https://ieeexplore.ieee.org/document/6547424,True,,6.0,,,,,,2013 IEEE 29th International Conference on Data Engineering Workshops (ICDEW),True,['resource-provisioning'],,,,,,,,,
891,Hard drive failure prediction using Decision Trees,"['Li, Jing', 'Stones, Rebecca J', 'Wang, Gang', 'Liu, Xiaoguang', 'Li, Zhongwei', 'Xu, Ming']",2017,"This paper proposes two hard drive failure prediction models based on Decision Trees (DTs) and Gradient Boosted Regression Trees (GBRTs) which perform well in prediction performance as well as stability and interpretability. The models are evaluated on a real-world dataset containing 121, 698 drives in total. Experimental results show the DT model predicts over 93% of failures at a false alarm rate under 0.01%, and the GBRT model can achieve about 90% failure detection rate without any false alarms. Moreover, the GBRT …",https://www.sciencedirect.com/science/article/abs/pii/S0951832016301569?via%3Dihub,True,,22.0,"['decision-tree', 'regression-tree']",,['novel-use'],['failure-prediction'],['hard-drive'],,True,['failure-management'],True,,['hardware-failure-prediction'],,,,,,
892,Proactive error prediction to improve storage system reliability,"['Mahdisoltani, Farzaneh', 'Stefanovici, Ioan', 'Schroeder, Bianca']",2017,"This paper proposes the use of machine learning techniques to make storage systems more reliable in the face of sector errors. Sector errors are partial drive failures, where individual sectors on a drive become unavailable, and occur at a high rate in both hard disk drives and solid state drives. The data in the affected sectors can only be recovered through redundancy in the system (eg another drive in the same RAID) and is lost if the error is encountered while the system operates in degraded mode, eg during RAID reconstruction.",https://www.usenix.org/conference/atc17/technical-sessions/presentation/mahdisoltani,True,,35.0,"['regression-tree', 'decision-tree', 'support-vector-machine', 'multilayer-perceptron', 'logistic-regression']",['host-metrics'],"['novel-use', 'comparison']",['failure-prediction'],"['hard-drive', 'ssd']",2017 $\{$USENIX$\}$ Annual Technical Conference ($\{$USENIX$\}$$\{$ATC$\}$ 17),True,['failure-management'],True,,['hardware-failure-prediction'],True,34.0,,,,['robustness']
893,Log-based Abnormal Task Detection and Root Cause Analysis for Spark,"['Siyang Lu', 'BingBing Rao', 'Xiang Wei', 'Byungchul Tak', 'Long Wang', 'Liqiang Wang']",2017,"Application delays caused by abnormal tasks arecommon problems in big data computing frameworks. Anabnormal task in Spark, which may run slowly withouterror or warning logs, not only reduces its resident node'sperformance, but also affects other nodes' efficiency.Spark log files report neither root causes of abnormal tasks,nor where and when abnormal scenarios happen. AlthoughSpark provides a “speculation” mechanism to detect stragglertasks, it can only detect tailed stragglers in each stage. Sincethe root causes of abnormal happening are complicated, thereare no effective ways to detect root causes.This paper proposes an approach to detect abnormality andanalyzes root causes using Spark log files. Unlike commononline monitoring or analysis tools, our approach is a pureoff-line method that can analyze abnormality accurately. Ourapproach consists of four steps. First, a parser preprocessesraw log files to generate structured log data. Second, ineach stage of Spark application, we choose features relatedto execution time and data locality of each task, as well asmemory usage and garbage collection of each node. Third,based on the selected features, we detect where and whenabnormalities happen. Finally, we analyze the problems usingweighted factors to decide the probability of root causes. In thispaper, we consider four potential root causes of abnormalities,which include CPU, memory, network, and disk. The proposedmethod has been tested on real-world Spark benchmarks.To simulate various scenario of root causes, we conductedinterference injections related to CPU, memory, network,and Disk. Our experimental results show that the proposedapproach is accurate on detecting abnormal tasks as well asfinding the root causes",https://ieeexplore.ieee.org/document/8029786,True,,16.0,,,,['failure-detection'],,2017 IEEE International Conference on Web Services (ICWS),True,['failure-management'],,,,,,,,,
894,Cheetah: A Dynamic Performance Optimization Approach on Heterogeneous Big Data Analytics Cluster,"['Haizhou Du', 'Shaohua Zhang', 'Ping Han', 'Keke Zhang', 'Bin Xu']",2019,"Any MapReduce-based big data analytics clusters systems like Hadoop, Spark face the substantial challenge of the Long Tail problem: a small subset of straggling tasks significantly impede parallel jobs completion. Specially, in heterogeneous environments some tasks become stragglers because of poor performance of some computing nodes, data skew, etc. Therefore, stragglers are well recognized as a major bottleneck in big data processing and hence the early detection and accurate identification of stragglers can have important impacts on big data processing. In past years speculative execution strategies that have been proposed to meet their challenges such as misjudgment or delaying of straggling tasks, improper selection of backup nodes, etc., which result in inaccurate and inefficient performance of the speculative execution. Encouraged by recent successes in applying reinforcement learning (RL) techniques to solve complex online control problems, we study RL can be used for automatic choose right strategy to address stragglers without human-intervention. In this paper, we present Cheetah: a novel dynamic optimal approach for speculative strategy which identifies stragglers by reinforcement learning and automatic chooses the best strategy to launch speculative tasks on heterogeneous cluster. We implement Cheetah with popular reinforcement learning frameworks, and deploy it on a testbed of heterogeneous Spark cluster with 4 nodes. According to the experimental results, Compared to existing approaches, Cheetah reduces the job completion time in different type applications while achieving superior performance. For example, it demonstrates up to almost 16% reduction average job completion time and improve 25% on accuracy over existing solutions on the heterogeneous Spark cluster.",https://ieeexplore.ieee.org/document/8904996,True,,0.0,,,,,,2019 5th International Conference on Big Data Computing and Communications (BIGCOM),True,['resource-provisioning'],,,,,,,,,
895,Neural network based approach for time to crash prediction to cope with software aging,"['Moona Yakhchi', 'Javier Alonso', 'Mahdi Fazeli', 'Amir Akhavan Bitaraf', 'Ahmad Patooqhy']",2015,"Recent studies have shown that software is one of the main reasons for computer systems unavailability. A growing accumulation of software errors with time causes a phenomenon called software aging. This phenomenon can result in system performance degradation and eventually system hang/crash. To cope with software aging, software rejuvenation has been proposed. Software rejuvenation is a proactive technique which leads to removing the accumulated software errors by stopping the system, cleaning up its internal state, and resuming its normal operation. One of the main challenges of software rejuvenation is accurately predicting the time to crash due to aging factors such as memory leaks. In this paper, different machine learning techniques are compared to accurately predict the software time to crash under different aging scenarios. Finally, by comparing the accuracy of different techniques, it can be concluded that the multilayer perceptron neural network has the highest prediction accuracy among all techniques studied.",https://ieeexplore.ieee.org/document/7111178,True,,2.0,,,,['failure-prediction'],,,True,['failure-management'],,,,,,,,,
896,Detection of Memory Leaks in C/C++ Code via Machine Learning,"['Artur Andrzejak', 'Felix Eichler', 'Mohammad Ghanavati']",2017,"Memory leaks are one of the primary causes of software aging. Despite of recent countermeasures in C/C++ such as smart pointers, leak-related defects remain a troublesome issue in C/C++ code, especially in legacy applications.We propose an approach for automatic detection of memory leaks in C/C++ programs based on characterizing memory allocation sites via the age distribution of the non-disposed memory chunks allocated by such a site (the so-called GenCount-technique introduced for Java by Vladimir Šor). We instrument malloc and free calls in C/C++ and collect for each allocation site data on the number of allocated memory fragments, their lifetimes, and sizes. Based on this data we compute feature vectors and train a machine learning classifier to differentiate between leaky and defect-free allocation sites.Our evaluation uses applications from SPEC CPU2006 suite with injected memory leaks resembling real leaks. The results show that even out-of-the-box classification algorithms can achieve high accuracy, with precision and recall values of 0.93 and 0.88, respectively.",https://ieeexplore.ieee.org/document/8109292,True,,6.0,,,,['failure-prediction'],,2017 IEEE International Symposium on Software Reliability Engineering Workshops (ISSREW),True,['failure-management'],,,,,,,,,
897,Software Defect Prediction via Convolutional Neural Network,"['Jian Li', 'Pinjia He', 'Jieming Zhu', 'Michael R. Lyu']",2017,"To improve software reliability, software defect prediction is utilized to assist developers in finding potential bugs and allocating their testing efforts. Traditional defect prediction studies mainly focus on designing hand-crafted features, which are input into machine learning classifiers to identify defective code. However, these hand-crafted features often fail to capture the semantic and structural information of programs. Such information is important in modeling program functionality and can lead to more accurate defect prediction. In this paper, we propose a framework called Defect Prediction via Convolutional Neural Network (DP-CNN), which leverages deep learning for effective feature generation. Specifically, based on the programs' Abstract Syntax Trees (ASTs), we first extract token vectors, which are then encoded as numerical vectors via mapping and word embedding. We feed the numerical vectors into Convolutional Neural Network to automatically learn semantic and structural features of programs. After that, we combine the learned features with traditional hand-crafted features, for accurate software defect prediction. We evaluate our method on seven open source projects in terms of F-measure in defect prediction. The experimental results show that in average, DP-CNN improves the state-of-the-art method by 12%.",https://ieeexplore.ieee.org/document/8009936,True,,84.0,['cnn'],['source-code'],['new-method'],['failure-prevention'],['source-code'],"2017 IEEE International Conference on Software Quality, Reliability and Security (QRS)",True,['failure-management'],True,,['software-defect-prediction'],True,14.0,,,,
898,An Improved Approach to Software Defect Prediction using a Hybrid Machine Learning Model,['Diana-Lucia Miholca'],2018,"Software defect prediction is an intricate but essential software testing related activity. As a solution to it, we have recently proposed HyGRAR, a hybrid classification model which combines Gradual Relational Association Rules (GRARs) with ANNs. ANNs were used to learn gradual relations that were then considered in a mining process so as to discover the interesting GRARs characterizing the defective and non-defective software entities, respectively. The classification of a new entity based on the discriminative GRARs was made through a non-adaptive heuristic method. In current paper, we propose to enhance HyGRAR through autonomously learning the classification methodology. Evaluation experiments performed on two open-source data sets indicate that the enhanced HyGRAR classifier outperforms the related approaches evaluated on the same two data sets.",https://ieeexplore.ieee.org/document/8750697,True,,0.0,,,,['failure-prediction'],,2018 20th International Symposium on Symbolic and Numeric Algorithms for Scientific Computing (SYNASC),True,['failure-management'],,,,,,,,,
899,Software Defect Prediction Using Random Forest Algorithm,"['Yan Naung Soe', 'Paulus Insap Santosa', 'Rudy Hartanto']",2018,"The software defect can cause the unnecessary effects on the software such as cost and quality. The prediction of the software defect can be useful for the development of good quality software. For the prediction, the PROMISE public dataset will be used and random forest (RF) algorithm will be applied with the RAPIDMINER machine learning tool. This paper will compare the performance evaluation upon the different number of trees in RF. As the results, the accuracy will be slightly increased if the number of trees will be more. The maximum accuracy is up to 99.59 and the minimum accuracy is 85.96. Another comparison is based on AUC curve that represents the most informative indicator of predictive accuracy within the field of software defect prediction. All of the results show that RF algorithm is effective in this prediction which is more suitable with the usage of hundred trees in the RF.",https://ieeexplore.ieee.org/document/8788881,True,,0.0,['random-forest'],,['novel-use'],['failure-prevention'],['source-code'],2018 12th South East Asian Technical University Consortium (SEATUC),True,['failure-management'],,,['software-defect-prediction'],,,,,,
900,Improving Performance in Software Defect Prediction Using Variational Autoencoder,"['Z. Eivazpour', 'Mohammad Reza Keyvanpour']",2019,"Software defect prediction (SDP) is a beneficial task to save limited resources in the software testing stage for improving software quality. However, the imbalanced distribution in defect datasets could be a challenge for often machine learning algorithms, an effect on the performance of the algorithms. To overcome this issue, oversampling techniques from the minority class has been adopted. In this work, we suggest a new oversampling method, which trained a variational autoencoder (VAE) to generate synthesized samples aimed for output mimicked minority samples that were then combined with training dataset into an augmented training dataset. In the experiments, we explored ten SDP datasets from the PROMISE freely accessible repository. We measured the performance of the proposed method by comparing it with state-of-the-art oversampling techniques including Random Over-Sampling, SMOTE, Borderline-SMOTE, and ADASYN. Based on the investigation results, the proposed method provides better mean performance of SDP models between all examined techniques.",https://ieeexplore.ieee.org/document/8734915,True,,1.0,['autoencoder'],,"['comparison', 'novel-use']",['failure-prevention'],['source-code'],,True,['failure-management'],,,['software-defect-prediction'],,,,,,['imbalance']
901,Software defect prediction via transfer learning based neural network,"['Qimeng Cao', 'Qing Sun', 'Qinghua Cao', 'Huobin Tan']",2015,"Software defect (Bug) prediction plays an important role in improving software quality. Many software defect prediction approaches have been proposed and achieved great effects in the real-world. However, the existing works are usually constrained in only one project, hence their effectiveness on cross-project defect prediction (cross-prediction) is usually poor. This is mainly because of the problem of class imbalance and feature distribution differences between the source and target projects. In this paper, we proposed an effective software defect prediction method called Transfer Component Analysis Neural Network (TCANN), by adequately considering the noise data, the class imbalance in data settings and transfer learning among cross-project. There are three parts in TCANN, aiming to solve the above mentioned three problems respectively. First, the Inter Quartile Range (IQR) based method is proposed for noise removal in datasets. Second, The transfer component analysis method is used to reduce the feature distribution differences between source and target data. Third, dynamic sampling neural network is proposed for dealing with class imbalance problem of the training dataset. Based on the classic open-source datasets collected by previous researchers, our experimental results show that TCANN improves the performance of both within-project and cross-project defect prediction in comparison with other methods.",https://ieeexplore.ieee.org/document/7366475,True,,13.0,,,,['failure-prediction'],,2015 First International Conference on Reliability Systems Engineering (ICRSE),True,['failure-management'],,,,,,,,,
902,Integrated Approach to Software Defect Prediction,"['Ebubeogu Amarachukwu Felix', 'Sai Peck Lee']",2017,"Software defect prediction provides actionable outputs to software teams while contributing to industrial success. Empirical studies have been conducted on software defect prediction for both cross-project and within-project defect prediction. However, existing studies have yet to demonstrate a method of predicting the number of defects in an upcoming product release. This paper presents such a method using predictor variables derived from the defect acceleration, namely, the defect density, defect velocity, and defect introduction time, and determines the correlation of each predictor variable with the number of defects. We report the application of an integrated machine learning approach based on regression models constructed from these predictor variables. An experiment was conducted on ten different data sets collected from the PROMISE repository, containing 22838 instances. The regression model constructed as a function of the average defect velocity achieved an adjusted R-square of 98.6%, with a p-value of <; 0.001. The average defect velocity is strongly positively correlated with the number of defects, with a correlation coefficient of 0.98. Thus, it is demonstrated that this technique can provide a blueprint for program testing to enhance the effectiveness of software development activities.",https://ieeexplore.ieee.org/document/8058420,True,,21.0,,,,['failure-prediction'],,,True,['failure-management'],,,,,,,,,
903,Convolutional Neural Networks over Control Flow Graphs for Software Defect Prediction,"['Anh Viet Phan', 'Minh Le Nguyen', 'Lam Thu Bui']",2017,"Existing defects in software components is unavoidable and leads to not only a waste of time and money but also many serious consequences. To build predictive models, previous studies focus on manually extracting features or using tree representations of programs, and exploiting different machine learning algorithms. However, the performance of the models is not high since the existing features and tree structures often fail to capture the semantics of programs. To explore deeply programs' semantics, this paper proposes to leverage precise graphs representing program execution flows, and deep neural networks for automatically learning defect features. Firstly, control flow graphs are constructed from the assembly instructions obtained by compiling source code; we thereafter apply multi-view multi-layer directed graph-based convolutional neural networks (DGCNNs) to learn semantic features. The experiments on four real-world datasets show that our method significantly outperforms the baselines including several other deep learning approaches.",https://ieeexplore.ieee.org/document/8371922,True,,20.0,['cnn'],,,['failure-prevention'],,2017 IEEE 29th International Conference on Tools with Artificial Intelligence (ICTAI),True,['failure-management'],,,['software-defect-prediction'],,,,,,
904,Transforming reactive auto-scaling into proactive auto-scaling,"['Laura R. Moore', 'Kathryn Bean', 'Tariq Ellahi']",2013,"Elasticity is a key characteristic of cloud platforms enabling resource to be acquired on-demand in response to time-varying workloads. We introduce a new elasticity management framework that takes as input commonly used reactive rule-based scaling strategies but offers in return proactive auto-scaling. The elasticity framework combines reactive and predictive auto-scaling techniques, and we discuss the specification and performance of these individual components. We present a case study, based on real datasets, to demonstrate that our framework is capable of making appropriate auto-scaling decisions that can improve resource utilization compared to that obtained from a purely reactive approach.",https://dl.acm.org/doi/abs/10.1145/2460756.2460758?download=true,True,,43.0,,,,['resource-consolidation'],,CloudDP '13: Proceedings of the 3rd International Workshop on Cloud Data and Platforms,True,['resource-provisioning'],,,,,,10.1145/2460756.2460758?download=true,,,
905,Fault Prediction using Early Lifecycle Data,"['Yue Jiang', 'Bojan Cukic', 'Tim Menzies']",2007,"The prediction of fault-prone modules in a software project has been the topic of many studies. In this paper, we investigate whether metrics available early in the development lifecycle can be used to identify fault-prone software modules. More precisely, we build predictive models using the metrics that characterize textual requirements. We compare the performance of requirements-based models against the performance of code-based models and models that combine requirement and code metrics. Using a range of modeling techniques and the data from three NASA projects, our study indicates that the early lifecycle metrics can play an important role in project management, either by pointing to the need for increased quality monitoring during the development or by using the models to assign verification and validation activities.",https://ieeexplore.ieee.org/document/4402215,True,,126.0,,['code-metrics'],,['failure-prevention'],['source-code'],,True,['failure-management'],,,['software-defect-prediction'],,,,,,
906,Intelligence System for Software Maintenance Severity Prediction,"['Sandhu, Parvinder Singh', 'Kumar, Sunil', 'Singh, Hardeep']",2007,"The software industry has been experiencing a software crisis, a difficulty of delivering software within budget, on time, and of good quality. This may happen due to number of defects present in the different modules of the project that may require maintenance. This necessitates the need of predicting maintenance urgency of the particular module in the software. In this paper, we have applied the different predictor models to NASA five public domain defect datasets coded in C, C++, Java and Perl programming languages. Twenty one software metrics of different datasets and Java Classes of thirty five algorithms belonging to the different learner categories of the WEKA project have been evaluated for the prediction of maintenance severity. The results of ten fold cross validation are recorded in terms of Accuracy, Mean Absolute Error (MAE) and Root Mean Squared Error (RMSE) for different project datasets. The results show that logistic model Trees (LMT) and Complimentary Naïve Bayes (CNB) based Model provide a relatively better prediction consistency compared to other models and hence, can be used for the maintenance severity prediction of the software. The developed system can also be used for analysis and to evaluate the influence of different factors on the maintenance severity of different software project modules.",https://www.semanticscholar.org/paper/Intelligence-System-for-Software-Maintenance-Sandhu-Kumar/274c4cb2f51afd4ae7fba338f1e65b29d98a3457,True,,20.0,"['linear-regression', 'naive-bayes', 'logistic-regression', 'decision-tree']",['software-metrics'],['comparison'],['failure-prevention'],['source-code'],,True,['failure-management'],,,['software-defect-prediction'],,,10.3844/jcssp.2007.281.288,,,
907,Object-oriented software fault prediction using neural networks,"['S. Kanmani', 'V. Rhymend Uthariaraj', 'V. Sankaranarayanan', 'P. Thambidurai']",2007,"This paper introduces two neural network based software fault prediction models using Object-Oriented metrics. They are empirically validated using a data set collected from the software modules developed by the graduate students of our academic institution. The results are compared with two statistical models using five quality attributes and found that neural networks do better. Among the two neural networks, Probabilistic Neural Networks outperform in predicting the fault proneness of the Object-Oriented modules developed.",https://dl.acm.org/doi/10.1016/j.infsof.2006.07.005,True,,31.0,,,,['failure-prediction'],,Information and Software Technology,True,['failure-management'],,,,,,10.1016/j.infsof.2006.07.005,,,
908,Predicting defect-prone software modules using support vector machines,"['Elish, Karim O', 'Elish, Mahmoud O']",2008,"Effective prediction of defect-prone software modules can enable software developers to focus quality assurance activities and allocate effort and resources more efficiently. Support vector machines (SVM) have been successfully applied for solving both classification and regression problems in many applications. This paper evaluates the capability of SVM in predicting defect-prone software modules and compares its prediction performance against eight statistical and machine learning models in the context of four NASA datasets. The results indicate that the prediction performance of SVM is generally better than, or at least, is competitive against the compared models.",https://www.sciencedirect.com/science/article/pii/S016412120700235X?via%3Dihub,True,,461.0,['support-vector-machine'],['code-metrics'],"['comparison', 'novel-use']",['failure-prevention'],['source-code'],,True,['failure-management'],True,,['software-defect-prediction'],True,3.0,,,,
909,Applying machine learning to software fault-proneness prediction,"['Gondra, Iker']",2008,"The importance of software testing to quality assurance cannot be overemphasized. The estimation of a module’s fault-proneness is important for minimizing cost and improving the effectiveness of the software testing process. Unfortunately, no general technique for estimating software fault-proneness is available. The observed correlation between some software metrics and fault-proneness has resulted in a variety of predictive models based on multiple metrics. Much work has concentrated on how to select the software metrics that are most likely to indicate fault-proneness. In this paper, we propose the use of machine learning for this purpose. Specifically, given historical data on software metric values and number of reported errors, an Artificial Neural Network (ANN) is trained. Then, in order to determine the importance of each software metric in predicting fault-proneness, a sensitivity analysis is performed on the trained ANN. The software metrics that are deemed to be the most critical are then used as the basis of an ANN-based predictive model of a continuous measure of fault-proneness. We also view fault-proneness prediction as a binary classification task (i.e., a module can either contain errors or be error-free) and use Support Vector Machines (SVM) as a state-of-the-art classification method. We perform a comparative experimental study of the effectiveness of ANNs and SVMs on a data set obtained from NASA’s Metrics Data Program data repository.",https://www.sciencedirect.com/science/article/pii/S0164121207001240,True,,220.0,"['multilayer-perceptron', 'support-vector-machine']",['software-metrics'],"['comparison', 'novel-use']",['failure-prevention'],['source-code'],,True,['failure-management'],,,['software-defect-prediction'],,,,,,
910,Application of reinforcement learning to routing in distributed wireless networks: a review,"['Hasan A. Al-Rawi', 'Ming Ann Ng', 'Kok-Lim Alvin Yau']",2015,"The dynamicity of distributed wireless networks caused by node mobility, dynamic network topology, and others has been a major challenge to routing in such networks. In the traditional routing schemes, routing decisions of a wireless node may solely depend on a predefined set of routing policies, which may only be suitable for a certain network circumstances. Reinforcement Learning (RL) has been shown to address this routing challenge by enabling wireless nodes to observe and gather information from their dynamic local operating environment, learn, and make efficient routing decisions on the fly. In this article, we focus on the application of the traditional, as well as the enhanced, RL models, to routing in wireless networks. The routing challenges associated with different types of distributed wireless networks, and the advantages brought about by the application of RL to routing are identified. In general, three types of RL models have been applied to routing schemes in order to improve network performance, namely Q-routing, multi-agent reinforcement learning, and partially observable Markov decision process. We provide an extensive review on new features in RL-based routing, and how various routing challenges and problems have been approached using RL. We also present a real hardware implementation of a RL-based routing scheme. Subsequently, we present performance enhancements achieved by the RL-based routing schemes. Finally, we discuss various open issues related to RL-based routing schemes in distributed wireless networks, which help to explore new research directions in this area. Discussions in this article are presented in a tutorial manner in order to establish a foundation for further research in this field.",https://dl.acm.org/doi/10.1007/s10462-012-9383-6,True,,57.0,,,,['configuration'],"['wireless', 'network']",Artificial Intelligence Review,True,['resource-provisioning'],,,,,,10.1007/s10462-012-9383-6,,,
911,Cooperative reinforcement learning approach for routing in ad hoc networks,"['Rahul Desai', 'B P Patil']",2015,"Most of the routing algorithms over ad hoc networks are based on the status of the link (up or down). They are not capable of adapting the run time changes such as traffic load, delay and delivery time to reach to the destination etc, thus though provides shortest path, these shortest path may not be optimum path to deliver the packets. Optimum path can only be achieved when quality of links within the network is detected on continuous basis instead of discrete time. Thus for achieving optimum routes we model ad hoc routing as a cooperative reinforcement learning problem. In this paper, agents are used to optimize the performance of a network on trial and error basis. This learning strategy is based work in swarm intelligence: those systems whose design is inspired by models of social insect behaviour. This paper describes the algorithm used in cooperative reinforcement learning approach and performs the analysis by comparing with existing routing protocols.",https://ieeexplore.ieee.org/document/7086962,True,,10.0,,,,,,2015 International Conference on Pervasive Computing (ICPC),True,['resource-provisioning'],,,,,,,,,
912,A Reinforcement Learning Network based Novel Adaptive Routing Algorithm for Wireless Ad-Hoc Network,"['Solanki, Jagrut', 'Chauhan, Anand']",2015,"Mobile communication has enjoyed an incredible rise in quality throughout the last decade. Network dependability is most important concern in wireless Ad-hoc network. a serious challenge that lies in MANET (Mobile Ad-hoc network) is that the unlimited mobility and lots of frequent failure because of link breakage. Standard routing algorithms are insufficient for Ad-hoc networks. as a results of major drawback in MANET is limited power provide, dynamic networking. In MANET each node works as a router and autonomously performs …",https://www.semanticscholar.org/paper/A-Reinforcement-Learning-Network-based-Novel-for-Solanki/c91c36c98131647233312b863a812529735ca708,True,,3.0,,,,,,,True,['resource-provisioning'],,,,,,,,,
913,QoS-Aware Adaptive Routing in Multi-layer Hierarchical Software Defined Networks: A Reinforcement Learning Approach,"['Shih-Chun Lin', 'Ian F. Akyildiz', 'Pu Wang', 'Min Luo']",2016,"Software-defined networks (SDNs) have been recognized as the next-generation networking paradigm that decouples the data forwarding from the centralized control. To realize the merits of dedicated QoS provisioning and fast route (re-)configuration services over the decoupled SDNs, various QoS requirements in packet delay, loss, and throughput should be supported by an efficient transportation with respect to each specific application. In this paper, a QoS-aware adaptive routing (QAR) is proposed in the designed multi-layer hierarchical SDNs. Specifically, the distributed hierarchical control plane architecture is employed to minimize signaling delay in large SDNs via three-levels design of controllers, i.e., the super, domain (or master), and slave controllers. Furthermore, QAR algorithm is proposed with the aid of reinforcement learning and QoS-aware reward function, achieving a time-efficient, adaptive, QoS-provisioning packet forwarding. Simulation results confirm that QAR outperforms the existing learning solution and provides fast convergence with QoS provisioning, facilitating the practical implementations in large-scale software service-defined networks.",https://ieeexplore.ieee.org/document/7557432,True,,89.0,,,,['configuration'],,2016 IEEE International Conference on Services Computing (SCC),True,['resource-provisioning'],,,,,,,,,
914,Failure Prediction Methodology for Improved Proactive Maintenance using Bayesian Approach,"['Abu-Samah, A', 'Shahzad, MK', 'Zamai, E', 'Said, A Ben']",2015,"Failure prediction is essential for predictive maintenance due to its ability to prevent failure occurrences and maintenance costs. At present, mathematical and statistical modeling are the prominent approaches used for failure predictions. These are based on equipment degradation physical models and machine learning methods, respectively. None of these approaches ensures failure predictions well before their occurrence to provide sufficient time to treat potential causes pro actively. Therefore, in this paper, we present a Bayesian based …",https://www.sciencedirect.com/science/article/pii/S2405896315017619?via%3Dihub,True,,15.0,,,,['failure-prediction'],,,True,['failure-management'],,,,,,,,,
915,An evaluation of linear models for host load prediction,"['P.A. Dinda', ""D.R. O'Hallaron""]",1999,"Evaluates linear models for predicting the Digital Unix five-second host load average from 1 to 30 seconds into the future. A detailed statistical study of a large number of long, fine-grain load traces from a variety of real machines leads to consideration of the Box-Jenkins (1994) models (AR, MA, ARMA, ARIMA), and the ARFIMA (autoregressive fractional integrated moving average) models (due to self-similarity). These models, as well as a simple windowed-mean scheme, are then rigorously evaluated by running a large number of randomized test cases on the load traces and by data-mining their results. The main conclusions are that the load is consistently predictable to a very useful degree, and that the simpler models, such as AR, are sufficient for performing this prediction.",https://ieeexplore.ieee.org/document/805285,True,,171.0,,,,['workload-prediction'],,Proceedings. The Eighth International Symposium on High Performance Distributed Computing (Cat. No. 99TH8469),True,['resource-provisioning'],,,,,,,,,
916,Load prediction using hybrid model for computational grid,"['Yongwei Wu', 'Yulai Yuan', 'Guangwen Yang', 'Weimin Zheng']",2007,"Due to the dynamic nature of grid environments, schedule algorithms always need assistance of a long-time-ahead load prediction to make decisions on how to use grid resources efficiently. In this paper, we present and evaluate a new hybrid model, which predicts the
n
-step-ahead load status by using interval values. This model integrates autoregressive (AR) model with confidence interval estimations to forecast the future load of a system. Meanwhile, two filtering technologies from signal processing field are also introduced into this model to eliminate data noise and enhance prediction accuracy. The results of experiments conducted on a real grid environment demonstrate that this new model is more capable of predicting
n
-step-ahead load in a computational grid than previous works. The proposed hybrid model performs well on prediction advance time for up to 50 minutes, with significant less prediction errors than conventional AR model. It also achieves an interval length acceptable for task scheduler.",https://ieeexplore.ieee.org/document/4354138,True,,48.0,,,,['workload-prediction'],,2007 8th IEEE/ACM International Conference on Grid Computing,True,['resource-provisioning'],,,,,,,,,
917,Multi-step-ahead host load prediction using autoencoder and echo state networks in cloud computing,"['Yang, Qiangpeng', 'Zhou, Yu', 'Yu, Yao', 'Yuan, Jie', 'Xing, Xianglei', 'Du, Sidan']",2015,"Cloud computing is a model for enabling ubiquitous, convenient, on-demand network access to a shared pool of configurable computing resources that can be rapidly provisioned and released with minimal management effort or service provider interaction. There are many proposals for resource management approaches for cloud infrastructures, but effective resource management is still a major challenge for the leading cloud infrastructure operators (eg, Amazon, Microsoft, Google), because the details of the underlying workloads and the …",https://link.springer.com/article/10.1007%2Fs11227-015-1426-8,True,,21.0,"['autoencoder', 'rnn']",,['novel-use'],['workload-prediction'],,,True,['resource-provisioning'],,,,,,,,,
918,Host Load Prediction Based on PSR and EA-GMDH for Cloud Computing System,"['Qiangpeng Yang', 'Chenglei Peng', 'Yao Yu', 'He Zhao', 'Yu Zhou', 'Ziqiang Wang', 'Sidan Du']",2013,"Host Load Prediction is one of the most effective measures to improve resource utilization in the Cloud systems. As the drastic fluctuation of the host load in the Cloud, accurate prediction of host load is still a challenge. In this paper, we propose a new prediction method which combines the Phase Space Reconstruction (PSR) method and the Group Method of Data Handling (GMDH) based on Evolutionary Algorithm (EA). Our proposed method could predict not only the mean load in consecutive future time intervals, but also the actual load in each consecutive future time interval. We evaluate our method using the host load trace in the Google data center with thousands of machines. According to the experiment results, our method outperforms the other methods by more than 60% in mean load prediction, and preforms well on actual load prediction over different time intervals, i.e. 0.5h to 3h.",https://ieeexplore.ieee.org/document/6686002,True,,9.0,,,,,,2013 International Conference on Cloud and Green Computing,True,['resource-provisioning'],,,,,,,,,
919,Predicting Short-Term Traffic Flow by Long Short-Term Memory Recurrent Neural Network,"['Yongxue Tian', 'Li Pan']",2015,"Intelligent Transportation System (ITS) is a significant part of smart city, and short-term traffic flow prediction plays an important role in intelligent transportation management and route guidance. A number of models and algorithms based on time series prediction and machine learning were applied to short-term traffic flow prediction and achieved good results. However, most of the models require the length of the input historical data to be predefined and static, which cannot automatically determine the optimal time lags. To overcome this shortage, a model called Long Short-Term Memory Recurrent Neural Network (LSTM RNN) is proposed in this paper, which takes advantages of the three multiplicative units in the memory block to determine the optimal time lags dynamically. The dataset from Caltrans Performance Measurement System (PeMS) is used for building the model and comparing LSTM RNN with several well-known models, such as random walk(RW), support vector machine(SVM), single layer feed forward neural network(FFNN) and stacked autoencoder(SAE). The results show that the proposed prediction model achieves higher accuracy and generalizes well.",https://ieeexplore.ieee.org/document/7463717,True,,89.0,['rnn'],,['novel-use'],['workload-prediction'],,2015 IEEE international conference on smart city/SocialCom/SustainCom (SmartCity),True,['resource-provisioning'],,,,,,,,,
920,Predicting computer system failures using support vector machines,"['Errin W. Fulp', 'Glenn A. Fink', 'Jereme N. Haack']",2008,"Mitigating the impact of computer failure is possible if accurate failure predictions are provided. Resources, applications, and services can be scheduled around predicted failure and limit the impact. Such strategies are especially important for multi-computer systems, such as compute clusters, that experience a higher rate failure due to the large number of components. However providing accurate predictions with sufficient lead time remains a challenging problem. This paper describes a new spectrum-kernel Support Vector Machine (SVM) approach to predict failure events based on system log files. These files containmessages that represent a change of system state. While a single message in the file may not be sufficient for predicting failure, a sequence or pattern of messages may be. The approach described in this paper will use a sliding window (sub-sequence) of messages to predict the likelihood of failure. The a frequency representation of the message sub-sequences observed are then used as input to the SVM. The SVM then associates the messages to a class of failed or non-failed system. Experimental results using actual system log files from a Linux-based compute cluster indicate the proposed spectrum-kernel SVM approach has promise and can predict hard disk failure with an accuracy of 73% two days in advance.",https://dl.acm.org/doi/10.5555/1855886.1855891,True,,89.0,['support-vector-machine'],['logs'],,['failure-prediction'],,WASL'08: Proceedings of the First USENIX conference on Analysis of system logs,True,['failure-management'],,,['system-failure-prediction'],,,10.5555/1855886.1855891,,,
921,Failure prediction based on log files using Random Indexing and Support Vector Machines,"['Ilenia Fronza', 'Alberto Sillitti', 'Giancarlo Succi', 'Mikko Terho', 'Jelena Vlasenko']",2013,"Research problem: The impact of failures on software systems can be substantial since the recovery process can require unexpected amounts of time and resources. Accurate failure predictions can help in mitigating the impact of failures. Resources, applications, and services can be scheduled to limit the impact of failures. However, providing accurate predictions sufficiently ahead is challenging. Log files contain messages that represent a change of system state. A sequence or a pattern of messages may be used to predict failures. Contribution: We describe an approach to predict failures based on log files using Random Indexing (RI) and Support Vector Machines (SVMs). Method: RI is applied to represent sequences: each operation is characterized in terms of its context. SVMs associate sequences to a class of failures or non-failures. Weighted SVMs are applied to deal with imbalanced datasets and to improve the true positive rate. We apply our approach to log files collected during approximately three months of work in a large European manufacturing company. Results: According to our results, weighted SVMs sacrifice some specificity to improve sensitivity. Specificity remains higher than 0.80 in four out of six analyzed applications. Conclusions: Overall, our approach is very reliable in predicting both failures and non-failures.",https://dl.acm.org/doi/10.1016/j.jss.2012.06.025,True,,88.0,['support-vector-machine'],['logs'],,['failure-prediction'],,Journal of Systems and Software,True,['failure-management'],True,,['system-failure-prediction'],True,47.0,10.1016/j.jss.2012.06.025,,,['recall']
922,Machine learning approaches for predicting software maintainability: a fuzzy-based transparent model,"['Moataz A. Ahmed', 'Hamdi A. Al-Jamimi']",2013,"Software quality is one of the most important factors for assessing the global competitive position of any software company. Thus, the quantification of the quality parameters and integrating them into the quality models is very essential.Many attempts have been made to precisely quantify the software quality parameters using various models such as Boehm's Model, McCall's Model and ISO/IEC 9126 Quality Model. A major challenge, although, is that effective quality models should consider two types of knowledge: imprecise linguistic knowledge from the experts and precise numerical knowledge from historical data.Incorporating the experts' knowledge poses a constraint on the quality model; the model has to be transparent.In this study, the authorspropose a process for developing fuzzy logic-based transparent quality prediction models.They applied the process to a case study where Mamdani fuzzy inference engine is used to predict software maintainability.Theycompared the Mamdani-based model with other machine learning approaches.The resultsshow that the Mamdani-based model is superior to all.",https://ieeexplore.ieee.org/document/6680574,True,,27.0,,,,['failure-prediction'],,,True,['failure-management'],,,,,,,,,
923,Failure Analysis of Jobs in Compute Clouds: A Google Cluster Case Study,"['Xin Chen', 'Charng-Da Lu', 'Karthik Pattabiraman']",2014,"In this paper, we analyze a workload trace from the Google cloud cluster and characterize the observed failures. The goal of our work is to improve the understanding of failures in compute clouds. We present the statistical properties of job and task failures, and attempt to correlate them with key scheduling constraints, node operations, and attributes of users in the cloud. We also explore the potential for early failure prediction, and anomaly detection for the jobs. Based on our results, we speculate that there are many opportunities to enhance the reliability of the applications running in the cloud, such as pro-active maintenance of nodes or limiting job resubmissions. We further find that resource usage patterns of the jobs can be leveraged by failure prediction techniques. Finally, we find that the termination statuses of jobs and tasks can be clustered into six dominant categories based on the user profiles.",https://ieeexplore.ieee.org/abstract/document/6982624,True,,59.0,['clustering'],['host-metrics'],['discussion'],['failure-prediction'],"['cluster', 'job']",2014 IEEE 25th International Symposium on Software Reliability Engineering,True,['failure-management'],True,,['system-failure-prediction'],,,,,,
924,Packing light: Portable workload performance prediction for the cloud,"['Jennie Duggan', 'Yun Chi', 'Hakan Hacigümüş', 'Shenghuo Zhu', 'Ugur Çetintemel']",2013,"We introduce a new learning-based solution for portable database workload performance prediction. The current state of the art addresses performance prediction for individual, static hardware configurations and thus cannot generalize to new platforms without additional training. In this work, we focus on analytical databases that might be deployed on different hardware configurations, possibly offered by various Infrastructure-as-a-Service (IaaS) providers in the cloud. Enabling workload performance predictions that can be ported across hardware configurations and IaaS offerings could significantly help cloud users with their service-purchase decisions and cloud providers with their provisioning decisions. Our solution is based on collaborative filtering modeling and prediction. We applied it to lightweight workload fingerprints that model the characteristics and behavior of concurrent query workloads for carefully selected, abstract hardware configurations. Our preliminary results are derived from experiments with TPC-H and TPC-DS benchmarks on the Amazon and Rackspace clouds. They demonstrate that our techniques can predict analytical workload throughput values for diverse hardware platforms with low training overhead and within approximately 30% of the correct figure.",https://ieeexplore.ieee.org/document/6547460,True,,23.0,,,,,,,True,['resource-provisioning'],,,,,,,,,
925,Prediction-Based Dynamic Resource Allocation for Video Transcoding in Cloud Computing,"['Fareed Jokhio', 'Adnan Ashraf', 'Sébastien Lafond', 'Ivan Porres', 'Johan Lilius']",2013,"This paper presents prediction-based dynamic resource allocation algorithms to scale video transcoding service on a given Infrastructure as a Service cloud. The proposed algorithms provide mechanisms for allocation and deallocation of virtual machines (VMs) to a cluster of video transcoding servers in a horizontal fashion. We use a two-step load prediction method, which allows proactive resource allocation with high prediction accuracy under real-time constraints. For cost-efficiency, our work supports transcoding of multiple on-demand video streams concurrently on a single VM, resulting in a reduced number of required VMs. We use video segmentation at group of pictures level, which splits video streams into smaller segments that can be transcoded independently of one another. The approach is demonstrated in a discrete-event simulation and an experimental evaluation involving two different load patterns.",https://ieeexplore.ieee.org/document/6498561,True,,52.0,,,,"['workload-prediction', 'resource-consolidation']",,"2013 21st Euromicro International Conference on Parallel, Distributed, and Network-Based Processing",True,['resource-provisioning'],,,,,,,,,
926,Vigilant: out-of-band detection of failures in virtual machines,"['Dan Pelleg', 'Muli Ben-Yehuda', 'Rick Harper', 'Lisa Spainhower', 'Tokunbo Adeshiyan']",2008,"What do our computer systems do all day? How do we make sure they continue doing it when failures occur? Traditional approaches to answering these questions often involve in-band monitoring agents. However in-band agents suffer from several drawbacks: they need to be written or customized for every workload (operating system and possibly also application), they comprise potential security liabilities, and are themselves affected by adverse conditions in the monitored systems.Virtualization technology makes it possible to encapsulate an entire operating system or application instance within a virtual object that can then be easily monitored and manipulated without any knowledge of the contents or behavior of that object. This can be done out-of-band, using general purpose agents that do not reside inside the object, and hence are not affected by the behavior of the object.This paper describes Vigilant, a novel way of monitoring virtual machines for problems. Vigilant requires no specialized agents inside a virtual object it is monitoring. Instead, it uses the hypervisor to directly monitor the resource requests and utilization of an object. Machine learning methods are then used to analyze the readings. Our experimental results show that problems can be detected out-of-band with high accuracy. Using Vigilant we demonstrate that out-of-band monitoring using virtualization and machine learning can accurately identify faults in the guest OS, while avoiding the many pitfalls associated with in-band monitoring.",https://dl.acm.org/doi/10.1145/1341312.1341319,True,,37.0,,,,['failure-detection'],,ACM SIGOPS Operating Systems Review,True,['failure-management'],,,,,,10.1145/1341312.1341319,,,
927,Modeling the Autoscaling Operations in Cloud with Time Series Data,"['Mehran N. A. H. Khan', 'Yan Liu', 'Hanieh Alipour', 'Samneet Singh']",2015,"Autoscaling involves complex cloud operations that automate the provisioning and de-provisioning of cloud resources to support continuous development of customer services. Autoscaling depends on a number of decisions derived by aggregating metrics at the infrastructure and the platform level. In this paper, we review existing autoscaling techniques deployed in leading cloud providers. We identify core features and entities of the autoscaling operations as variables. We model these variables that quantify the interactions between these entities and incorporate workload time series data to calibrate the model. Hence the model allows proactive analysis of workload patterns and estimation of the responsiveness of the autoscaling operations. We demonstrate the use of this model with Google cluster trace data.",https://ieeexplore.ieee.org/document/7371434,True,,9.0,,,,,,,True,['resource-provisioning'],,,,,,,,,
928,Using Machine Learning Algorithms for Cloud Client Prediction Models in a Web VM Resource Provisioning Environment,"['Ajila, Samuel Adesoye', 'Bankole, Akindele A']",2016,"In order to meet Service Level Agreement (SLA) requirements, efficient scaling of Virtual Machine (VM) resources in cloud computing needs to be provisioned ahead due to the instantiation time required by the VM. One way to do this is by predicting future resource demands. The existing research on VM resource provisioning are either reactive in their approach or use only non-business level metrics. In this research, a Cloud client prediction model for TPC-W benchmark web application is developed and evaluated using three …",https://journals.scholarpublishing.org/index.php/TMLAI/article/view/1690,True,,13.0,"['multilayer-perceptron', 'linear-regression', 'support-vector-machine']","['sla', 'configuration']","['comparison', 'novel-use']",['workload-prediction'],"['vm', 'ec2']",,True,['resource-provisioning'],,,,,,,,,
929,"Gandalf: An Intelligent, End-To-End Analytics Service for Safe Deployment in Cloud-Scale Infrastructure","['Li, Ze', 'Cheng, Qian', 'Hsieh, Ken', 'Dang, Yingnong', 'Huang, Peng', 'Singh, Pankaj', 'Yang, Xinsheng', 'Lin, Qingwei', 'Wu, Youjiang', 'Levy, Sebastien', 'others']",2020,"Modern cloud systems have a vast number of components that continuously undergo updates. Deploying these frequent updates quickly without breaking the system is challenging. In this paper, we present Gandalf, an end-to-end analytics service for safe deployment in a large-scale system infrastructure. Gandalf enables rapid and robust impact assessment of software rollouts to catch bad rollouts before they cause widespread outages. Gandalf monitors and analyzes various fault signals. It will correlate each signal against all …",https://www.semanticscholar.org/paper/Gandalf%3A-An-Intelligent%2C-End-To-End-Analytics-for-Li-Cheng/42b8082b06dd29875d6649c686e7e2dd9bc4850f,True,,0.0,"['correlation', 'clustering', 'logistic-regression']","['events', 'kpis', 'structure']",['new-method'],['failure-detection'],"['vm', 'node']",17th $\{$USENIX$\}$ Symposium on Networked Systems Design and Implementation ($\{$NSDI$\}$ 20),True,['failure-management'],,,['anomaly-detection'],,,,,,
930,Convolutional neural networks for unsupervised anomaly detection in text data,"['Gorokhov, Oleg', 'Petrovskiy, Mikhail', 'Mashechkin, Igor']",2017,"In this paper, we discuss the problem of anomaly detection in text data using convolutional neural network (CNN). Recently CNNs have become one of the most popular and powerful tools for various machine learning tasks. CNN's main advantage is an ability to extract complicated hidden features from high dimensional data with complex structure. Usually CNNs are applied in supervised learning mode. On the other hand, unsupervised anomaly detection is an important problem in many applications, including computer security …",https://link.springer.com/chapter/10.1007/978-3-319-68935-7_54,True,,7.0,,,,['failure-detection'],,International Conference on Intelligent Data Engineering and Automated Learning,True,['failure-management'],,,,,,,,,
931,Prediction of software development faults in PL/SQL files using neural network models,"['Quah, Tong-Seng', 'Thwin, Mie Mie Thet']",2004,"Database application constitutes one of the largest and most important software domains in the world. Some classes or modules in such applications are responsible for database operations. Structured Query Language (SQL) is used to communicate with database middleware in these classes or modules. It can be issued interactively or embedded in a host language. This paper aims to predict the software development faults in PL/SQL files using SQL metrics. Based on actual project defect data, the SQL metrics are empirically …",https://www.researchgate.net/publication/222300032_Prediction_of_software_development_faults_in_PLSQL_files_using_neural_network_models,True,,22.0,,,,['failure-prediction'],,,True,['failure-management'],,,,,,,,,
932,Neural-network-based approaches for software reliability estimation using dynamic weighted combinational models,"['Su, Yu-Shen', 'Huang, Chin-Yu']",2007,"Software reliability is the probability of failure-free software operation for a specified period of time in a specified environment. During the last three decades, many software reliability growth models (SRGMs) have been proposed and analyzed for measuring software reliability growth. SRGMs are mathematical models that represent software failures as a random process and can be used to evaluate development status during testing. However, most of SRGMs depend on some assumptions or distributions. In this paper, we propose an …",https://www.sciencedirect.com/science/article/pii/S0164121206001737?via%3Dihub,True,,181.0,,,,['failure-prevention'],['software'],,True,['failure-management'],,,['software-defect-prediction'],,,,,,
933,Prediction of software reliability using connectionist models,"['N. Karunanithi', 'D. Whitley', 'Y.K. Malaiya']",1992,"The usefulness of connectionist models for software reliability growth prediction is illustrated. The applicability of the connectionist approach is explored using various network models, training regimes, and data representation methods. An empirical comparison is made between this approach and five well-known software reliability growth models using actual data sets from several different software projects. The results presented suggest that connectionist models may adapt well across different data sets and exhibit a better predictive accuracy. The analysis shows that the connectionist approach is capable of developing models of varying complexity.<
>",https://ieeexplore.ieee.org/document/148475,True,,272.0,['multilayer-perceptron'],,"['comparison', 'novel-use']",['failure-prevention'],['software'],,True,['failure-management'],,,['software-defect-prediction'],,,,,,
934,Evolutionary neural network modeling for software cumulative failure time prediction,"['Tian, Liang', 'Noore, Afzel']",2005,An evolutionary neural network modeling approach for software cumulative failure time prediction based on multiple-delayed-input single-output architecture is proposed. Genetic algorithm is used to globally optimize the number of the delayed input neurons and the number of neurons in the hidden layer of the neural network architecture. Modification of Levenberg–Marquardt algorithm with Bayesian regularization is used to improve the ability to predict software cumulative failure time. The performance of our proposed approach has been compared using real-time control and flight dynamic application data sets. Numerical results show that both the goodness-of-fit and the next-step-predictability of our proposed approach have greater accuracy in predicting software cumulative failure time compared to existing approaches,https://www.sciencedirect.com/science/article/pii/S0951832004000857,True,,167.0,"['multilayer-perceptron', 'genetic-programming']",,['new-method'],['failure-prediction'],,,True,['failure-management'],True,,['system-failure-prediction'],,,,,,
935,On-line prediction of software reliability using an evolutionary connectionist model,"['Tian, Liang', 'Noore, Afzel']",2004,"An on-line adaptive software reliability prediction model using evolutionary connectionist approach based on multiple-delayed-input single-output architecture is proposed. Based on the currently available software failure time data, genetic algorithm is used to globally optimize the number of the delayed input neurons and the number of neurons in the hidden layer of the neural network architecture. Bayesian regularization is applied to our network training scheme to improve the generalization capability. The corresponding optimized …",https://www.sciencedirect.com/science/article/pii/S016412120400144X,True,,103.0,"['genetic-programming', 'multilayer-perceptron']",['events'],['new-method'],['failure-prediction'],['software'],,True,['failure-management'],,,['system-failure-prediction'],,,,,,
936,Automated Identification of Failure Causes in System Logs,"['Leonardo Mariani', 'Fabrizio Pastore']",2008,"Log files are commonly inspected by system administrators and developers to detect suspicious behaviors and diagnose failure causes. Since size of log files grows fast, thus making manual analysis impractical, different automatic techniques have been proposed to analyze log files. Unfortunately, accuracy and effectiveness of these techniques are often limited by the unstructured nature of logged messages and the variety of data that can be logged.This paper presents a technique to automatically analyze log files and retrieve important information to identify failure causes. The technique automatically identifies dependencies between events and values in logs corresponding to legal executions, generates models of legal behaviors and compares log files collected during failing executions with the generated models to detect anomalous event sequences that are presented to users. Experimental results show the effectiveness of the technique in supporting developers and testers to identify failure causes.",https://ieeexplore.ieee.org/document/4700316,True,,113.0,['automaton'],['logs'],,['failure-detection'],,,True,['failure-management'],,,['anomaly-detection'],,,,,,
937,Mining invariants from console logs for system problem detection,"['Jian-Guang Lou', 'Qiang Fu', 'Shengqi Yang', 'Ye Xu', 'Jiang Li']",2010,"Detecting execution anomalies is very important to the maintenance and monitoring of large-scale distributed systems. People often use console logs that are produced by distributed systems for troubleshooting and problem diagnosis. However, manually inspecting console logs for the detection of anomalies is unfeasible due to the increasing scale and complexity of distributed systems. Therefore, there is great demand for automatic anomaly detection techniques based on log analysis. In this paper, we propose an unstructured log analysis technique for anomaly detection, with a novel algorithm to automatically discover program invariants in logs. At first, a log parser is used to convert the unstructured logs to structured logs. Then, the structured log messages are further grouped to log message groups according to the relationship among log parameters. After that, the program invariants are automatically mined from the log message groups. The mined invariants can reveal the inherent linear characteristics of program work flows. With these learned invariants, our technique can automatically detect anomalies in logs. Experiments on Hadoop show that the technique can effectively detect execution anomalies. Compared with the state of art, our approach can not only detect numerous real problems with high accuracy but also provide intuitive insight into the problems.",https://dl.acm.org/doi/10.5555/1855840.1855864,True,,134.0,,['logs'],,['failure-detection'],,USENIXATC'10: Proceedings of the 2010 USENIX conference on USENIX annual technical conference,True,['failure-management'],,,['anomaly-detection'],,,10.5555/1855840.1855864,,,
938,Mining anomalies using traffic feature distributions,"['Anukool Lakhina', 'Mark Crovella', 'Christophe Diot']",2005,"The increasing practicality of large-scale flow capture makes it possible to conceive of traffic analysis methods that detect and identify a large and diverse set of anomalies. However the challenge of effectively analyzing this massive data source for anomaly diagnosis is as yet unmet. We argue that the distributions of packet features (IP addresses and ports) observed in flow traces reveals both the presence and the structure of a wide range of anomalies. Using entropy as a summarization tool, we show that the analysis of feature distributions leads to significant advances on two fronts: (1) it enables highly sensitive detection of a wide range of anomalies, augmenting detections by volume-based methods, and (2) it enables automatic classification of anomalies via unsupervised learning. We show that using feature distributions, anomalies naturally fall into distinct and meaningful clusters. These clusters can be used to automatically classify anomalies and to uncover new anomaly types. We validate our claims on data from two backbone networks (Abilene and Geant) and conclude that feature distributions show promise as a key element of a fairly general network anomaly diagnosis framework.",https://dl.acm.org/doi/10.1145/1080091.1080118,True,,1481.0,"['dimensionality-reduction', 'clustering']","['packet-content', 'network-metrics']",['new-method'],['failure-detection'],['network'],"SIGCOMM '05: Proceedings of the 2005 conference on Applications, technologies, architectures, and protocols for computer communications",True,['failure-management'],True,,['anomaly-detection'],True,55.0,10.1145/1080091.1080118,,,
939,Unsupervised Traffic Flow Classification Using a Neural Autoencoder,"['Jonas Höchst', 'Lars Baumgärtner', 'Matthias Hollick', 'Bernd Freisleben']",2017,"To cope with the varying delay and bandwidth requirements of today's mobile applications, mobile wireless networks can profit from classifying and predicting mobile application traffic. State-of-the-art traffic classification approaches have various disadvantages: port-based classification methods can be circumvented by choosing non-standard ports, protocol fingerprinting can be confused by the use of encryption, and current supervised learning methods for analyzing the statistical properties of network flows try to detect predefined classes, such as e-mail or FTP traffic, learned during training. In this paper, we present a novel approach to unsupervised traffic flow classification using statistical properties of flows and clustering based on a neural auto encoder. A novel time interval based feature vector construction and a semi-automatic cluster labeling method facilitate traffic flow classification independent of known traffic classes. An experimental evaluation on real data captured over a period of four months is presented. The obtained results show that 7 different classes of mobile traffic flows are detected with an average precision of 80% and an average recall of 75%.",https://ieeexplore.ieee.org/document/8109399,True,,16.0,"['multilayer-perceptron', 'autoencoder', 'clustering']",,,['failure-detection'],,2017 IEEE 42nd Conference on Local Computer Networks (LCN),True,['failure-management'],,,['traffic-classification'],,,,,,
940,Support Vector Machines for TCP traffic classification,"['Este, Alice', 'Gringoli, Francesco', 'Salgarelli, Luca']",2009,"Support Vector Machines (SVM) represent one of the most promising Machine Learning (ML) tools that can be applied to the problem of traffic classification in IP networks. In the case of SVMs, there are still open questions that need to be addressed before they can be generally applied to traffic classifiers. Having being designed essentially as techniques for binary classification, their generalization to multi-class problems is still under research. Furthermore, their performance is highly susceptible to the correct optimization of …",https://www.sciencedirect.com/science/article/abs/pii/S1389128609001649,True,,316.0,['support-vector-machine'],,['novel-use'],['failure-detection'],['network'],,True,['failure-management'],True,,['traffic-classification'],True,66.0,,,,
941,A method for classification of network traffic based on C5.0 Machine Learning Algorithm,"['Tomasz Bujlow', 'Tahir Riaz', 'Jens Myrup Pedersen']",2012,"Monitoring of the network performance in highspeed Internet infrastructure is a challenging task, as the requirements for the given quality level are service-dependent. Backbone QoS monitoring and analysis in Multi-hop Networks requires therefore knowledge about types of applications forming current network traffic. To overcome the drawbacks of existing methods for traffic classification, usage of C5.0 Machine Learning Algorithm (MLA) was proposed. On the basis of statistical traffic information received from volunteers and C5.0 algorithm we constructed a boosted classifier, which was shown to have ability to distinguish between 7 different applications in test set of 76,632-1,622,710 unknown cases with average accuracy of 99.3-99.9%. This high accuracy was achieved by using high quality training data collected by our system, a unique set of parameters used for both training and classification, an algorithm for recognizing flow direction and the C5.0 itself. Classified applications include Skype, FTP, torrent, web browser traffic, web radio, interactive gaming and SSH. We performed subsequent tries using different sets of parameters and both training and classification options. This paper shows how we collected accurate traffic data, presents arguments used in classification process, introduces the C5.0 classifier and its options, and finally evaluates and compares the obtained results.",https://ieeexplore.ieee.org/document/6167418,True,,125.0,['decision-tree'],"['packet-content', 'network-metrics']",['novel-use'],['failure-detection'],"['web', 'network']","2012 international conference on computing, networking and communications (ICNC)",True,['failure-management'],,,['traffic-classification'],,,,,,
942,Internet traffic classification using bayesian analysis techniques,"['Andrew W. Moore', 'Denis Zuev']",2005,"Accurate traffic classification is of fundamental importance to numerous other network activities, from security monitoring to accounting, and from Quality of Service to providing operators with useful forecasts for long-term provisioning. We apply a Naïve Bayes estimator to categorize traffic by application. Uniquely, our work capitalizes on hand-classified network data, using it as input to a supervised Naïve Bayes estimator. In this paper we illustrate the high level of accuracy achievable with the \Naive Bayes estimator. We further illustrate the improved accuracy of refined variants of this estimator.Our results indicate that with the simplest of Naïve Bayes estimator we are able to achieve about 65% accuracy on per-flow classification and with two powerful refinements we can improve this value to better than 95%; this is a vast improvement over traditional techniques that achieve 50--70%. While our technique uses training data, with categories derived from packet-content, all of our training and testing was done using header-derived discriminators. We emphasize this as a powerful aspect of our approach: using samples of well-known traffic to allow the categorization of traffic using commonly available information alone.",https://dl.acm.org/doi/10.1145/1064212.1064220,True,,1633.0,"['naive-bayes', 'entropy-selection']",['network-metrics'],['novel-use'],['failure-detection'],['network'],SIGMETRICS '05: Proceedings of the 2005 ACM SIGMETRICS international conference on Measurement and modeling of computer systems,True,['failure-management'],True,,['traffic-classification'],True,64.0,10.1145/1064212.1064220,,,
943,Experience Report: System Log Analysis for Anomaly Detection,"['Shilin He', 'Jieming Zhu', 'Pinjia He', 'Michael R. Lyu']",2016,"Anomaly detection plays an important role in management of modern large-scale distributed systems. Logs, which record system runtime information, are widely used for anomaly detection. Traditionally, developers (or operators) often inspect the logs manually with keyword search and rule matching. The increasing scale and complexity of modern systems, however, make the volume of logs explode, which renders the infeasibility of manual inspection. To reduce manual effort, many anomaly detection methods based on automated log analysis are proposed. However, developers may still have no idea which anomaly detection methods they should adopt, because there is a lack of a review and comparison among these anomaly detection methods. Moreover, even if developers decide to employ an anomaly detection method, re-implementation requires a nontrivial effort. To address these problems, we provide a detailed review and evaluation of six state-of-the-art log-based anomaly detection methods, including three supervised methods and three unsupervised methods, and also release an open-source toolkit allowing ease of reuse. These methods have been evaluated on two publicly-available production log datasets, with a total of 15,923,592 log messages and 365,298 anomaly instances. We believe that our work, with the evaluation results as well as the corresponding findings, can provide guidelines for adoption of these methods and provide references for future development.",https://ieeexplore.ieee.org/abstract/document/7774521,True,,103.0,"['logistic-regression', 'rule-mining', 'decision-tree', 'dimensionality-reduction', 'clustering']",['logs'],"['discussion', 'comparison']",['failure-detection'],,,True,['failure-management'],,,['anomaly-detection'],,,,,,
944,Web traffic anomaly detection using C-LSTM neural networks,"['Kim, Tae-Young', 'Cho, Sung-Bae']",2018,"Web traffic refers to the amount of data that is sent and received by people visiting online websites. Web traffic anomalies represent abnormal changes in time series traffic, and it is important to perform detection quickly and accurately for the efficient operation of complex computer networks systems. In this paper, we propose a C-LSTM neural network for effectively modeling the spatial and temporal information contained in traffic data, which is a one-dimensional time series signal. We also provide a method for automatically extracting robust features of spatial-temporal information from raw data. Experiments demonstrate that our C-LSTM method can extract more complex features by combining a convolutional neural network (CNN), long short-term memory (LSTM), and deep neural network (DNN). The CNN layer is used to reduce the frequency variation in spatial information; the LSTM layer is suitable for modeling time information; and the DNN layer is used to map data into a more separable space. Our C-LSTM method also achieves nearly perfect anomaly detection performance for web traffic data, even for very similar signals that were previously considered to be very difficult to classify. Finally, the C-LSTM method outperforms other state-of-the-art machine learning techniques on Yahoo's well-known Webscope S5 dataset, achieving an overall accuracy of 98.6% and recall of 89.7% on the test dataset.",https://www.sciencedirect.com/science/article/abs/pii/S0957417418302288,True,,41.0,"['rnn', 'cnn']",,,['failure-detection'],['network'],,True,['failure-management'],True,,['traffic-classification'],,,,,,
945,Traffic Anomaly Detection Using KMeans Clustering,"['M{\\""u}nz, Gerhard', 'Li, Sa', 'Carle, Georg']",2007,"Data mining techniques make it possible to search large amounts of data for characteristic rules and patterns. If applied to network monitoring data recorded on a host or in a network, they can be used to detect intrusions, attacks and/or anomalies. This paper gives an introduction to Network Data Mining, ie the application of data mining methods to packet and flow data captured in a network, including a comparative overview of existing approaches. Furthermore, we present a novel flow-based anomaly detection scheme based on the K …",http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.323.6870,True,,17.0,,,,['failure-detection'],,GI/ITG Workshop MMBnet,True,['failure-management'],,,,,,,,,
946,Metric selection and anomaly detection for cloud operations using log and metric correlation analysis,"['Farshchi, Mostafa', 'Schneider, Jean-Guy', 'Weber, Ingo', 'Grundy, John']",2018,"Cloud computing systems provide the facilities to make application services resilient against failures of individual computing resources. However, resiliency is typically limited by a cloud consumer's use and operation of cloud resources. In particular, system operations have been reported as one of the leading causes of system-wide outages. This applies specifically to DevOps operations, such as backup, redeployment, upgrade, customized scaling, and migration–which are executed at much higher frequencies now than a decade …",https://www.sciencedirect.com/science/article/pii/S0164121217300596,True,,27.0,,['logs'],['new-method'],['failure-detection'],,,True,['failure-management'],,,['anomaly-detection'],,,,,,
947,Experience Report: Log-Based Behavioral Differencing,"['Maayan Goldstein', 'Danny Raz', 'Itai Segall']",2017,"Monitoring systems and ensuring the required service level is an important operation task. However, doing this based on external visible data, such as systems logs, is very difficult since it is very hard to extract from the logged data the exact state and the root cause to the actions taken by the system. Yet, identifying behavioral changes of complex systems can be used for early identification of problems and allow proactive correction measurements. Since it is practically impossible to perform this task manually, there is a critical need for a methodology that can analyze logs, automatically create a behavioral model, and compare the behavior to the expected behavior.In this paper we propose a novel approach for comparison between serviceexecutions as exhibited in their log files. The behavior is captured by FiniteState Automaton models (FSAs), enhanced with performance related data, bothmined from the logs. Our tool then computes the difference between the current model and behavioral models created when the service was known to operate well. A visual framework that graphically presents and emphasizes the changes in the behavior is then used to trace their root cause. We evaluate our approach over real telecommunication logs.",https://ieeexplore.ieee.org/document/8109094,True,,13.0,['automaton'],['logs'],['novel-use'],['failure-detection'],,,True,['failure-management'],,,['anomaly-detection'],,,,,,
948,Failure Diagnosis for Distributed Systems Using Targeted Fault Injection,"['Cuong Pham', 'Long Wang', 'Byung Chul Tak', 'Salman Baset', 'Chunqiang Tang', 'Zbigniew Kalbarczyk', 'Ravishankar K. Iyer']",2016,"This paper introduces a novel approach to automating failure diagnostics in distributed systems by combining fault injection and data analytics. We use fault injection to populate the database of failures for a target distributed system. When a failure is reported from production environment, the database is queried to find “matched” failures generated by fault injections. Relying on the assumption that similar faults generate similar failures, we use information from the matched failures as hints to locate the actual root cause of the reported failures. In order to implement this approach, we introduce techniques for (i) reconstructing end-to-end execution flows of distributed software components, (ii) computing the similarity of the reconstructed flows, and (iii) performing precise fault injection at pre-specified executing points in distributed systems. We have evaluated our approach using an OpenStack cloud platform, a popular cloud infrastructure management system. Our experimental results showed that this approach is effective in determining the root causes, e.g., fault types and affected components, for 71-100 percent of tested failures. Furthermore, it can provide fault locations close to actual ones and can easily be used to find and fix actual root causes. We have also validated this technique by localizing real bugs that occurred in OpenStack.",https://ieeexplore.ieee.org/document/7484300,True,,16.0,,,,['root-cause-analysis'],,,True,['failure-management'],,,,,,,,,
949,Structural Event Detection from Log Messages,"['Fei Wu', 'Pranay Anchuri', 'Zhenhui Li']",2017,"A wide range of modern web applications are only possible because of the composable nature of the web services they are built upon. It is, therefore, often critical to ensure proper functioning of these web services. As often, the server-side of web services is not directly accessible, several log message based analysis have been developed to monitor the status of web services. Existing techniques focus on using clusters of messages (log patterns) to detect important system events. We argue that meaningful system events are often representable by groups of cohesive log messages and the relationships among these groups. We propose a novel method to mine structural events as directed workflow graphs (where nodes represent log patterns, and edges represent relations among patterns). The structural events are inclusive and correspond to interpretable episodes in the system. The problem is non-trivial due to the nature of log data: (i) Individual log messages contain limited information, and (ii) Log messages in a large scale web system are often interleaved even though the log messages from individual components are ordered. As a result, the patterns and relationships mined directly from the messages and their ordering can be erroneous and unreliable in practice. Our solution is based on the observation that meaningful log patterns and relations often form workflow structures that are connected. Our method directly models the overall quality of structural events. Through both qualitative and quantitative experiments on real world datasets, we demonstrate the effectiveness and the expressiveness of our event detection method.",https://dl.acm.org/doi/10.1145/3097983.3098124,True,,12.0,,,,['failure-detection'],,KDD '17: Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,True,['failure-management'],,,,,,10.1145/3097983.3098124,,,
950,Log20: Fully Automated Optimal Placement of Log Printing Statements under Specified Overhead Threshold,"['Xu Zhao', 'Kirk Rodrigues', 'Yu Luo', 'Michael Stumm', 'Ding Yuan', 'Yuanyuan Zhou']",2017,"When systems fail in production environments, log data is often the only information available to programmers for postmortem debugging. Consequently, programmers' decision on where to place a log printing statement is of crucial importance, as it directly affects how effective and efficient postmortem debugging can be. This paper presents Log20, a tool that determines a near optimal placement of log printing statements under the constraint of adding less than a specified amount of performance overhead. Log20 does this in an automated way without any human involvement. Guided by information theory, the core of our algorithm measures how effective each log printing statement is in disambiguating code paths. To do so, it uses the frequencies of different execution paths that are collected from a production environment by a low-overhead tracing library. We evaluated Log20 on HDFS, HBase, Cassandra, and ZooKeeper, and observed that Log20 is substantially more efficient in code path disambiguation compared to the developers' manually placed log printing statements. Log20 can also output a curve showing the trade-off between the informativeness of the logs and the performance slowdown, so that a developer can choose the right balance.",https://dl.acm.org/doi/10.1145/3132747.3132778,True,,38.0,"['entropy-selection', 'optimization']",['logs'],['new-method'],['failure-detection'],"['hadoop', 'yarn']",SOSP '17: Proceedings of the 26th Symposium on Operating Systems Principles,True,['failure-management'],True,,['log-enhancement'],True,69.0,10.1145/3132747.3132778,,,
951,A Petri net approach to fault detection and diagnosis in distributed systems (Parts 1 and 2),"['Aghasaryan, Armen', 'Fabre, E', 'Benveniste, Albert', 'Boubour, R', 'Jard, C']",1997,"For pt. I see ibid., p. 720-5 (1997). We present an original construction of stochastic Petri nets (PN) dedicated to large distributed discrete event systems. Its main characteristic is to provide statistically independent behaviors to concurrent (parallel) processes of the system. We end up with"" hybrid"" model where only some events are randomized, and that can't be described by a standard Markov dynamics. Equivalently, time is only partially ordered in such systems. Then assuming that every fired transition produces a random label we …",https://www.researchgate.net/publication/2299302_A_Petri_net_approach_to_fault_detection_and_diagnosis_in_distributed_systems_Parts_1_and_2,True,,32.0,,,,['failure-detection'],,Proceedings of the 36th IEEE Conference on Decision and Control,True,['failure-management'],,,,,,,,,
952,Dynamic Anomalography: Tracking Network Anomalies Via Sparsity and Low Rank,"['Morteza Mardani', 'Gonzalo Mateos', 'Georgios B. Giannakis']",2012,"In the backbone of large-scale networks, origin-to-destination (OD) traffic flows experience abrupt unusual changes known as traffic volume anomalies, which can result in congestion and limit the extent to which end-user quality of service requirements are met. As a means of maintaining seamless end-user experience in dynamic environments, as well as for ensuring network security, this paper deals with a crucial network monitoring task termed dynamic anomalography. Given link traffic measurements (noisy superpositions of unobserved OD flows) periodically acquired by backbone routers, the goal is to construct an estimated map of anomalies in real time, and thus summarize the network `health state' along both the flow and time dimensions. Leveraging the low intrinsic-dimensionality of OD flows and the sparse nature of anomalies, a novel online estimator is proposed based on an exponentially-weighted least-squares criterion regularized with the sparsity-promoting
l
1
-norm of the anomalies, and the nuclear norm of the nominal traffic matrix. After recasting the non-separable nuclear norm into a form amenable to online optimization, a real-time algorithm for dynamic anomalography is developed and its convergence established under simplifying technical assumptions. For operational conditions where computational complexity reductions are at a premium, a lightweight stochastic gradient algorithm based on Nesterov's acceleration technique is developed as well. Comprehensive numerical tests with both synthetic and real network data corroborate the effectiveness of the proposed online algorithms and their tracking capabilities, and demonstrate that they outperform state-of-the-art approaches developed to diagnose traffic anomalies.",https://ieeexplore.ieee.org/abstract/document/6376091,True,,118.0,,,,['failure-detection'],['network'],,True,['failure-management'],,,['anomaly-detection'],,,,,,
953,Automated Classification of Network Traffic Anomalies,"['Fernandes, Guilherme', 'Owezarski, Philippe']",2009,"Network traffic anomalies detection and characterization has been a hot topic of research for many years. Although the field is very advanced in the detection of network traffic anomalies, accurate automated classification is still a very challenging and unmet problem. This paper presents a new algorithm for automated classification of network traffic anomalies. The algorithm relies on three steps:(i) after an anomaly has been detected, identify all (or most) related packets or flow records;(ii) use these packets or flow records to derive several …",https://link.springer.com/chapter/10.1007/978-3-642-05284-2_6,True,,41.0,,,,['failure-detection'],,International Conference on Security and Privacy in Communication Systems,True,['failure-management'],,,,,,,,,
954,A methodology for root-cause analysis in component based systems,"['Kui Wang', 'Carol Fung', 'Chao Ding', 'Polo Pei', 'Shaohan Huang', 'Zhongzhi Luan', 'Depei Qian']",2015,"In component based enterprise systems, anomaly detectors are commonly deployed on application-level components, but not on lower-level functional components. When anomaly alarms are triggered, system managers are expected to handle them in a timely manner to avoid cascading failures. Excessive large volume of anomaly alarms makes them impractical to handle manually. Most existing root cause analysis methods are based on the assumption that all components are monitored and analysis are performed based on the time correlation of the generated alarms. However, full monitoring coverage may not be practical due to cost and complexity. In this paper, we present RCSF, a root cause analysis method that targets at systems where only application-level components are monitored by anomaly detectors. The method analyzes the components performance log on functional components and seek for most probable fault propagation sequences based on anomaly analysis. We evaluate the RCSF method based on real enterprise system data and compare it with some baseline methods. Experimental results show that our proposed method can effectively anchor the root causes of failures by providing a short list of most probable causes, and the performance is significantly improved compared to the baseline methods.",https://ieeexplore.ieee.org/document/7404741?arnumber=7404741,True,,5.0,,,,['root-cause-analysis'],,2015 IEEE 23rd International Symposium on Quality of Service (IWQoS),True,['failure-management'],,,,,,,,,
955,HiLighter: Automatically Building Robust Signatures of Performance Behavior for Small- and Large-Scale Systems,"[""Bod{\\'\\i}k, Peter"", 'Goldszmidt, Moises', 'Fox, Armando']",2008,"Previous work showed that statistical analysis techniques could successfully be used to construct compact signatures of distinct operational problems in Internet server systems. Because signatures are amenable to well-known similarity search techniques, they can be used as a way to index past problems and identify particular operational problems as new or recurrent. In this paper we use a different statistical technique for constructing signatures (logistic regression with L1 regularization) that improves on previous work in two ways. First …",https://www.semanticscholar.org/paper/HiLighter%3A-Automatically-Building-Robust-Signatures-Bod%C3%ADk-Goldszmidt/f80fa53279c98b34093ae0f759dddb0d22bc52cf,True,,25.0,,,,['root-cause-analysis'],,SysML,True,['failure-management'],,,,,,,,,
956,Toward Self-Healing Multitier Services,"['Brian Cook', 'Shivnath Babu', 'George Candea', 'Songyun Duan']",2007,"Are self-heating database-centric multitier services Utopia or just a hard puzzle? We argue for the latter and aim to identify the missing pieces of this puzzle. We advocate robust and scalable learning-based approaches to self-healing that we expect to work well for a large class of multitier services. We identify performance-availability problems (PAPs) as the most relevant target for self-healing, and argue that PAPs are best addressed macroscopically. outside the realm of individual tiers. Finally, we lay out a research agenda for learning-based approaches to self-healing, to enable wider deployment of self-healing multi-tier services.",https://ieeexplore.ieee.org/document/4401025,True,,26.0,,,,['remediation'],,2007 IEEE 23rd International Conference on Data Engineering Workshop,True,['failure-management'],,,,,,,,,
957,Fingerpointing correlated failures in replicated systems,"['Soila Pertet', 'Rajeev Gandhi', 'Priya Narasimhan']",2007,"Replicated systems are often hosted over underlying group communication protocols that provide totally ordered, reliable delivery of messages. In the face of a performance problem at a single node, these protocols can cause correlated performance degradations at even non-faulty nodes, leading to potential red herrings in failure diagnosis. We propose a fingerpointing approach that combines node-level (local) anomaly detection, followed by system-wide (global) fingerpointing. The local anomaly detection relies on threshold-based analyses of system metrics, while global fingerpointing is based on the hypothesis that the root-cause of the failure is the node with an ""odd-man-out"" view of the anomalies. We compare the results of applying three classifiers - a heuristic algorithm, an unsupervised learner (k-means clustering), and a supervised learner (k-nearest-neighbor) - to finger-point the faulty node.",https://dl.acm.org/doi/abs/10.5555/1361442.1361451,True,,27.0,,,,['failure-detection'],,SYSML'07: Proceedings of the 2nd USENIX workshop on Tackling computer systems problems with machine learning techniques,True,['failure-management'],,,,,,10.5555/1361442.1361451,,,
958,Ensembles of models for automated diagnosis of system performance problems,"['S. Zhang', 'I. Cohen', 'M. Goldszmidt', 'J. Symons', 'A. Fox']",2005,"Violations of service level objectives (SLO) in Internet services are urgent conditions requiring immediate attention. Previously we explored (I. Cohen et al., 2004) an approach for identifying which low-level system properties were correlated to high-level SLO violations (the metric attribution problem). The approach is based on automatically inducing models from data using pattern recognition and probability modeling techniques. In this paper we extend our approach to adapt to changing workloads and external disturbances by maintaining an ensemble of probabilistic models, adding new models when existing ones do not accurately capture current system behavior. Using realistic workloads on an implemented prototype system, we show that the ensemble of models captures the performance behavior of the system accurately under changing workloads and conditions. We fuse information from the models in the ensemble to identify likely causes of the performance problem, with results comparable to those produced by an oracle that continuously changes the model based on advance knowledge of the workload. The cost of inducing new models and managing the ensembles is negligible, making our approach both immediately practical and theoretically appealing.",https://ieeexplore.ieee.org/document/1467838,True,,180.0,['bayesian-network'],"['software-metrics', 'kpis', 'host-metrics']",,['failure-detection'],,2005 International Conference on Dependable Systems and Networks (DSN'05),True,['failure-management'],,,['anomaly-detection'],,,,,,
959,Analysis of software rejuvenation using Markov Regenerative Stochastic Petri Net,"['S. Garg', 'A. Puliafito', 'M. Telek', 'K.S. Trivedi']",1995,"In a client-server type system, the server software is required to run continuously for very long periods. Due to repeated and potentially faulty usage by many clients, such software ""ages"" with time and eventually fails. (Huang et al., 1995) proposed a technique called ""software rejuvenation"" in which the software is periodically stopped and then restarted in a ""robust"" state after proper maintenance. This ""renewal"" of software prevents (or at least postpones) the crash failure. As the time lost (or the cost incurred) due to the software failure is typically more than the time lost (or the cost incurred) due to rejuvenation, the technique reduces the expected unavailability of the software. We present a quantitative analysis of software rejuvenation. The behavior of the system is represented through a Markov Regenerative Stochastic Petri Net (MRSPN) model which is solved both for steady state as well as transient conditions. We provide a closed-form analytical solution for the steady state expected down time (and the expected cost incurred) due to system unavailability. We also evaluate the optimal rejuvenation interval which minimizes the expected unavailability of the software.",https://ieeexplore.ieee.org/abstract/document/497656,True,,262.0,['petri-net'],,['new-method'],['failure-prevention'],,Proceedings of Sixth International Symposium on Software Reliability Engineering. ISSRE'95,True,['failure-management'],,,['rejuvenation'],,,,,,
960,Diagnosing network-wide traffic anomalies,"['Anukool Lakhina', 'Mark Crovella', 'Christophe Diot']",2004,"Anomalies are unusual and significant changes in a network's traffic levels, which can often span multiple links. Diagnosing anomalies is critical for both network operators and end users. It is a difficult problem because one must extract and interpret anomalous patterns from large amounts of high-dimensional, noisy data.In this paper we propose a general method to diagnose anomalies. This method is based on a separation of the high-dimensional space occupied by a set of network traffic measurements into disjoint subspaces corresponding to normal and anomalous network conditions. We show that this separation can be performed effectively by Principal Component Analysis.Using only simple traffic measurements from links, we study volume anomalies and show that the method can: (1) accurately detect when a volume anomaly is occurring; (2) correctly identify the underlying origin-destination (OD) flow which is the source of the anomaly; and (3) accurately estimate the amount of traffic involved in the anomalous OD flow.We evaluate the method's ability to diagnose (i.e., detect, identify, and quantify) both existing and synthetically injected volume anomalies in real traffic from two backbone networks. Our method consistently diagnoses the largest volume anomalies, and does so with a very low false alarm rate.",https://dl.acm.org/doi/10.1145/1015467.1015492,True,,1424.0,"['dimensionality-reduction', 'similarity-matching']",['network-metrics'],['new-method'],"['root-cause-analysis', 'failure-detection']",['network'],"SIGCOMM '04: Proceedings of the 2004 conference on Applications, technologies, architectures, and protocols for computer communications",True,['failure-management'],True,,"['root-cause-diagnosis', 'anomaly-detection']",True,54.0,10.1145/1015467.1015492,,,
961,On predictability of system anomalies in real world,"['Yongmin Tan', 'Xiaohui Gu']",2010,"As computer systems become increasingly complex, system anomalies have become major concerns in system management. In this paper, we present a comprehensive measurement study to quantify the predictability of different system anomalies. Online anomaly prediction allows the system to foresee impending anomalies so as to take proper actions to mitigate anomaly impact. Our anomaly prediction approach combines feature value prediction with statistical classification methods. We conduct extensive measurement study to investigate anomalous behavior of three systems in the real world: PlanetLab, SMART hard drive data, and IBM System S. We observe that real world system anomalies do exhibit predictability, which can be predicted with high accuracy and significant lead time.",https://ieeexplore.ieee.org/document/5581600,True,,35.0,,,,['failure-detection'],,"2010 IEEE International Symposium on Modeling, Analysis and Simulation of Computer and Telecommunication Systems",True,['failure-management'],,,,,,,,,
962,Experimental evaluation of n-tier systems: Observation and analysis of multi-bottleneck,"['Simon Malkowski', 'Markus Hedwig', 'Calton Pu']",2009,"In many areas such as e-commerce, mission-critical N-tier applications have grown increasingly complex. They are characterized by non-stationary workloads (e.g., peak load several times the sustained load) and complex dependencies among the component servers. We have studied N-tier applications through a large number of experiments using the RUBiS and RUBBoS benchmarks. We apply statistical methods such as kernel density estimation, adaptive filtering, and change detection through multiple-model hypothesis tests to analyze more than 200 GB of recorded data. Beyond the usual single-bottlenecks, we have observed more intricate bottleneck phenomena. For instance, in several configurations all system components show average resource utilization significantly below saturation, but overall throughput is limited despite addition of more resources. More concretely, our analysis shows experimental evidence of multi-bottleneck cases with low average resource utilization where several resources saturate alternatively, indicating a clear lack of independence in their utilization. Our data corroborates the increasing awareness of the need for more sophisticated analytical performance models to describe N-tier applications that do not rely on independent resource utilization assumptions. We also present a preliminary taxonomy of multi-bottlenecks found in our experimentally observed data.",https://ieeexplore.ieee.org/document/5306791,True,,62.0,,,,['failure-detection'],,2009 IEEE International Symposium on Workload Characterization (IISWC),True,['failure-management'],,,,,,,,,
963,Online performance anomaly prediction and prevention for complex distributed systems,"['Tan, Yongmin', 'others']",2012,"TAN, YONGMIN. Online Performance Anomaly Prediction and Prevention for Complex dummy Distributed Systems.(Under the direction of Dr. Xiaohui (Helen) Gu.) Real world distributed systems (eg, cloud computing infrastructures, enterprise data centers, massive data processing systems) have become increasingly complex as they grow in both scale and functionality. However, such complexity makes these systems vulnerable to performance anomalies caused by various faults such as resource contentions, performance …",https://dl.acm.org/doi/book/10.5555/2519690,True,,4.0,,,,['failure-detection'],,,True,['failure-management'],,,,,,10.5555/2519690,,,
964,Detecting Bottleneck in n-Tier IT Applications Through Analysis,"['Jung, Gueyoung', 'Swint, Galen', 'Parekh, Jason', 'Pu, Calton', 'Sahai, Akhil']",2006,"As the complexity of large-scale enterprise applications increases, providing performance verification through staging becomes an important part of reducing business risks associated with violating sophisticated service-level agreement (SLA). Currently, performance verification during the staging process is accomplished through either an expensive, cumbersome manual approach or ad hoc automation. This paper describes an automation approach as part of the Elba project supporting monitoring and performance …",https://link.springer.com/chapter/10.1007%2F11907466_13,True,,9.0,,,,['failure-detection'],,International Workshop on Distributed Systems: Operations and Management,True,['failure-management'],,,,,,,,,
965,"A scalable, non-parametric anomaly detection framework for Hadoop","['Li Yu', 'Zhiling Lan']",2013,"In this paper, we present a scalable and practical problem diagnosis framework for Hadoop environments. Our design features a decentralized approach based on hierarchical grouping and a novel non-parametric diagnostic mechanism. We evaluate our framework under various Hadoop workloads. The experimental results show that our design outperforms traditional methods significantly in the context of complex anomaly patterns and high anomaly probability.",https://dl.acm.org/doi/10.1145/2494621.2494643,True,,25.0,,,,['failure-detection'],,CAC '13: Proceedings of the 2013 ACM Cloud and Autonomic Computing Conference,True,['failure-management'],,,,,,10.1145/2494621.2494643,,,
966,Exploring time and frequency domains for accurate and automated anomaly detection in cloud computing systems,"['Qiang Guan', 'Song Fu', 'Nathan DeBardeleben', 'Sean Blanchard']",2013,"Cloud computing has become increasingly popular by obviating the need for users to own and maintain complex computing infrastructures. However, due to their inherent complexity and large scale, production cloud computing systems are prone to various runtime problems caused by hardware and software faults and environmental factors. Autonomic anomaly detection is crucial for understanding emergent, cloud-wide phenomena and self-managing cloud resources for system-level dependability assurance. To detect anomalous cloud …",https://ieeexplore.ieee.org/document/6820866,True,,16.0,,,,['failure-detection'],,2013 IEEE 19th Pacific Rim International Symposium on Dependable Computing,True,['failure-management'],,,,,,,,,
967,A hybrid anomaly detection framework in cloud computing using one-class and two-class support vector machines,"['Fu, Song', 'Liu, Jianguo', 'Pannu, Husanbir']",2012,"Modern production utility clouds contain thousands of computing and storage servers. Such a scale combined with ever-growing system complexity of their components and interactions, introduces a key challenge for anomaly detection and resource management for highly dependable cloud computing. Autonomic anomaly detection is a crucial technique for understanding emergent, cloud-wide phenomena and self-managing cloud resources for system level dependability assurance. We propose a new hybrid self-evolving anomaly detection framework using one-class and two-class support vector machines. Experimental results in an institute wide cloud computing system show that the detection accuracy of the algorithm improves as it evolves and it can achieve 92.1% detection sensitivity and 83.8% detection specificity, which makes it well suitable for building highly dependable clouds.",https://link.springer.com/chapter/10.1007/978-3-642-35527-1_60,True,,6.0,,,,['failure-detection'],,International Conference on Advanced Data Mining and Applications,True,['failure-management'],,,,,,,,,
968,Experience mining Google’s production console logs,"['Wei Xu', 'Ling Huang', 'Armando Fox', 'David Patterson', 'Michael Jordan']",2010,"We describe our early experience in applying our console log mining techniques [19, 20] to logs from production Google systems with thousands of nodes. This data set is five orders of magnitude in size and contains almost 20 times as many messages types as the Hadoop data set we used in [19]. It also has many properties that are unique to large scale production deployments (e.g., the system stays on for several months and multiple versions of the software can run concurrently). Our early experience shows that our techniques, including source code based log parsing, state and sequence based feature creation and problem detection, work well on this production data set. We also discuss our experience in using our log parser to assist the log sanitization.",https://dl.acm.org/doi/10.5555/1928991.1928999,True,,55.0,,,,['failure-detection'],,SLAML'10: Proceedings of the 2010 workshop on Managing systems via log analysis and machine learning techniques,True,['failure-management'],,,,,,10.5555/1928991.1928999,,,
969,Cherrypick: adaptively unearthing the best cloud configurations for big data analytics,"['Omid Alipourfard', 'Hongqiang Harry Liu', 'Jianshu Chen', 'Shivaram Venkataraman', 'Minlan Yu', 'Ming Zhang']",2017,"Picking the right cloud configuration for recurring big data analytics jobs running in clouds is hard, because there can be tens of possible VM instance types and even more cluster sizes to pick from. Choosing poorly can significantly degrade performance and increase the cost to run a job by 2-3x on average, and as much as 12x in the worst-case. However, it is challenging to automatically identify the best configuration for a broad spectrum of applications and cloud configurations with low search cost. CherryPick is a system that leverages Bayesian Optimization to build performance models for various applications, and the models are just accurate enough to distinguish the best or close-to-the-best configuration from the rest with only a few test runs. Our experiments on five analytic applications in AWS EC2 show that CherryPick has a 45-90% chance to find optimal configurations, otherwise near-optimal, saving up to 75% search cost compared to existing solutions.",https://dl.acm.org/doi/10.5555/3154630.3154669,True,,190.0,,,,['configuration'],,NSDI'17: Proceedings of the 14th USENIX Conference on Networked Systems Design and Implementation,True,['resource-provisioning'],,,,,,10.5555/3154630.3154669,,,
970,Time series anomaly detection for trustworthy services in cloud computing systems,"['Chengqiang Huang', 'Geyong Min', 'Yulei Wu', 'Yiming Ying', 'Ke Pei', 'Zuochang Xiang']",2017,"As a powerful architecture for large-scale computation, cloud computing has revolutionized the way that computing infrastructure is abstracted and utilized. Coupled with the challenges caused by Big Data, the rocketing development of cloud computing boosts the complexity of system management and maintenance, resulting in weakened trustworthiness of cloud services. To cope with this problem, a compelling method, i.e., Support Vector Data Description (SVDD), is investigated in this paper for detecting anomalous performance metrics of cloud services. Although competent in general anomaly detection, SVDD suffers from unsatisfactory false alarm rate and computational complexity in time series anomaly detection, which considerably hinders its practical applications. Therefore, this paper proposes a relaxed form of linear programming SVDD (RLPSVDD) and presents important insights into parameter selection for practical time series anomaly detection in order to monitor the operations of cloud services. Experiments on the Iris dataset and the Yahoo benchmark datasets validate the effectiveness of our approaches. Furthermore, the comparison of RLPSVDD and the methods obtained from Twitter, Numenta, Etsy and Yahoo, shows the overall preference for RLPSVDD in time series anomaly detection.",https://ieeexplore.ieee.org/document/7937930,True,,28.0,,,,['failure-detection'],,,True,['failure-management'],,,['anomaly-detection'],,,,,,
971,AVA: automated interpretation of dynamically detected anomalies,"['Anton Babenko', 'Leonardo Mariani', 'Fabrizio Pastore']",2009,"Dynamic analysis techniques have been extensively adopted to discover causes of observed failures. In particular, anomaly detection techniques can infer behavioral models from observed legal executions and compare failing executions with the inferred models to automatically identify the likely anomalous events that caused observed failures. Unfortunately the output of these techniques is limited to a set of independent suspicious anomalous events that does not capture the structure and the rationale of the differences between the correct and the failing executions. Thus, testers spend a relevant amount of time and effort to investigate executions and interpret these differences, reducing effectiveness of anomaly detection techniques. In this paper, we present Automata Violations Analyzer (AVA), a technique to automatically produce candidate interpretations of detected failures from anomalies identified by anomaly detection techniques. Interpretations capture the rationale of the differences between legal and failing executions with user understandable patterns that simplify identification of failure causes. The empirical validation with synthetic cases and third-party systems shows that AVA produces useful interpretations.",https://dl.acm.org/doi/10.1145/1572272.1572300,True,,36.0,['automaton'],"['logs', 'events']",['new-method'],['root-cause-analysis'],,ISSTA '09: Proceedings of the eighteenth international symposium on Software testing and analysis,True,['failure-management'],,,['root-cause-diagnosis'],,,10.1145/1572272.1572300,,,
972,DAPA: diagnosing application performance anomalies for virtualized infrastructures,"['Hui Kang', 'Xiaoyun Zhu', 'Jennifer L. Wong']",2012,"As cloud service providers leverage server virtualization to host applications in virtual machines (VMs), they must ensure proper allocation of resource capacities in order to satisfy the contracted service level agreements (SLAs) with the application owners. However, the ever-growing number of virtual and physical machines within such infrastructure creates greater challenges in quickly and effectively localizing the system bottlenecks that lead to SLA violations. This paper describes DAPA, a new performance diagnostic framework to help system administrators analyze application performance anomalies and identify potential causes of SLA violations. DAPA incorporates several customized statistical techniques to capture the quantitative relationship between the application performance and virtualized system metrics. We have built a prototype implementation of DAPA on a cluster of virtualized systems to diagnose a set of SLA violations for an enterprise application. Preliminary evaluation results show that DAPA is able to localize the most suspicious attributes of the virtual machines and physical hosts that are related to the SLA violations.",https://dl.acm.org/doi/10.5555/2228283.2228294,True,,33.0,,,,['root-cause-analysis'],['vm'],"Hot-ICE'12: Proceedings of the 2nd USENIX conference on Hot Topics in Management of Internet, Cloud, and Enterprise Networks and Services",True,['failure-management'],,,,,,10.5555/2228283.2228294,,,
973,X-ray: automating root-cause diagnosis of performance anomalies in production software,"['Mona Attariyan', 'Michael Chow', 'Jason Flinn']",2012,"Troubleshooting the performance of production software is challenging. Most existing tools, such as profiling, tracing, and logging systems, reveal what events occurred during performance anomalies. However, users of such toolsmust infer why these events occurred; e.g., that their execution was due to a root cause such as a specific input request or configuration setting. Such inference often requires source code and detailed application knowledge that is beyond system administrators and end users. This paper introduces performance summarization, a technique for automatically diagnosing the root causes of performance problems. Performance summarization instruments binaries as applications execute. It first attributes performance costs to each basic block. It then uses dynamic information flow tracking to estimate the likelihood that a block was executed due to each potential root cause. Finally, it summarizes the overall cost of each potential root cause by summing the per-block cost multiplied by the cause-specific likelihood over all basic blocks. Performance summarization can also be performed differentially to explain performance differences between two similar activities. X-ray is a tool that implements performance summarization. Our results show that X-ray accurately diagnoses 17 performance issues in Apache, lighttpd, Postfix, and PostgreSQL, while adding 2.3% average runtime overhead.",https://dl.acm.org/doi/10.5555/2387880.2387910,True,,211.0,['graph-mining'],"['host-metrics', 'traces']",['new-method'],['root-cause-analysis'],"['apache', 'postfix', 'postgreSQL', 'lightpd']",OSDI'12: Proceedings of the 10th USENIX conference on Operating Systems Design and Implementation,True,['failure-management'],True,,['root-cause-diagnosis'],True,81.0,10.5555/2387880.2387910,,,
974,Locating causes of program failures,"['Holger Cleve', 'Andreas Zeller']",2005,"Which is the defect that causes a software failure? By comparing the program states of a failing and a passing run, we can identify the state differences that cause the failure. However, these state differences can occur all over the program run. Therefore, we focus in space on those variables and values that are relevant for the failure, and in time on those moments where cause transitions occur---moments where new relevant variables begin being failure causes: ""Initially, variable argc was 3; therefore, at shell_sort(), variable [2] was 0, and therefore, the program failed."" In our evaluation, cause transitions locate the failure-inducing defect twice as well as the best methods known so far.",https://dl.acm.org/doi/10.1145/1062455.1062522,True,,707.0,['search'],['runs'],['new-method'],['root-cause-analysis'],['source-code'],ICSE '05: Proceedings of the 27th international conference on Software engineering,True,['failure-management'],True,,['fault-localization'],True,74.0,10.1145/1062455.1062522,,,
975,Tree-based methods for classifying software failures,"['P. Francis', 'D. Leon', 'M. Minch', 'A. Podgurski']",2004,"Recent research has addressed the problem of providing automated assistance to software developers in classifying reported instances of software failures so that failures with the same cause are grouped together. In this paper, two new tree-based techniques are presented for refining an initial classification of failures. One of these techniques is based on the use of dendrograms, which are rooted trees used to represent the results of hierarchical cluster analysis. The second technique employs a classification tree constructed to recognize failed executions. With both techniques, the tree representation is used to guide the refinement process. We also report the results of experimentally evaluating these techniques on several subject programs.",https://ieeexplore.ieee.org/document/1383139,True,,93.0,,,,['failure-detection'],,15th International Symposium on Software Reliability Engineering,True,['failure-management'],,,,,,,,,
976,Tejo: A Supervised Anomaly Detection Scheme for NewSQL Databases,"['Silvestre, Guthemberg', 'Sauvanaud, Carla', 'Ka{\\^a}niche, Mohamed', 'Kanoun, Karama']",2015,"The increasing availability of streams of data and the need of auto-tuning applications have made big data mainstream. NewSQL databases have become increasingly important to ensure fast data processing for the emerging stream processing platforms. While many architectural improvements have been made on NewSQL databases to handle fast data processing, anomalous events on the underlying, complex cloud environments may undermine their performance. In this paper, we present Tejo, a supervised anomaly …",https://www.researchgate.net/publication/282652510_Tejo_A_Supervised_Anomaly_Detection_Scheme_for_NewSQL_Databases,True,,2.0,,,,['failure-detection'],,International Workshop on Software Engineering for Resilient Systems,True,['failure-management'],,,,,,,,,
977,Failure Detection in Large-Scale Internet Services by Principal Subspace Mapping,"['Haifeng Chen', 'Guofei Jiang', 'Kenji Yoshihira']",2007,"Fast and accurate failure detection is becoming essential in managing large-scale Internet services. This paper proposes a novel detection approach based on the subspace mapping between the system inputs and internal measurements. By exploring these contextual dependencies, our detector can initiate repair actions accurately, increasing the availability of the system. Although a classical statistical method, the canonical correlation analysis (CCA), is presented in the paper to achieve subspace mapping, we also propose a more advanced technique, the principal canonical correlation analysis (PCCA), to improve the performance of the CCA-based detector. PCCA extracts a principal subspace from internal measurements that is not only highly correlated with the inputs but also a significant representative of the original measurements. Experimental results on a Java 2 platform, enterprise edition (J2EE)- based Web application demonstrate that such property of PCCA is especially beneficial to failure detection tasks.",https://ieeexplore.ieee.org/document/4302740,True,,22.0,,,,['failure-detection'],,,True,['failure-management'],,,,,,,,,
978,Mining for misconfigured machines in grid systems,"['Noam Palatin', 'Arie Leizarowitz', 'Assaf Schuster', 'Ran Wolff']",2006,"Grid systems are proving increasingly useful for managing the batch computing jobs of organizations. One well-known example is Intel, whose internally developed NetBatch system manages tens of thousands of machines. The size, heterogeneity, and complexity of grid systems make them very difficult, however, to configure. This often results in misconfigured machines, which may adversely affect the entire system.We investigate a distributed data mining approach for detection of misconfigured machines. Our Grid Monitoring System (GMS) non-intrusively collects data from all sources (log files, system services, etc.) available throughout the grid system. It converts raw data to semantically meaningful data and stores this data on the machine it was obtained from, limiting incurred overhead and allowing scalability. Afterwards, when analysis is requested, a distributed outliers detection algorithm is employed to identify misconfigured machines. The algorithm itself is implemented as a recursive workflow of grid jobs. It is especially suited to grid systems, in which the machines might be unavailable most of the time and often fail altogether.",https://dl.acm.org/doi/10.1145/1150402.1150488,True,,46.0,,,,['failure-detection'],,KDD '06: Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining,True,['failure-management'],,,,,,10.1145/1150402.1150488,,,
979,An Anomaly Detection Algorithm of Cloud Platform Based on Self-Organizing Maps,"['Liu, Jun', 'Chen, Shuyu', 'Zhou, Zhen', 'Wu, Tianshu']",2016,"Virtual machines (VM) on a Cloud platform can be influenced by a variety of factors which can lead to decreased performance and downtime, affecting the reliability of the Cloud platform. Traditional anomaly detection algorithms and strategies for Cloud platforms have some flaws in their accuracy of detection, detection speed, and adaptability. In this paper, a dynamic and adaptive anomaly detection algorithm based on Self-Organizing Maps (SOM) for virtual machines is proposed. A unified modeling method based on SOM to detect the machine performance within the detection region is presented, which avoids the cost of modeling a single virtual machine and enhances the detection speed and reliability of large-scale virtual machines in Cloud platform. The important parameters that affect the modeling speed are optimized in the SOM process to significantly improve the accuracy of the SOM modeling and therefore the anomaly detection accuracy of the virtual machine.",https://www.researchgate.net/publication/301251852_An_Anomaly_Detection_Algorithm_of_Cloud_Platform_Based_on_Self-Organizing_Maps,True,,16.0,['som'],,,['failure-detection'],['vm'],,True,['failure-management'],,,['anomaly-detection'],,,,,,
980,A Self-healing Framework for QoS-Aware Web Service Composition via Case-Based Reasoning,"['Li, Guoqiang', 'Liao, Lejian', 'Song, Dandan', 'Wang, Jingang', 'Sun, Fuzhen', 'Liang, Guangcheng']",2013,"The self-healing ability is very important for service-oriented systems in a highly dynamic environment. Case-based reasoning is adopted to cope with the component services failing to meet with the functional and nonfunctional requirements in this paper. We take previous failure instances as cases which are stored in a case base. When a new fault occurs, its symptoms are extracted and matched against the case base to look for the most similar case. A case representation and a similarity function are proposed. Meanwhile, a novel …",https://link.springer.com/chapter/10.1007/978-3-642-37401-2_64,True,,16.0,,,,['remediation'],,Asia-Pacific Web Conference,True,['failure-management'],,,,,,,,,
981,E2EProf: Automated End-to-End Performance Management for Enterprise Systems,"['Sandip Agarwala', 'Fernando Alegre', 'Karsten Schwan', 'Jegannathan Mehalingham']",2007,"Distributed systems are becoming increasingly complex, caused by the prevalent use of Web services, multi-tier architectures, and grid computing, where dynamic sets of components interact with each other across distributed and heterogeneous computing infrastructures. For these applications to be able to predictably and efficiently deliver services to end users, it is therefore, critical to understand and control their runtime behavior. In a datacenter environment, for instance, understanding the end-to-end dynamic behavior of certain IT subsystems, from the time requests are made to when responses are generated and finally, received, is a key prerequisite for improving application response, to provide required levels of performance, or to meet service level agreements (SLAs). The E2EProf toolkit enables the efficient and nonintrusive capture and analysis of end-to-end program behavior for complex enterprise applications. E2EProf permits an enterprise to recognize and analyze performance problems when they occur - online, to take corrective actions as soon as possible and wherever necessary along the paths currently taken by user requests - end-to-end, and to do so without the need to instrument applications - nonintrusively. Online analysis exploits a novel signal analysis algorithm, termed pathmap, which dynamically detects the causal paths taken by client requests through application and backend servers and annotates these paths with end-to-end latencies and with the contributions to these latencies from different path components. Thus, with pathmap, it is possible to dynamically identify the bottlenecks present in selected servers or services and to detect the abnormal or unusual performance behaviors indicative of potential problems or overloads. Pathmap and the E2EProf toolkit successfully detect causal request paths and associated performance bottlenecks in the RUBiS ebay-like multi-tier Web application and in one of the datacenter of our industry partner, Del...",https://ieeexplore.ieee.org/document/4273026,True,,80.0,,,,"['workload-prediction', 'resource-consolidation']",,37th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN'07),True,['resource-provisioning'],,,,,,,,,
982,CloudPD: Problem determination and diagnosis in shared dynamic clouds,"['Bikash Sharma', 'Praveen Jayachandran', 'Akshat Verma', 'Chita R. Das']",2013,"In this work, we address problem determination in virtualized clouds. We show that high dynamism, resource sharing, frequent reconfiguration, high propensity to faults and automated management introduce significant new challenges towards fault diagnosis in clouds. Towards this, we propose CloudPD, a fault management framework for clouds. CloudPD leverages (i) a canonical representation of the operating environment to quantify the impact of sharing; (ii) an online learning process to tackle dynamism; (iii) a correlation-based performance models for higher detection accuracy; and (iv) an integrated end-to-end feedback loop to synergize with a cloud management ecosystem. Using a prototype implementation with cloud representative batch and transactional workloads like Hadoop, Olio and RUBiS, it is shown that CloudPD detects and diagnoses faults with low false positives (<; 16%) and high accuracy of 88%, 83% and 83%, respectively. In an enterprise trace-based case study, CloudPD diagnosed anomalies within 30 seconds and with an accuracy of 77%, demonstrating its effectiveness in real-life operations.",https://ieeexplore.ieee.org/abstract/document/6575298,True,,80.0,"['pattern-matching', 'similarity-matching', 'markov-model', 'clustering']","['host-metrics', 'kpis']","['novel-use', 'new-method']","['root-cause-analysis', 'failure-detection']","['vm', 'cloud']",2013 43rd Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN),True,['failure-management'],True,,"['root-cause-diagnosis', 'anomaly-detection']",True,59.0,,,,
983,"Mining Behavior Graphs for ""Backtrace"" of Noncrashing Bugs","['Liu, Chao', 'Yan, Xifeng', 'Yu, Hwanjo', 'Han, Jiawei', 'Yu, Philip S']",2005,"Analyzing the executions of a buggy software program is essentially a data mining process. Although many interesting methods have been developed to trace crashing bugs (such as memory violation and core dumps), it is still difficult to analyze noncrashing bugs (such as logical errors). In this paper, we develop a novel method to classify the structured traces of program executions using software behavior graphs. By analyzing the correct and incorrect executions, we have made good progress at the isolation of program regions that may lead …",https://www.researchgate.net/publication/220907162_Mining_Behavior_Graphs_for_Backtrace_of_Noncrashing_Bugs,True,,98.0,['graph-mining'],,,['failure-detection'],,Proceedings of the 2005 SIAM International Conference on Data Mining,True,['failure-management'],,,['anomaly-detection'],,,,,,
984,Data Analysis of Minimally-Structured Heterogeneous Logs: An experimental study of log template extraction and anomaly detection based on Recurrent Neural Network and Naive Bayes,"['Liu, Chang']",2016,"Nowadays, the ideas of continuous integration and continuous delivery are under heavy usage in order to achieve rapid software development speed and quick product delivery to the customers with good quality. During the process of modern software development, the testing stage has always been with great significance so that the delivered software is meeting all the requirements and with high quality, maintainability, sustainability, scalability, etc. The key assignment of software testing is to find bugs from every test and solve them …",https://www.semanticscholar.org/paper/Data-Analysis-of-Minimally-Structured-Heterogeneous-Liu/658cd8b3c27ed1787e243b060295bb97a1e3ed5e,True,,2.0,"['rnn', 'naive-bayes']",,"['new-method', 'comparison', 'novel-use']",['failure-detection'],,,True,['failure-management'],,,['log-enhancement'],,,,,,
985,Log File Anomaly Detection,"['Yang, Tian', 'Agrawal, Vikas']",2016,"Analysis of log files pertaining to a failed run can be a tedious task, especially if the file runs into thousands of lines. Using the recent development in text analysis using deep neural networks, we present a method to reduce effort needed to analyze the log file by …",https://www.semanticscholar.org/paper/Log-File-Anomaly-Detection-Yang-Agrawal/f81d9cbb8ecb01be8c043af5fbab19c11c0245ba,True,,5.0,,,,['failure-detection'],,,True,['failure-management'],,,,,,,,,
986,Fault Prediction Modeling for Software Quality Estimation: Comparing Commonly Used Techniques,"['Khoshgoftaar, Taghi M', 'Seliya, Naeem']",2003,High-assurance and complex mission-critical software systems are heavily dependent on reliability of their underlying software applications. An early software fault prediction is a proven technique in achieving high software reliability. Prediction models based on software metrics can predict number of faults in software modules. Timely predictions of such models can be used to direct cost-effective quality enhancement efforts to modules that are likely to have a high number of faults. We evaluate the predictive performance of six commonly used …,https://link.springer.com/article/10.1023/A:1024424811345,True,,168.0,,,,['failure-prevention'],,,True,['failure-management'],,,['software-defect-prediction'],,,,,,
987,A neural network approach for early detection of program modules having high risk in the maintenance phase,"['Taghi M. Khoshgoftaar', 'David L. Lanning']",1995,"A neural network model is developed to classify program modules as either high or low risk based on multiple criterion variables. The inputs to the model include a selection of software complexity metrics collected from a telecommunications system. Two criterion variables are used for class determination: the number of changes to enhance the program modules, and the number of changes required to remove faults from the modules. The data were deliberately biased to magnify differences in metrics values between the discriminant …",https://dl.acm.org/doi/abs/10.1016/0164-1212%2894%2900130-F,True,,167.0,,,,['failure-prevention'],,Journal of Systems and Software,True,['failure-management'],,,['software-defect-prediction'],,,10.1016/0164-1212%2894%2900130-F,,,
988,Characterizing the behavior of a program using multiple-length N-grams,['Carla Marceau'],2001,"Some recent advances in intrusion detection are based on detecting anomalies in program behavior, as characterized by the sequence of kernel calls the program makes. Specifically, traces of kernel calls are collected during a training period. The substrings of fixed length N (for some N) of those traces are called N-grams. The set of N-grams occurring during normal execution has been found to discriminate effectively between normal behavior of a program and the behavior of the program under attack. The N-gram characterization, while effective …",https://dl.acm.org/doi/10.1145/366173.366197,True,,60.0,,,,['failure-detection'],,NSPW '00: Proceedings of the 2000 workshop on New security paradigms,True,['failure-management'],,,,,,10.1145/366173.366197,,,
989,Real-time failure prediction in online services,"['Mohammed Shatnawi', 'Mohamed Hefeeda']",2015,"Current data mining techniques used to create failure predictors for online services require massive amounts of data to build, train, and test the predictors. These operations are tedious, time consuming, and are not done in real-time. Also, the accuracy of the resulting predictor is highly compromised by changes that affect the environment and working conditions of the predictor. We propose a new approach to creating a dynamic failure predictor for online services in real-time and keeping its accuracy high during the services run-time changes. We use synthetic transactions during the run-time lifecycle to generate current data about the service. This data is used in its ephemeral state to build, train, test, and maintain an up-to-date failure predictor. We implemented the proposed approach in a large-scale online ad service that processes billions of requests each month in six data centers distributed in three continents. We show that the proposed predictor is able to maintain failure prediction accuracy as high as 86% during online service changes, whereas the accuracy of the state-of-the-art predictors may drop to less than 10%.",https://ieeexplore.ieee.org/document/7218516,True,,16.0,['rule-mining'],['events'],['new-method'],['failure-prediction'],,2015 IEEE Conference on Computer Communications (INFOCOM),True,['failure-management'],,,['system-failure-prediction'],,,,,,
990,DeepDive: transparently identifying and managing performance interference in virtualized environments,"['Dejan Novaković', 'Nedeljko Vasić', 'Stanko Novaković', 'Dejan Kostić', 'Ricardo Bianchini']",2013,"We describe the design and implementation of Deep-Dive, a system for transparently identifying and managing performance interference between virtual machines (VMs) co-located on the same physical machine in Infrastructure-as-a-Service cloud environments. DeepDive successfully addresses several important challenges, including the lack of performance information from applications, and the large overhead of detailed interference analysis. We first show that it is possible to use easily-obtainable, low-level metrics to clearly discern when interference is occurring and what resource is causing it. Next, using realistic workloads, we show that DeepDive quickly learns about interference across co-located VMs. Finally, we show DeepDive's ability to deal efficiently with interference when it is detected, by using a low-overhead approach to identifying a VM placement that alleviates interference.",https://dl.acm.org/doi/10.5555/2535461.2535489,True,,52.0,,,,['scheduling'],,USENIX ATC'13: Proceedings of the 2013 USENIX conference on Annual Technical Conference,True,['resource-provisioning'],,,,,,10.5555/2535461.2535489,,,
991,Anomaly Detection and Diagnosis for Cloud services: Practical experiments and lessons learned,"['Sauvanaud, Carla', 'Ka{\\^a}niche, Mohamed', 'Kanoun, Karama', 'Lazri, Kahina', 'Silvestre, Guthemberg Da Silva']",2018,"The dependability of cloud computing services is a major concern of cloud providers. In particular, anomaly detection techniques are crucial to detect anomalous service behaviors that may lead to the violation of service level agreements (SLAs) drawn with users. This paper describes an anomaly detection system (ADS) designed to detect errors related to the erroneous behavior of the service, and SLA violations in cloud services. One major objective is to help providers to diagnose the anomalous virtual machines (VMs) on which a service is …",https://www.researchgate.net/publication/322892457_Anomaly_Detection_and_Diagnosis_for_Cloud_services_Practical_experiments_and_lessons_learned,True,,13.0,,,,['failure-detection'],,,True,['failure-management'],,,,,,,,,
992,Multivariate Online Anomaly Detection Using Kernel Recursive Least Squares,"['Tarem Ahmed', 'Mark Coates', 'Anukool Lakhina']",2007,"High-speed backbones are regularly affected by various kinds of network anomalies, ranging from malicious attacks to harmless large data transfers. Different types of anomalies affect the network in different ways, and it is difficult to know a priori how a potential anomaly will exhibit itself in traffic statistics. In this paper we describe an online, sequential, anomaly detection algorithm, that is suitable for use with multivariate data. The proposed algorithm is based on the kernel version of the recursive least squares algorithm. It assumes no model for network traffic or anomalies, and constructs and adapts a dictionary of features that approximately spans the subspace of normal behaviour. The algorithm raises an alarm immediately upon encountering a deviation from the norm. Through comparison with existing block-based offline methods based upon Principal Component Analysis, we demonstrate that our online algorithm is equally effective but has much faster time-to-detection and lower computational complexity. We also explore minimum volume set approaches in identifying the region of normality.",https://ieeexplore.ieee.org/document/4215661,True,,81.0,,,,['failure-detection'],,IEEE INFOCOM 2007-26th IEEE International Conference on Computer Communications,True,['failure-management'],,,,,,,,,
993,Real-time detection of performance anomalies for cloud services,"['Olumuyiwa Ibidunmoye', 'Thijs Metsch', 'Erik Elmroth']",2016,"Service performance degradation and downtimes are a common on the Internet today. Many on-line services (e.g. Amazon.com, Spotify, and Netflix, etc.) report huge loss in revenue and traffic per episode. This is perhaps due to the correlation between performance and end-users's satisfaction.",https://ieeexplore.ieee.org/document/7590412,True,,10.0,['autoregression'],,['novel-use'],['failure-detection'],,2016 IEEE/ACM 24th International Symposium on Quality of Service (IWQoS),True,['failure-management'],,,['anomaly-detection'],,,,,,['online']
994,An Anomaly Detection Framework for Detecting Anomalous Virtual Machines under Cloud Computing Environment,"['Wang, GuiPing', 'Wang, JiaWei']",2016,"A variety of faults may cause performance degradation or even downtime of virtual machines (VMs) under Cloud environment, thus lowering the dependability of Cloud platform. Detecting anomalous VMs before real failures occur is an important means to improve the dependability of Cloud platform. Since the performance or state of VMs may be affected by the environmental factors, this article proposes an environment-aware anomaly detection framework (termed EaAD) for VMs under Cloud environment. EaAD partitions all the VMs in …",https://www.researchgate.net/publication/296938547_An_Anomaly_Detection_Framework_for_Detecting_Anomalous_Virtual_Machines_under_Cloud_Computing_Environment,True,,8.0,,,,['failure-detection'],,,True,['failure-management'],,,,,,,,,
995,An unsupervised anomaly detection approach using energy-based spatiotemporal graphical modeling,"['Liu, Chao', 'Ghosal, Sambuddha', 'Jiang, Zhanhong', 'Sarkar, Soumik']",2017,"This paper presents a new data-driven framework for unsupervised system-wide anomaly detection for modern distributed complex systems within which there exists a strong connectivity among sub-systems, operating in diverse modes and encountering a large variety of anomalies. The framework is based on a spatiotemporal feature extraction scheme built on the concept of symbolic dynamics for discovering and representing causal interactions among subsystems. The extracted features from the spatiotemporal pattern …",https://www.researchgate.net/publication/320385140_An_unsupervised_anomaly_detection_approach_using_energy-based_spatiotemporal_graphical_modeling,True,,22.0,,,,['failure-detection'],,,True,['failure-management'],,,,,,,,,
996,Investigation and practical assessment of alarm correlation methods for the use in GSM access networks,['H. Wietgrefe'],2002,"This paper compares and assesses several alarm correlation methods for their suitability and performance in global systems for mobile communications (GSM). The assessment criteria used reflect the special circumstances found in these networks. A high importance is given to the aspects related to a practical and feasible network management. Of the neural networks investigated, the cascade correlation learning algorithm performs best. This approach is compared with correlation techniques proposed in the literature: rule-based diagnosis, model based diagnosis and alarm correlation using codebooks. It is shown that for alarm correlation in a GSM access network the proposed cascade correlation approach is superior to the other correlation techniques.",https://ieeexplore.ieee.org/document/1015597,True,,45.0,,,,['root-cause-analysis'],,NOMS 2002. IEEE/IFIP Network Operations and Management Symposium.'Management Solutions for the New Communications World'(Cat. No. 02CH37327),True,['failure-management'],,,,,,,,,
997,BP NEURAL NETWORK-BASED EFFECTIVE FAULT LOCALIZATION,"['Wong, W Eric', 'Qi, Yu']",2009,"In program debugging, fault localization identifies the exact locations of program faults. Finding these faults using an ad-hoc approach or based only on programmers' intuitive guesswork can be very time consuming. A better way is to use a well-justified method, supported by case studies for its effectiveness, to automatically identify and prioritize suspicious code for an examination of possible fault locations. To do so, we propose the use of a back-propagation (BP) neural network, a machine learning model which has been successfully applied to software risk analysis, cost prediction, and reliability estimation, to help programmers effectively locate program faults. A BP neural network is suitable for learning the input-output relationship from a set of data, such as the inputs and the corresponding outputs of a program. We first train a BP neural network with the coverage data (statement coverage in our case) and the execution result (success or failure) collected from executing a program, and then we use the trained network to compute the suspiciousness of each executable statement, in terms of its likelihood of containing faults. Suspicious code is ranked in descending order based on its suspiciousness. Programmers will examine such code from the top of the rank to identify faults. Four case studies on different programs (the Siemens suite, the Unix suite, grep and gzip) are conducted. Our results suggest that a BP neural network-based fault localization method is effective in locating program faults.",https://www.worldscientific.com/doi/10.1142/S021819400900426X,True,,55.0,,,,['root-cause-analysis'],,,True,['failure-management'],,,,,,,,,
998,Using Machine Learning to Support Debugging with Tarantula,"['Lionel C. Briand', 'Yvan Labiche', 'Xuetao Liu']",2007,"Using a specific machine learning technique, this paper proposes a way to identify suspicious statements during debugging. The technique is based on principles similar to Tarantula but addresses its main flaw: its difficulty to deal with the presence of multiple faults as it assumes that failing test cases execute the same fault(s). The improvement we present in this paper results from the use of C4.5 decision trees to identify various failure conditions based on information regarding the test cases' inputs and outputs. Failing test cases executing under similar conditions are then assumed to fail due to the same fault(s). Statements are then considered suspicious if they are covered by a large proportion of failing test cases that execute under similar conditions. We report on a case study that demonstrates improvement over the original Tarantula technique in terms of statement ranking. Another contribution of this paper is to show that failure conditions as modeled by a C4.5 decision tree accurately predict failures and can therefore be used as well to help debugging.",https://ieeexplore.ieee.org/abstract/document/4402205,True,,90.0,,,,['root-cause-analysis'],,The 18th IEEE International Symposium on Software Reliability (ISSRE'07),True,['failure-management'],,,['rca-others'],,,,,,
999,Software Fault Localization Using N-gram Analysis,"['Nessa, Syeda', 'Abedin, Muhammad', 'Wong, W Eric', 'Khan, Latifur', 'Qi, Yu']",2008,"A major portion of software development effort is spent in testing and debugging. Execution sequence collected in the testing phase can be a rich source of information for locating the fault in the program, but the exact execution sequence of a program, i.e., the actual order of execution of the statements in the program, is seldom used due to the huge volume. In this study, we apply data mining techniques on this data to reduce the debugging time by narrowing down the possible location of the fault. Our method applies N-gram analysis to rank the executable statements of a software by level of suspicion. We conducted three case studies to demonstrate the effectiveness of our proposed method. We also present comparison with other approaches, and illustrate the potential of our method.",https://link.springer.com/chapter/10.1007/978-3-540-88582-5_51,True,,15.0,,,,['root-cause-analysis'],,"International Conference on Wireless Algorithms, Systems, and Applications",False,['failure-management'],,,,,,,,,
1000,Formal concept analysis enhances fault localization in software,"['Peggy Cellier', 'Mireille Ducassé', 'Sébastien Ferré', 'Olivier Ridoux']",2008,"Recent work in fault localization crosschecks traces of correct and failing execution traces. The implicit underlying technique is to search for association rules which indicate that executing a particular source line will cause the whole execution to fail. This technique, however, has limitations. In this article, we first propose to consider more expressive association rules where several lines imply failure. We then propose to use Formal Concept Analysis (FCA) to analyze the resulting numerous rules in order to improve the readability of the information contained in the rules. The main contribution of this article is to show that applying two data mining techniques, association rules and FCA, produces better results than existing fault localization techniques.",https://dl.acm.org/doi/10.5555/1787746.1787766,True,,49.0,,,,['root-cause-analysis'],,ICFCA'08: Proceedings of the 6th international conference on Formal concept analysis,True,['failure-management'],,,,,,10.5555/1787746.1787766,,,
1001,Multiple Fault Localization with Data Mining,"['Cellier, Peggy', ""Ducass{\\'e}, Mireille"", ""Ferr{\\'e}, S{\\'e}bastien"", 'Ridoux, Olivier']",2011,"We propose an interactive fault localization method based on two data mining techniques, formal concept analysis and association rules. A lattice formalizes the partial ordering and the dependencies between the sets of program elements (eg, lines) that are most likely to lead to program execution failures. The paper provides an algorithm to traverse that lattice starting from the most suspect places. The main contribution is that the algorithm is able to deal with any number of faults within a single execution of a test suite. In addition, a stopping …",https://www.hal.inserm.fr/IRISA/hal-01119562,True,,26.0,"['rule-mining', 'formal-concept-analysis']",['test-cases'],['novel-use'],['root-cause-analysis'],['source-code'],SEKE,True,['failure-management'],,,['fault-localization'],,,,,,
1002,"Data mining and cross-checking of execution traces: a re-interpretation of Jones, Harrold and Stasko test information","['Tristan Denmat', 'Mireille Ducassé', 'Olivier Ridoux']",2005,"The current trend in debugging and testing is to cross-check information collected during several executions. Jones et al., for example, propose to use the instruction coverage of passing and failing runs in order to visualize suspicious statements. This seems promising but lacks a formal justification. In this paper, we show that the method of Jones et al. can be re-interpreted as a data mining procedure. More particularly, they define an indicator which characterizes association rules between data. With this formal framework we are able to explain intrinsic limitations of the above indicator.",https://dl.acm.org/doi/10.1145/1101908.1101979,True,,57.0,,,,['root-cause-analysis'],,ASE '05: Proceedings of the 20th IEEE/ACM international Conference on Automated software engineering,True,['failure-management'],,,,,,10.1145/1101908.1101979,,,
1003,Model-Based Debugging or How to Diagnose Programs Automatically,"['Franz Wotawa', 'Markus Stumptner', 'Wolfgang Mayer']",2002,"We describe the extension of the well-known model-based diagnosis approach to the location of errors in imperative programs (exhibited on a subset of the Java language). The source program is automatically converted to a logical representation (called model). Given this model and a particular test case or set of test cases, a program-independent search algorithm determines a the minimal sets of statements whose incorrectness can explain incorrect outcomes when the program is executed on the test cases, and which can then be indicated to the developer by the system. We analyze example cases and discuss empirical results from a Java debugger implementation incorporating our approach. The use of AI techniques is more flexible than traditional debugging techniques such as algorithmic debugging and program slicing.",https://dl.acm.org/doi/10.5555/646864.708248,True,,70.0,,,,['root-cause-analysis'],,IEA/AIE '02: Proceedings of the 15th international conference on Industrial and engineering applications of artificial intelligence and expert systems: developments in applied artificial intelligence,True,['failure-management'],,,,,,10.5555/646864.708248,,,
1004,Software fault localization using semi-supervised learning,"['Wei, Zheng', 'Xiaoxue, Wu', 'Xin, Tan', 'Yaopeng, Peng', 'Shuai, Yang']",2015,"In order to improve the efficiency of software fault localization, supervised learning methods are widely used in automatic software fault localization. But these methods mostly ignore a very important fact: in order to train a good performance of the classifier through supervised learning method, there must be a large number of labeled samples. While in the actual project, to obtain a large number of labeled samples is quite difficult; even if it can be done, the cost is very high. In order to solve this problem, we propose a semi supervised learning …",https://www.researchgate.net/publication/282981823_Software_fault_localization_using_semi-supervised_learning,True,,1.0,,,,['root-cause-analysis'],,,True,['failure-management'],,,,,,,,,
1005,Statistical Debugging: A Hypothesis Testing-Based Approach,"['Chao Liu', 'Long Fei', 'Xifeng Yan', 'Jiawei Han', 'S.P. Midkiff']",2006,"Manual debugging is tedious, as well as costly. The high cost has motivated the development of fault localization techniques, which help developers search for fault locations. In this paper, we propose a new statistical method, called SOBER, which automatically localizes software faults without any prior knowledge of the program semantics. Unlike existing statistical approaches that select predicates correlated with program failures, SOBER models the predicate evaluation in both correct and incorrect executions and regards a predicate as fault-relevant if its evaluation pattern in incorrect executions significantly diverges from that in correct ones. Featuring a rationale similar to that of hypothesis testing, SOBER quantifies the fault relevance of each predicate in a principled way. We systematically evaluate SOBER under the same setting as previous studies. The result clearly demonstrates the effectiveness: SOBER could help developers locate 68 out of the 130 faults in the Siemens suite by examining no more than 10 percent of the code, whereas the cause transition approach proposed by Holger et al. [2005] and the statistical approach by Liblit et al. [2005] locate 34 and 52 faults, respectively. Moreover, the effectiveness of SOBER is also evaluated in an ""imperfect world"", where the test suite is either inadequate or only partially labeled. The experiments indicate that SOBER could achieve competitive quality under these harsh circumstances. Two case studies with grep 2.2 and bc 1.06 are reported, which shed light on the applicability of SOBER on reasonably large programs",https://ieeexplore.ieee.org/document/1717474,True,,326.0,"['correlation', 'rule-mining']",['runs'],['new-method'],['root-cause-analysis'],,,True,['failure-management'],False,,['fault-localization'],,,,,,
1006,SOBER: statistical model-based bug localization,"['Chao Liu', 'Xifeng Yan', 'Long Fei', 'Jiawei Han', 'Samuel P. Midkiff']",2005,"Automated localization of software bugs is one of the essential issues in debugging aids. Previous studies indicated that the evaluation history of program predicates may disclose important clues about underlying bugs. In this paper, we propose a new statistical model-based approach, called SOBER, which localizes software bugs without any prior knowledge of program semantics. Unlike existing statistical debugging approaches that select predicates correlated with program failures, SOBER models evaluation patterns of predicates in both correct and incorrect runs respectively and regards a predicate as bug-relevant if its evaluation pattern in incorrect runs differs significantly from that in correct ones. SOBER features a principled quantification of the pattern difference that measures the bug-relevance of program predicates.We systematically evaluated our approach under the same setting as previous studies. The result demonstrated the power of our approach in bug localization: SOBER can help programmers locate 68 out of 130 bugs in the Siemens suite when programmers are expected to examine no more than 10% of the code, whereas the best previously reported is 52 out of 130. Moreover, with the assistance of SOBER, we found two bugs in bc 1.06 (an arbitrary precision calculator on UNIX/Linux), one of which has never been reported before.",https://dl.acm.org/doi/10.1145/1095430.1081753,True,,457.0,[],['runs'],['new-method'],['root-cause-analysis'],,ACM SIGSOFT Software Engineering Notes,True,['failure-management'],True,,['fault-localization'],True,75.0,10.1145/1095430.1081753,,,
1007,Abstract interpretation of programs for model-based debugging,"['Wolfgang Mayer', 'Markus Stumptner']",2007,"Developing model-based automatic debugging strategies has been an active research area for several years. We present a model-based debugging approach that is based on Abstract Interpretation, a technique borrowed from program analysis. The Abstract Interpretation mechanism is integrated with a classical model-based reasoning engine. We test the approach on sample programs and provide the first experimental comparison with earlier models used for debugging. The results show that the Abstract Interpretation based model provides more precise explanations than previous models or standard non-model based approaches.",https://dl.acm.org/doi/10.5555/1625275.1625350,True,,21.0,,,,['root-cause-analysis'],,IJCAI'07: Proceedings of the 20th international joint conference on Artifical intelligence,True,['failure-management'],,,,,,10.5555/1625275.1625350,,,
1008,A Crosstab-based Statistical Method for Effective Fault Localization,"['Eric Wong', 'Tingting Wei', 'Yu Qi', 'Lei Zhao']",2008,"Fault localization is the most expensive activity in program debugging. Traditional ad-hoc methods can be time-consuming and ineffective because they rely on programmers' intuitive guesswork, which may neither be accurate nor reliable. A better solution is to utilize a systematic and statistically well-defined method to automatically identify suspicious code that should be examined for possible fault locations. We present a crosstab-based statistical method using the coverage information of each executable statement and the execution result (success or failure) with respect to each test case. A crosstab is constructed for each executable statement and a statistic is computed to determine the suspiciousness of the corresponding statement. Statements with a higher suspiciousness are more likely to contain bugs and should be examined before those with a lower suspiciousness. Three case studies using the Siemens suite, the Space program, and the Unix suite, respectively, are conducted. Our results suggest that the crosstab-based method is effective in fault localization and performs better (in terms of a smaller percentage of executable statements that have to be examined until the first statement containing the fault is reached) than other methods such as Tarantula. The difference in efficiency (computational time) between these two methods is very small.",https://ieeexplore.ieee.org/document/4539531,True,,166.0,['correlation'],['test-cases'],['new-method'],['root-cause-analysis'],['software'],"2008 1st international conference on software testing, verification, and validation",True,['failure-management'],,,['fault-localization'],,,,,,
1009,Isolating cause-effect chains from computer programs,['Andreas Zeller'],2002,"Consider the execution of a failing program as a sequence of program states. Each state induces the following state, up to the failure. Which variables and values of a program state are relevant for the failure? We show how the Delta Debugging algorithm isolates the relevant variables and values by systematically narrowing the state difference between a passing run and a failing run---by assessing the outcome of altered executions to determine wether a change in the program state makes a difference in the test outcome. Applying Delta Debugging to multiple states of the program automatically reveals the cause-effect chain of the failure---that is, the variables and values that caused the failure.In a case study, our prototype implementation successfully isolated the cause-effect chain for a failure of the GNU C compiler: ""Initially, the C program to be compiled contained an addition of 1.0; this caused an addition operator in the intermediate RTL representation; this caused a cycle in the RTL tree---and this caused the compiler to crash.""",https://dl.acm.org/doi/10.1145/587051.587053,True,,701.0,['graph-mining'],"['runs', 'source-code']",,['root-cause-analysis'],['software'],SIGSOFT '02/FSE-10: Proceedings of the 10th ACM SIGSOFT symposium on Foundations of software engineering,True,['failure-management'],True,,['fault-localization'],True,73.0,10.1145/587051.587053,,,
1010,Experimental evaluation of using dynamic slices for fault location,"['Xiangyu Zhang', 'Haifeng He', 'Neelam Gupta', 'Rajiv Gupta']",2005,"Dynamic slicing algorithms have been considered to aid in debugging for many years. However, as far as we know, no detailed studies on evaluating the benefits of using dynamic slicing for detecting faulty statements in programs have been carried out. We have developed a dynamic slicing framework that uses dynamic instrumentation to efficiently collect dynamic slices and reduced ordered Binary Decision Diagrams (roBDDs) to compactly store them. We have used the above framework to implement three variants of dynamic slicing algorithms including: data slicing, full slicing, and relevant slicing algorithms. We have carried out detailed experiments to evaluate these algorithms. Our results show that full slices and relevant slices can considerably reduce the subset of program statements that need to be examined to locate faulty statements. We expect that the observations presented here will enable development of new slicing based algorithms for automated debugging.",https://dl.acm.org/doi/abs/10.1145/1085130.1085135,True,,83.0,,,,['root-cause-analysis'],,AADEBUG'05: Proceedings of the sixth international symposium on Automated analysis-driven debugging,True,['failure-management'],,,,,,10.1145/1085130.1085135,,,
1011,Structure and chance: melding logic and probability for software debugging,"['Lisa Burnell', 'Eric Horvitz']",1995,"Software errors abound in the world of computing. Sophisticated computer programs rank high on the list of the most complex systems ever created by humankind. The complexity of a program or a set of interacting programs makes it extremely difficult to perform offline verification of run-time behavior. Thus, the creation and maintenance of program code is often linked to a process of incremental refinement and ongoing detection and correction of errors. To be sure, the detection and repair of program errors is an inescapable part of the process of software development. However, run-time software errors may be discovered in fielded applications days, months, or even years after the software was last modified—especially in applications composed of a plethora of separate programs created and updated by different people at different times. In such complex applications, software errors are revealed through the run-time interaction of hundreds of distinct processes competing for limited memory and CPU resources. Software developers and support engineers responsible for correcting software problems face difficult challenges in tracking down the source of run-time errors in complex applications. The information made available to engineers about the nature of a failure often leaves open a wide range of possibilities that must be sifted through carefully in searching for an underlying error.",https://dl.acm.org/doi/10.1145/203330.203338,True,,75.0,,,,['root-cause-analysis'],,Communications of the ACM,True,['failure-management'],,,,,,10.1145/203330.203338,,,
1012,Mining Temporal Specifications for Error Detection,"['Weimer, Westley', 'Necula, George C']",2005,"Specifications are necessary in order to find software bugs using program verification tools. This paper presents a novel automatic specification mining algorithm that uses information about error handling to learn temporal safety rules. Our algorithm is based on the observation that programs often make mistakes along exceptional control-flow paths, even when they behave correctly on normal execution paths. We show that this focus improves the effectiveness of the miner for discovering specifications beneficial for bug finding. We …",https://link.springer.com/chapter/10.1007/978-3-540-31980-1_30,True,,264.0,,,,['failure-prevention'],['source-code'],International Conference on Tools and Algorithms for the Construction and Analysis of Systems,True,['failure-management'],,,['software-defect-prediction'],,,,,,
1013,"A Controller Architecture for Anomaly Detection, Root Cause Analysis and Self-Adaptation for Cluster Architectures","['Samir, Areeg', 'Pahl, Claus']",2019,"Service-based cloud computing allows applications to be deployed and managed through third-party provided services , making typically virtualised resources available. However, often there is no direct access to platform-level execution parameters of a provided service, and only some quality properties can be directly observed while others remain hidden from the service consumer. We introduce a controller architecture for autonomous, self-adaptive anomaly remediation in this semi-hidden setting. The controller determines the possible causes of consumer-observed anomalies in an underlying provider-controlled infrastructure. We use Hidden Markov Models to map observed performance anomalies into hidden resources, and to identify the root causes of the observed anomalies. We apply the model to a clustered computing resource environment that is based on three layers of aggregated resources.",https://www.researchgate.net/publication/333235851_A_Controller_Architecture_for_Anomaly_Detection_Root_Cause_Analysis_and_Self-Adaptation_for_Cluster_Architectures,True,,8.0,['markov-model'],"['sla', 'host-metrics']","['new-method', 'comparison']","['root-cause-analysis', 'failure-detection', 'remediation']",['cluster'],Proceedings of the Eleventh International Conference on Adaptive and Self-Adaptive Systems and Applications,True,['failure-management'],True,,"['recovery', 'root-cause-diagnosis', 'anomaly-detection']",True,92.0,,,,
1014,AI Planning and Combinatorial Optimization for Web Service Composition in Cloud Computing,"['Zou, Guobing', 'Chen, Yixin', 'Yang, Y', 'Huang, Ruoyun', 'Xu, You']",2010,"In recent years, there has been an increasing interest in web service composition due to its importance in practical applications. At the same time, cloud computing is gradually evolving as a widely used computing platform where many different web services are published and available in cloud data centers. The issue is that traditional service composition methods mainly focus on how to find service composition sequence in a single cloud, but not from a multi-cloud service base. It is challenging to efficiently find a composition solution in a multiple cloud base because it involves not only service composition but also combinatorial optimization. In this paper, we first propose a framework of service composition in multi-cloud base environments. Next, three different cloud combination methods are presented to select a cloud combination subject to not only finding feasible composition sequence, but also containing minimum clouds.",https://www.researchgate.net/publication/269087035_AI_Planning_and_Combinatorial_Optimization_for_Web_Service_Composition_in_Cloud_Computing,True,,69.0,['optimization'],,['new-method'],['service-composition'],,Proc international conference on cloud computing and virtualization,True,['resource-provisioning'],,,,,,10.5176/978-981-08-5837-7_166,,,
1015,Application of artificial intelligence techniques for root cause analysis of customer support calls,"['Parthasarathy, Sailashri']",2017,"Dell Technologies seeks to use the advancements in the field of artificial intelligence to improve its products and services. This thesis aims to implement artificial intelligence techniques in the context of Dell's Client Solutions Division, specifically to analyze the root cause of customer calls so actions can be taken to remedy them. This improves the customer experience while reducing the volume of calls, and hence costs, to Dell. This thesis evaluated the external vendor landscape for text analytics, developed an internal proof-of …",https://www.semanticscholar.org/paper/Application-of-artificial-intelligence-techniques-Parthasarathy/5a3568e5c0b6db9ae587f7cce81bd48c69374d56,True,,0.0,"['language-modeling', 'logistic-regression']","['tickets', 'logs']",,['root-cause-analysis'],,,True,['failure-management'],,,['root-cause-diagnosis'],,,,,,
1016,A Coding Approach to Event Correlation,"['Kliger, Shmuel', 'Yemini, Shaula', 'Yemini, Yechiam', 'Ohsie, David', 'Stolfo, Salvatore']",1995,This paper describes a novel approach to event correlation in networks based on coding techniques. Observable symptom events are viewed as a code that identifies the problems that caused them; correlation is performed by decoding the set of observed symptoms. The …,https://link.springer.com/chapter/10.1007%2F978-0-387-34890-2_24,True,,290.0,['codebook'],['events'],['new-method'],['root-cause-analysis'],,International Symposium on Integrated Network Management,True,['failure-management'],False,,['root-cause-diagnosis'],,,,,,
1017,Combinatorial designs in multiple faults localization for battlefield networks,"['M.A. Fecko', 'M. Steinder']",2001,We present an application of combinatorial designs and variance analysis to correlating events in the midst of multiple network faults. The network fault model is based on the probabilistic dependency graph that accounts for the uncertainty about the state of network elements. Orthogonal arrays help reduce the exponential number of failure configurations to a small subset on which further analysis is performed. The preliminary results show that statistical analysis can pinpoint the probable causes of the observed symptoms with high accuracy and significant level of confidence. An example demonstrates how multiple soft link failures are localized in MIL-STD 188-220's datalink layer to explain the end-to-end connectivity problems in the network layer This technique can be utilized for the networks operating in an unreliable environment such as wireless and/or military networks.,https://ieeexplore.ieee.org/document/985975,True,,26.0,,,,['root-cause-analysis'],,2001 MILCOM Proceedings Communications for Network-Centric Operations: Creating the Information Force (Cat. No. 01CH37277),True,['failure-management'],,,,,,,,,
1018,Effective Fault Localization using Code Coverage,"['W. Eric Wong', 'Yu Qi', 'Lei Zhao', 'Kai-Yuan Cai']",2007,"Localizing a bug in a program can be a complex and time- consuming process. In this paper we propose a code coverage-based fault localization method to prioritize suspicious code in terms of its likelihood of containing program bugs. Code with a higher risk should be examined before that with a lower risk, as the former is more suspicious (i.e., more likely to contain program bugs) than the latter. We also answer a very important question: how can each additional test case that executes the program successfully help locate program bugs? We propose that with respect to a piece of code, the aid introduced by the first successful test that executes it in computing its likelihood of containing a bug is larger than or equal to that of the second successful test that executes it, which is larger than or equal to that of the third successful test that executes it, etc. A tool, chiDebug, was implemented to automate the computation of the risk of the code and the subsequent prioritization of suspicious code for locating program bugs. A case study using the Siemens suite was also conducted. Data collected from our study support the proposal described above. They also indicate that our method (in particular Heuristics III (c), (d), and (e)) can effectively reduce the search domain for locating program bugs.",https://ieeexplore.ieee.org/document/4291037,True,,174.0,['heuristics'],"['test-cases', 'runs']",['new-method'],['root-cause-analysis'],['software'],31st Annual International Computer Software and Applications Conference (COMPSAC 2007),True,['failure-management'],True,,['fault-localization'],,,,,,
1019,Software Fault Localization via Mining Execution Graphs,"['Parsa, Saeed', 'Naree, Somaye Arabi', 'Koopaei, Neda Ebrahimi']",2011,"Software fault localization has attracted a lot of attention recently. Most existing methods focus on finding a single suspicious statement of code which is likelihood of containing bugs. Despite the accuracy of such methods, developers have trouble understanding the context of the bug, given each bug location in isolation. There is a high possibility of locating bug contexts through finding discriminative execution sub-paths between failing and passing executions. Representing each execution of a program as a graph, discriminative …",https://link.springer.com/chapter/10.1007/978-3-642-21887-3_46,True,,10.0,,,,['root-cause-analysis'],,International Conference on Computational Science and Its Applications,True,['failure-management'],,,,,,,,,
1020,Evolving Human Competitive Spectra-Based Fault Localisation Techniques,"['Yoo, Shin']",2012,Spectra-Based Fault Localisation (SBFL) aims to assist debugging by applying risk evaluation formulæ (sometimes called suspiciousness metrics) to program spectra and ranking statements according to the predicted risk. Designing a risk evaluation formula is often an intuitive process done by human software engineer. This paper presents a Genetic Programming (GP) approach for evolving risk assessment formulæ. The empirical evaluation using 92 faults from four Unix utilities produces promising results. Equations …,https://link.springer.com/chapter/10.1007/978-3-642-33119-0_18,True,,35.0,['genetic-programming'],,,['root-cause-analysis'],,International Symposium on Search Based Software Engineering,True,['failure-management'],,,,,,,,,
1021,Generating test data for both path coverage and fault detection using genetic algorithms,"['Gong, Dunwei', 'Zhang, Yan']",2013,"The aim of software testing is to find faults in a program under test, so generating test data that can expose the faults of a program is very important. To date, current studies on generating test data for path coverage do not perform well in detecting low probability faults on the covered path. The automatic generation of test data for both path coverage and fault detection using genetic algorithms is the focus of this study. To this end, the problem is first formulated as a bi-objective optimization problem with one constraint whose objectives are …",https://link.springer.com/article/10.1007/s11704-013-3024-3,True,,9.0,['genetic-programming'],,,['root-cause-analysis'],,,True,['failure-management'],,,['rca-others'],,,,,,
1022,Local versus global lessons for defect prediction and effort estimation,"['Tim Menzies', 'Andrew Butcher', 'David Cok', 'Andrian Marcus', 'Lucas Layman', 'Forrest Shull', 'Burak Turhan', 'Thomas Zimmermann']",2012,"Existing research is unclear on how to generate lessons learned for defect prediction and effort estimation. Should we seek lessons that are global to multiple projects or just local to particular projects? This paper aims to comparatively evaluate local versus global lessons learned for effort estimation and defect prediction. We applied automated clustering tools to effort and defect datasets from the PROMISE repository. Rule learners generated lessons learned from all the data, from local projects, or just from each cluster. The results indicate that the lessons learned after combining small parts of different data sources (i.e., the clusters) were superior to either generalizations formed over all the data or local lessons formed from particular projects. We conclude that when researchers attempt to draw lessons from some historical data source, they should 1) ignore any existing local divisions into multiple sources, 2) cluster across all available data, then 3) restrict the learning of lessons to the clusters from other sources that are nearest to the test data.",https://ieeexplore.ieee.org/document/6363444,True,,185.0,,,"['new-method', 'discussion']",['failure-prevention'],['software'],,True,['failure-management'],,,['software-defect-prediction'],,,,,,
1023,Evaluating defect prediction approaches: a benchmark and an extensive comparison,"['D’Ambros, Marco', 'Lanza, Michele', 'Robbes, Romain']",2012,"Reliably predicting software defects is one of the holy grails of software engineering. Researchers have devised and implemented a plethora of defect/bug prediction approaches varying in terms of accuracy, complexity and the input data they require. However, the absence of an established benchmark makes it hard, if not impossible, to compare approaches. We present a benchmark for defect prediction, in the form of a publicly available dataset consisting of several software systems, and provide an extensive comparison of well-known bug prediction approaches, together with novel approaches we devised. We evaluate the performance of the approaches using different performance indicators: classification of entities as defect-prone or not, ranking of the entities, with and without taking into account the effort to review an entity. We performed three sets of experiments aimed at (1) comparing the approaches across different systems, (2) testing whether the differences in performance are statistically significant, and (3) investigating the stability of approaches across different learners. Our results indicate that, while some approaches perform better than others in a statistically significant manner, external validity in defect prediction is still an open problem, as generalizing results to different contexts/learners proved to be a partially unsuccessful endeavor.",https://link.springer.com/article/10.1007/s10664-011-9173-9,True,,378.0,"['naive-bayes', 'logistic-regression', 'decision-tree']","['code-metrics', 'code-history']","['discussion', 'comparison']",['failure-prevention'],['source-code'],Empirical Software Engineering,True,['failure-management'],True,True,['software-defect-prediction'],True,10.0,10.1007/s10664-011-9173-9,,,
1024,"Defect prediction from static code features: current results, limitations, new approaches","['Menzies, Tim', 'Milton, Zach', 'Turhan, Burak', 'Cukic, Bojan', 'Jiang, Yue', 'Bener, Ay{\\c{s}}e']",2010,"Building quality software is expensive and software quality assurance (QA) budgets are limited. Data miners can learn defect predictors from static code features which can be used to control QA resources; e.g. to focus on the parts of the code predicted to be more defective. Recent results show that better data mining technology is not leading to better defect predictors. We hypothesize that we have reached the limits of the standard learning goal of maximizing area under the curve (AUC) of the probability of false alarms and probability of detection “AUC(pd, pf)”; i.e. the area under the curve of a probability of false alarm versus probability of detection. Accordingly, we explore changing the standard goal. Learners that maximize “AUC(effort, pd)” find the smallest set of modules that contain the most errors. WHICH is a meta-learner framework that can be quickly customized to different goals. When customized to AUC(effort, pd), WHICH out-performs all the data mining methods studied here. More importantly, measured in terms of this new goal, certain widely used learners perform much worse than simple manual methods. Hence, we advise against the indiscriminate use of learners. Learners must be chosen and customized to the goal at hand. With the right architecture (e.g. WHICH), tuning a learner to specific local business goals can be a simple task.",https://link.springer.com/article/10.1007/s10515-010-0069-5,True,,331.0,,,['discussion'],['failure-prevention'],['software'],,True,['failure-management'],,,['software-defect-prediction'],,,,,,
1025,Reducing Features to Improve Code Change-Based Bug Prediction,"['Shivkumar Shivaji', 'E. James Whitehead', 'Ram Akella', 'Sunghun Kim']",2012,"Machine learning classifiers have recently emerged as a way to predict the introduction of bugs in changes made to source code files. The classifier is first trained on software history, and then used to predict if an impending change causes a bug. Drawbacks of existing classifier-based bug prediction techniques are insufficient performance for practical use and slow prediction times due to a large number of machine learned features. This paper investigates multiple feature selection techniques that are generally applicable to classification-based bug prediction methods. The techniques discard less important features until optimal classification performance is reached. The total number of features used for training is substantially reduced, often to less than 10 percent of the original. The performance of Naive Bayes and Support Vector Machine (SVM) classifiers when using this technique is characterized on 11 software projects. Naive Bayes using feature selection provides significant improvement in buggy F-measure (21 percent improvement) over prior change classification bug prediction results (by the second and fourth authors [28]). The SVM's improvement in buggy F-measure is 9 percent. Interestingly, an analysis of performance for varying numbers of features shows that strong performance is achieved at even 1 percent of the original number of features.",https://ieeexplore.ieee.org/abstract/document/6226427,True,,190.0,"['support-vector-machine', 'naive-bayes']",,"['comparison', 'novel-use']",['failure-prevention'],['source-code'],,True,['failure-management'],,,['software-defect-prediction'],,,,,,
1026,A family of code coverage-based heuristics for effective fault localization,"['W. Eric Wong', 'Vidroha Debroy', 'Byoungju Choi']",2010,"Locating faults in a program can be very time-consuming and arduous, and therefore, there is an increased demand for automated techniques that can assist in the fault localization process. In this paper a code coverage-based method with a family of heuristics is proposed in order to prioritize suspicious code according to its likelihood of containing program bugs. Highly suspicious code (i.e., code that is more likely to contain a bug) should be examined before code that is relatively less suspicious; and in this manner programmers can identify and repair faulty code more efficiently and effectively. We also address two important issues: first, how can each additional failed test case aid in locating program faults; and second, how can each additional successful test case help in locating program faults. We propose that with respect to a piece of code, the contribution of the first failed test case that executes it in computing its likelihood of containing a bug is larger than or equal to that of the second failed test case that executes it, which in turn is larger than or equal to that of the third failed test case that executes it, and so on. This principle is also applied to the contribution provided by successful test cases that execute the piece of code. A tool, @gDebug, was implemented to automate the computation of the suspiciousness of the code and the subsequent prioritization of suspicious code for locating program faults. To validate our method case studies were performed on six sets of programs: Siemens suite, Unix suite, space, grep, gzip, and make. Data collected from the studies are supportive of the above claim and also suggest Heuristics III(a), (b) and (c) of our method can effectively reduce the effort spent on fault localization.",https://dl.acm.org/doi/10.1016/j.jss.2009.09.037,True,,37.0,['multilayer-perceptron'],,,['root-cause-analysis'],,Journal of Systems and Software,True,['failure-management'],,,,,,10.1016/j.jss.2009.09.037,,,
1027,Context-aware statistical debugging: from bug predictors to faulty control flow paths,"['Lingxiao Jiang', 'Zhendong Su']",2007,"Effective bug localization is important for realizing automated debugging. One attractive approach is to apply statistical techniques on a collection of evaluation profiles of program properties to help localize bugs. Previous research has proposed various specialized techniques to isolate certain program predicates as bug predictors. However, because many bugs may not be directly associated with these predicates, these techniques are often ineffective in localizing bugs. Relevant control flow paths that may contain bug locations are more informative than stand-alone predicates for discovering and understanding bugs. In this paper, we propose an approach to automatically generate such faulty control flow paths that link many bug predictors together for revealing bugs. Our approach combines feature selection (to accurately select failure-related predicates as bug predictors), clustering (to group correlated predicates), and control flow graph traversal in a novel way to help generate the paths. We have evaluated our approach on code including the Siemens test suite and rhythmbox (a large music management application for GNOME). Our experiments show that the faulty control flow paths are accurate, useful for localizing many bugs, and helped to discover previously unknown errors in rhythmbox",https://dl.acm.org/doi/10.1145/1321631.1321660,True,,65.0,['support-vector-machine'],,,['root-cause-analysis'],,ASE '07: Proceedings of the twenty-second IEEE/ACM international conference on Automated software engineering,True,['failure-management'],,,,,,10.1145/1321631.1321660,,,
1028,A Novel Framework for Locating Software Faults Using Latent Divergences,"['Roychowdhury, Shounak', 'Khurshid, Sarfraz']",2011,"Fault localization, ie, identifying erroneous lines of code in a buggy program, is a tedious process, which often requires considerable manual effort and is costly. Recent years have seen much progress in techniques for automated fault localization, specifically using program spectra–executions of failed and passed test runs provide a basis for isolating the faults. Despite the progress, fault localization in large programs remains a challenging problem, because even inspecting a small fraction of the lines of code in a large problem …",https://link.springer.com/chapter/10.1007/978-3-642-23808-6_4,True,,2.0,,,,['root-cause-analysis'],,Joint European Conference on Machine Learning and Knowledge Discovery in Databases,True,['failure-management'],,,,,,,,,
1029,Monitoring and diagnosing software requirements,"['Wang, Yiqiao', 'Mcilraith, Sheila A', 'Yu, Yijun', 'Mylopoulos, John']",2008,"We propose a framework adapted from Artificial Intelligence theories of action and diagnosis for monitoring and diagnosing failures of software requirements. Software requirements are specified using goal models where they are associated with preconditions and postconditions. The monitoring component generates log data that contains the truth values of specified pre/post-conditions, as well as system action executions. Such data can be generated at different levels of granularity, depending on diagnostic feedback. The …",https://link.springer.com/article/10.1007/s10515-008-0042-8,True,,75.0,"['rule-mining', 'constraint-solving']",['logs'],['new-method'],['root-cause-analysis'],,,True,['failure-management'],,,['fault-localization'],,,,,,
1030,Probabilistic fault diagnosis in communication systems through incremental hypothesis updating,"['M. Steinder', 'A. S. Sethi']",2004,"This paper presents a probabilistic event-driven fault localization technique, which uses a probabilistic symptom-fault map as a fault propagation model. The technique isolates the most probable set of faults through incremental updating of a symptom-explanation hypothesis. At any time, it provides a set of alternative hypotheses, each of which is a complete explanation of the set of symptoms observed thus far. The hypotheses are ranked according to a measure of their goodness. The technique allows multiple simultaneous independent faults to be identified and incorporates both negative and positive symptoms in the analysis. As shown in a simulation study, the technique offers close-to-optimal accuracy and is resilient both to noise in the symptom data and to inaccuracies of the probabilistic fault propagation model.",https://www.sciencedirect.com/science/article/pii/S1389128604000817,True,,90.0,['bayesian-network'],,,['root-cause-analysis'],,Computer Networks: The International Journal of Computer and Telecommunications Networking,True,['failure-management'],,,,,,10.1016/j.comnet.2004.01.007,,,
1031,"Gestalt: Fast, Unified Fault Localization for Networked Systems","['Radhika Niranjan Mysore', 'Ratul Mahajan', 'Amin Vahdat', 'George Varghese']",2014,"We show that the performance of existing fault localization algorithms differs markedly for different networks; and no algorithm simultaneously provides high localization accuracy and low computational overhead. We develop a framework to explain these behaviors by anatomizing the algorithms with respect to six important characteristics of real networks, such as uncertain dependencies, noise, and covering relationships. We use this analysis to develop Gestalt, a new algorithm that combines the best elements of existing ones and includes a new technique to explore the space of fault hypotheses. We run experiments on three real, diverse networks. For each, Gestalt has either significantly higher localization accuracy or an order of magnitude lower running time. For example, when applied to the Lync messaging system that is used widely within corporations, Gestalt localizes faults with the same accuracy as Sherlock, while reducing fault localization time from days to 23 seconds.",https://dl.acm.org/doi/10.5555/2643634.2643662,True,,31.0,,,,['root-cause-analysis'],,USENIX ATC'14: Proceedings of the 2014 USENIX conference on USENIX Annual Technical Conference,True,['failure-management'],,,,,,10.5555/2643634.2643662,,,
1032,Fault detection and localization in distributed systems using invariant relationships,"['Abhishek B. Sharma', 'Haifeng Chen', 'Min Ding', 'Kenji Yoshihira', 'Guofei Jiang']",2013,"Recent advances in sensing and communication technologies enable us to collect round-the-clock monitoring data from a wide-array of distributed systems including data centers, manufacturing plants, transportation networks, automobiles, etc. Often this data is in the form of time series collected from multiple sensors (hardware as well as software based). Previously, we developed a time-invariant relationships based approach that uses Auto-Regressive models with eXogenous input (ARX) to model this data. A tool based on our approach has been effective for fault detection and capacity planning in distributed systems. In this paper, we first describe our experience in applying this tool in real-world settings. We also discuss the challenges in fault localization that we face when using our tool, and present two approaches - a spatial approach based on invariant graphs and a temporal approach based on expected broken invariant patterns - that we developed to address this problem.",https://ieeexplore.ieee.org/document/6575304,True,,45.0,,,,['root-cause-analysis'],,2013 43rd Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN),True,['failure-management'],,,,,,,,,
1033,The DStar Method for Effective Software Fault Localization,"['W. Eric Wong', 'Vidroha Debroy', 'Ruizhi Gao', 'Yihao Li']",2013,"Effective debugging is crucial to producing reliable software. Manual debugging is becoming prohibitively expensive, especially due to the growing size and complexity of programs. Given that fault localization is one of the most expensive activities in program debugging, there has been a great demand for fault localization techniques that can help guide programmers to the locations of faults. In this paper, a technique named DStar (D*) is proposed which can suggest suspicious locations for fault localization automatically without requiring any prior information on program structure or semantics. D* is evaluated across 24 programs, and is compared to 38 different fault localization techniques. Both single-fault and multi-fault programs are used. Results indicate that D* is more effective at locating faults than all the other techniques it is compared to. An empirical evaluation is also conducted to illustrate how the effectiveness of D* increases as the exponent * grows, and then levels off when the exponent * exceeds a critical value. Discussions are presented to support such observations.",https://ieeexplore.ieee.org/document/6651713,True,,158.0,,['test-cases'],,['root-cause-analysis'],['source-code'],,True,['failure-management'],True,,['fault-localization'],True,77.0,,,,['no-knowledge-required']
1034,Experimental Evaluation of Hybrid Algorithm in Spectrum based Fault Localization,"['Park, A Jonghee', 'Kim, B Jeongho', 'Lee, C Eunseok']",2014,"During debugging process in software development cycle, fault localization is inevitable work. Diverse approaches have been proposed, such as program slicing, machine learning, and data mining for fault localization. In this paper we propose an effective hybrid fault localization algorithm based on a spectrum that enables fault detection in every statement. This algorithm distinguishes the location of a bug that causes a false positive score through the relationship between a test case and statement hit information. We also provide a fault …",https://www.semanticscholar.org/paper/Experimental-Evaluation-of-Hybrid-Algorithm-in-Park-Kim/93d2495000cec06dd3a88a62dff218a896bcc3ea,True,,1.0,,,,['root-cause-analysis'],,Proceedings of the International Conference on Software Engineering Research and Practice (SERP),True,['failure-management'],,,,,,,,,
1035,Design and Realization of a Network Fault Diagnosing Expert System Based on Genetic Algorithm,"['DAI, Zhong-jian', 'SU, Li-min']",2005,"An intelligent fault diagnosis expert system is developed by using object-oriented approach and Visual (C++ 6.0) language. The frame of system and the function of each module are also introduced. In the system the theory of genetic algorithm is presented to process the problem of knowledge acquisition, and expert system is used for network fault diagnosis, this improves the system capability of restraining noises, and makes system be able to do self-learning. The results of the experiment on the campus network show that the system can …",http://en.cnki.com.cn/Article_en/CJFDTotal-BJLG200501010.htm,True,,4.0,['genetic-programming'],,,['root-cause-analysis'],['network'],,True,['failure-management'],,,['fault-localization'],,,,,,
1036,Network Fault Diagnoses Expert System Model Based on Decision Tree,"['Qu, ZY', 'Gao, YF', 'Nie, Xin']",2008,"In view of inherent insufficiency of traditional network fault diagnoses expert system in the aspect of knowledge acquisition,together with the characteristics of decision tree algorithm and expert system,this paper puts forward a new network fault diagnoses expert system model based on decision tree,discusses the reasoning mechanism,the reasoning algorithm,and the method of knowledge acquisition of the model. Experimental results demonstrate that this model has the characteristics of simple process,low computation complexity and strong adaptability.",https://www.semanticscholar.org/paper/Network-Fault-Diagnoses-Expert-System-Model-Based-Zhao-yang/41f7a572a66b38f135328a1d9c7d62c44e08c232,True,,9.0,,,,['root-cause-analysis'],,,True,['failure-management'],,,,,,,,,
1037,Network application layer fault detection system based on SVM,"['Li, Qianmu', 'Xu, Manwu', 'Zhang, Hong', 'Liu, Fengyu']",2006,"A framework of SVM based network fault detection system of application layer was proposed. The function, mechanism and realization of the components of this framework were discussed. By means of distance metric of heterogeneous datasets, the feature data of network were preprocessed. Based on guaranteed estimators, the size of test set was estimated. Thus the bad train result for lack of examples was not only avoided, but the training time was also reduced and the efficiency of training was improved. During the training, by means of fuzzy mathematics, considering the effect of different network data features to the classification, a weight method was brought forward. It improved the accuracy of network fault detection. The problem of low detection accuracy of some types of faults for the imbalance of training examples was researched. A method of increasing the proportion of the examples of these types was proposed. It improved the detection accuracy of these types of faults.",https://www.researchgate.net/publication/296609986_Network_application_layer_fault_detection_system_based_on_SVM,True,,2.0,['support-vector-machine'],,['novel-use'],['failure-detection'],"['application', 'network']",,True,['failure-management'],,,,,,,,,
1038,Application of rough set theory in network fault diagnosis,"['Yuqing Peng', 'Gengqian Liu', 'Tao Lin', 'Hengshan Geng']",2005,"In this paper rough set theory is researched and applied in computer network fault diagnosis. Original MIB( information base of management) data from network which reflect network fault are collected first, and a reduction algorithm based on attribute significance and attribute frequency is implemented on the MIB data, which removing inconsistent or erroneous MIB data. Based on attribute core and user preference attribute set, the algorithm makes not only use of advantage of these two algorithm, but also the universality of core, user background knowledge, and domain experience. At the same time, the minimal support degree and minimal belief degree is introduced into rough set theory for decision rules discovery and get decision rules.",https://ieeexplore.ieee.org/document/1489022,True,,12.0,['rule-mining'],,,['root-cause-analysis'],['network'],,True,['failure-management'],,,['fault-localization'],,,,,,
1039,An Artificial Intelligence Approach to Network Fault Management,"['G{\\""u}rer, Denise W', 'Khan, Irfan', 'Ogier, Richard', 'Keffer, Renee']",1996,"Traditionally, network management activities, such as fault management, have been performed with direct human involvement. However, these activities are becoming more demanding and data intensive, due to the heterogeneous nature and increasing size of networks today. For these reasons, it is becoming necessary to automate network management activities. Artificial intelligence technologies can play an important role in the problem solving and reasoning techniques that are employed in fault management. Expert systems have been successfully applied to some types of fault management. However, these systems are not flexible enough for today's evolving network needs. We propose a hybrid AI solution that employs both neural networks and case-based reasoning techniques for the fault management of heterogeneous distributed information networks.",https://www.researchgate.net/publication/2731474_An_Artificial_Intelligence_Approach_to_Network_Fault_Management,True,,101.0,"['multilayer-perceptron', 'case-based-reasoning']",,,['root-cause-analysis'],,,True,['failure-management'],,,,,,,,,
1040,Detection and Localization of Network Black Holes,"['R. R. Kompella', 'J. Yates', 'A. Greenberg', 'A. C. Snoeren']",2007,"Internet backbone networks are under constant flux, struggling to keep up with increasing demand. The pace of technology change often outstrips the deployment of associated fault monitoring capabilities that are built into today's IP protocols and routers. Moreover, some of these new technologies cross networking layers, raising the potential for unanticipated interactions and service disruptions that the built-in monitoring systems cannot detect. In such instances, failures may cause data packets to be silently dropped inside the network without triggering any alarms or responses (e.g., the failure is not routed around). So-called ""silent failures"" or ""black holes"" represent a critical threat to today's rapidly evolving networks. In this paper, we present a simple and effective method to detect and diagnose such silent failures. Our method uses active measurement between edge routers to raise alarms whenever end-to-end connectivity is disrupted, regardless of the cause. These alarms feed localization agents that employ spatial correlation techniques to isolate the root-cause of failure. Using data from two real systems deployed on sections of a tier-I ISP network, we successfully detect and localize three known black holes. Further, we present simulation results demonstrating that our system accurately and precisely (both greater than 80% according to our metrics) localizes a variety of failures classes.",https://ieeexplore.ieee.org/document/4215834,True,,182.0,"['correlation', 'graph-mining']",[],['new-method'],"['root-cause-analysis', 'failure-detection']",['network'],IEEE INFOCOM 2007-26th IEEE International Conference on Computer Communications,True,['failure-management'],,,"['anomaly-detection', 'fault-localization', 'root-cause-diagnosis']",,,,,,
1041,Streaming pattern discovery in multiple time-series,"['Spiros Papadimitriou', 'Jimeng Sun', 'Christos Faloutsos']",2005,"In this paper, we introduce SPIRIT (Streaming Pattern dIscoveRy in multIple Time-series). Given n numerical data streams, all of whose values we observe at each time tick t, SPIRIT can incrementally find correlations and hidden variables, which summarise the key trends in the entire stream collection. It can do this quickly, with no buffering of stream values and without comparing pairs of streams. Moreover, it is any-time, single pass, and it dynamically detects changes. The discovered trends can also be used to immediately spot potential anomalies, to do efficient forecasting and, more generally, to dramatically simplify further data processing. Our experimental evaluation and case studies show that SPIRIT can incrementally capture correlations and discover trends, efficiently and effectively.",https://dl.acm.org/doi/10.5555/1083592.1083674,True,,95.0,['dimensionality-reduction'],,,"['root-cause-analysis', 'failure-detection']",,VLDB '05: Proceedings of the 31st international conference on Very large data bases,True,['failure-management'],,,['anomaly-detection'],,,10.5555/1083592.1083674,,,
1042,Minerals: using data mining to detect router misconfigurations,"['Franck Le', 'Sihyung Lee', 'Tina Wong', 'Hyong S. Kim', 'Darrell Newcomb']",2006,"Recent studies have shown that router misconfigurations are common and have dramatic consequences for the operations of networks. Not only can misconfigurations compromise the security of a single network, they can even cause global disruptions in Internet connectivity. Several solutions have been proposed that can detect a number of problems in real configuration files. However, these solutions share a common limitation: they are rule-based. Rules are assumed to be known beforehand, and violations of these rules are deemed misconfigurations. As policies typically differ among networks, rule-based approaches are limited in the scope of mistakes they can detect. In this paper, we address the problem of router misconfigurations using data mining. We apply association rules mining to the configuration files of routers across an administrative domain to discover local, network-specific policies. Deviations from these local policies are potential misconfigurations. We have evaluated our scheme on configuration files from a large state-wide network provider, a large university campus and a high-performance research network, and found promising results. We discovered a number of errors that were confirmed and later corrected by the network engineers. These errors would have been difficult to detect with current rule-based approaches.",https://dl.acm.org/doi/10.1145/1162678.1162681,True,,34.0,['rule-mining'],,,['root-cause-analysis'],['router'],MineNet '06: Proceedings of the 2006 SIGCOMM workshop on Mining network data,True,['failure-management'],,,,,,10.1145/1162678.1162681,,,
1043,Discovering actionable patterns in event data,"['J. L. Hellerstein', 'S. Ma', 'C.-S. Perng']",2002,"Applications such as those for systems management and intrusion detection employ an automated real-time operation system in which sensor data are collected and processed in real time. Although such a system effectively reduces the need for operation staff, it requires constructing and maintaining correlation rules. Currently, rule construction requires experts to identify problem patterns, a process that is time-consuming and error-prone. In this paper, we propose reducing this burden by mining historical data that are readily available. Specifically, we first present efficient algorithms to mine three types of important patterns from historical event data: event bursts, periodic patterns, and mutually dependent patterns. We then discuss a framework for efficiently mining events that have multiple attributes. Last, we present Event Correlation Constructor—a tool that validates and extends correlation knowledge.",https://ieeexplore.ieee.org/document/5386872,True,,131.0,['rule-mining'],"['events', 'logs']",,['root-cause-analysis'],,,True,['failure-management'],,,['root-cause-diagnosis'],,,,,,
1044,Predicting Bugs in Source Code Changes with Incremental Learning Method,"['Yuan, Zi', 'Yu, Lili', 'Liu, Chao', 'Zhang, Linghua']",2013,"Software is constructed by a series of changes and each change has the risk to introduce bugs. Predicting the existence of bugs in source code changes could help developers detect and fix bugs immediately upon the completion of a change, which accelerates the bug fixing process and save the limited time and human resources effectively. However, because of altering nature in the underlying bug generation process, the concept used to depict the bug introducing patterns is drifting, which makes it difficult to predict latent bugs of source code changes accurately, especially in the long-term prediction scenario. In order to deal with this problem, a feature-based incremental learning framework is proposed. It is comprised of three components:(1) an incremental discretization method, which is used to transform the quantitive features in the corpus incrementally, (2) an incremental feature selection method, which is always keeping a subset with the most informative features, and (3) an incremental classification algorithm, which updates the classifier dynamically and considers the current best subset of features during prediction. This proposed approach is evaluated on three famous open source systems, Eclipse, Mozilla and jedit. The results show that our approach performs better than the non-incremental method in dealing with concept drift, with the consideration of keeping the value of both precision and recall stable at a suitable level over time. We also implement a prototype with this learning framework and apply it to a real software development scenario.",https://www.semanticscholar.org/paper/Predicting-Bugs-in-Source-Code-Changes-with-Method-Yuan-Yu/fd19980574d03350fd4dfc8e1c0b2983a1928e20,True,,4.0,['naive-bayes'],"['code-history', 'source-code']",,['failure-prediction'],['software'],,True,['failure-management'],,,,,,,,,
1045,Efficient ticket routing by resolution sequence mining,"['Qihong Shao', 'Yi Chen', 'Shu Tao', 'Xifeng Yan', 'Nikos Anerousis']",2008,"IT problem management calls for quick identification of resolvers to reported problems. The efficiency of this process highly depends on ticket routing---transferring problem ticket among various expert groups in search of the right resolver to the ticket. To achieve efficient ticket routing, wise decision needs to be made at each step of ticket transfer to determine which expert group is likely to be, or to lead to the resolver. In this paper, we address the possibility of improving ticket routing efficiency by mining ticket resolution sequences alone, without accessing ticket content. To demonstrate this possibility, a Markov model is developed to statistically capture the right decisions that have been made toward problem resolution, where the order of the Markov model is carefully chosen according to the conditional entropy obtained from ticket data. We also design a search algorithm, called Variable-order Multiple active State search(VMS), that generates ticket transfer recommendations based on our model. The proposed framework is evaluated on a large set of real-world problem tickets. The results demonstrate that VMS significantly improves human decisions: Problem resolvers can often be identified with fewer ticket transfers.",https://dl.acm.org/doi/10.1145/1401890.1401964,True,,89.0,"['search', 'markov-model']",['tickets'],,['remediation'],,KDD '08: Proceedings of the 14th ACM SIGKDD international conference on Knowledge discovery and data mining,True,['failure-management'],True,,['triage'],True,87.0,10.1145/1401890.1401964,,,
1046,Optimal online deterministic algorithms and adaptive heuristics for energy and performance efficient dynamic consolidation of virtual machines in Cloud data centers,"['Anton Beloglazov', 'Rajkumar Buyya']",2012,"The rapid growth in demand for computational power driven by modern service applications combined with the shift to the Cloud computing model have led to the establishment of large-scale virtualized data centers. Such data centers consume enormous amounts of electrical energy resulting in high operating costs and carbon dioxide emissions. Dynamic consolidation of virtual machines (VMs) using live migration and switching idle nodes to the sleep mode allows Cloud providers to optimize resource usage and reduce energy consumption. However, the obligation of providing high quality of service to customers leads to the necessity in dealing with the energy-performance trade-off, as aggressive consolidation may lead to performance degradation. Because of the variability of workloads experienced by modern applications, the VM placement should be optimized continuously in an online manner. To understand the implications of the online nature of the problem, we conduct a competitive analysis and prove competitive ratios of optimal online deterministic algorithms for the single VM migration and dynamic VM consolidation problems. Furthermore, we propose novel adaptive heuristics for dynamic consolidation of VMs based on an analysis of historical data from the resource usage by VMs. The proposed algorithms significantly reduce energy consumption, while ensuring a high level of adherence to the service level agreement. We validate the high efficiency of the proposed algorithms by extensive simulations using real-world workload traces from more than a thousand PlanetLab VMs. Copyright © 2011 John Wiley & Sons, Ltd.",https://dl.acm.org/doi/10.1002/cpe.1867,True,,1419.0,"['linear-regression', 'optimization']","['kpis', 'sla', 'host-metrics']",['new-method'],"['resource-consolidation', 'power-management']",['vm'],Concurrency and Computation: Practice & Experience,True,['resource-provisioning'],,,,,,10.1002/cpe.1867,,,
1047,Dynamic Virtual Machine Management via Approximate Markov Decision Process,"['Zhenhua Han', 'Haisheng Tan', 'Guihai Chen', 'Rui Wang', 'Yifan Chen', 'Francis C.M. Lau']",2016,"Efficient virtual machine (VM) management can dramatically reduce energy consumption in data centers. Existing VM management algorithms fall into two categories based on whether the VMs' resource demands are assumed to be static or dynamic. The former category fails to maximize the resource utilization as they cannot adapt to the dynamic nature of VMs' resource demands. Most approaches in the latter category are heuristical and lack theoretical performance guarantees. In this work, we formulate dynamic VM management as a large-scale Markov Decision Process (MDP) problem and derive an optimal solution. Our analysis of real-world data traces supports our choice of the modeling approach. However, solving the large-scale MDP problem suffers from the curse of dimensionality. Therefore, we further exploit the special structure of the problem and propose an approximate MDP-based dynamic VM management method, called MadVM. We prove the convergence of MadVM and analyze the bound of its approximation error. Moreover, MadVM can be implemented in a distributed system, which should suit the needs of real data centers. Extensive simulations based on two real-world workload traces show that MadVM achieves significant performance gains over two existing baseline approaches in power consumption, resource shortage and the number of VM migrations. Specifically, the more intensely the resource demands fluctuate, the more MadVM outperforms.",https://ieeexplore.ieee.org/document/7524384,True,,43.0,['decision-process'],['metrics'],,['power-management'],['vm'],,True,['resource-provisioning'],,,,,,,,,
1048,FChain: Toward Black-Box Online Fault Localization for Cloud Systems,"['Hiep Nguyen', 'Zhiming Shen', 'Yongmin Tan', 'Xiaohui Gu']",2013,"Distributed applications running inside cloud systems are prone to performance anomalies due to various reasons such as resource contentions, software bugs, and hardware failures. One big challenge for diagnosing an abnormal distributed application is to pinpoint the faulty components. In this paper, we present a black-box online fault localization system called FChain that can pinpoint faulty components immediately after a performance anomaly is detected. FChain first discovers the onset time of abnormal behaviors at different components by distinguishing the abnormal change point from many change points caused by normal workload fluctuations. Faulty components are then pinpointed based on the abnormal change propagation patterns and inter-component dependency relationships. FChain performs runtime validation to further filter out false alarms. We have implemented FChain on top of the Xen platform and tested it using several benchmark applications (RUBiS, Hadoop, and IBM System S). Our experimental results show that FChain can quickly pinpoint the faulty components with high accuracy within a few seconds. FChain can achieve up to 90% higher precision and 20% higher recall than existing schemes. FChain is non-intrusive and light-weight, which imposes less than 1% overhead to the cloud system.",https://ieeexplore.ieee.org/document/6681572,True,,69.0,"['markov-model', 'case-based-reasoning']","['network-metrics', 'kpis', 'host-metrics']","['comparison', 'new-method']","['failure-detection', 'root-cause-analysis']","['rubis', 'hadoop', 'xen']",2013 IEEE 33rd International Conference on Distributed Computing Systems,True,['failure-management'],True,,"['anomaly-detection', 'fault-localization']",True,70.0,,,,
1049,Workload-aware anomaly detection for Web applications,"['Wang, Tao', 'Wei, Jun', 'Zhang, Wenbo', 'Zhong, Hua', 'Huang, Tao']",2014,"The failure of Web applications often affects a large population of customers, and leads to severe economic loss. Anomaly detection is essential for improving the reliability of Web applications. Current approaches model correlations among metrics, and detect anomalies when the correlations are broken. However, dynamic workloads cause the metric correlations to change over time. Moreover, modeling various metric correlations are difficult in complex Web applications. This paper addresses these problems and proposes an online …",https://www.sciencedirect.com/science/article/pii/S0164121213000721,True,,37.0,['clustering'],,,['failure-detection'],,,True,['failure-management'],,,"['anomaly-detection', 'workload-prediction']",,,10.1016/j.jss.2013.03.060,,,
1050,adaptive workload prediction of grid performance in confidence windows,"['Yongwei Wu', 'Kai Hwang', 'Yulai Yuan', 'Weiming Zheng']",2009,"Predicting grid performance is a complex task because heterogeneous resource nodes are involved in a distributed environment. Long execution workload on a grid is even harder to predict due to heavy load fluctuations. In this paper, we use Kalman filter to minimize the prediction errors. We apply Savitzky-Golay filter to train a sequence of confidence windows. The purpose is to smooth the prediction process from being disturbed by load fluctuations. We present a new adaptive hybrid method (AHModel) for load prediction guided by trained confidence windows. We test the effectiveness of this new prediction scheme with real-life workload traces on the AuverGrid and Grid5000 in France. Both theoretical and experimental results are reported in this paper. As the lookahead span increases from 10 to 50 steps (5 minutes per step), the AHModel predicts the grid workload with a mean-square error (MSE) of 0.04-0.73 percent, compared with 2.54-30.2 percent in using the static point value autoregression (AR) prediction method. The significant gain in prediction accuracy makes the new model very attractive to predict Grid performance. The model was proved especially effective to predict large workload that demands very long execution time, such as exceeding 4 hours on the Grid5000 over 5,000 processors. With minor changes of some system parameters, the AHModel can apply to other computational grids as well. At the end, we discuss extended research issues and tool development for Grid performance prediction.",https://ieeexplore.ieee.org/document/5226619,True,,72.0,,,,['workload-prediction'],,,True,['resource-provisioning'],,,,,,,,,
1051,multi-model prediction for enhancing content locality in elastic server infrastructures,"['Tirado, Juan M', 'Higuero, Daniel', 'Isaila, Florin', 'Carretero, Jesus']",2011,"Infrastructures serving on-line applications experience dynamic workload variations depending on diverse factors such as popularity, marketing, periodic patterns, fads, trends, events, etc. Some predictable factors such as trends, periodicity or scheduled events allow for proactive resource provisioning in order to meet fluctuations in workloads. However, proactive resource provisioning requires prediction models forecasting future workload patterns. This paper proposes a multi-model prediction approach, in which data are grouped …",https://ieeexplore.ieee.org/document/6152728,True,,23.0,,,,,,2011 18th International Conference on High Performance Computing,True,['resource-provisioning'],,,,,,,,,
1052,Failure prediction for HPC systems and applications: Current situation and open issues,"['Gainaru, Ana', 'Cappello, Franck', 'Snir, Marc', 'Kramer, William']",2013,"As large-scale systems evolve towards post-petascale computing, it is crucial to focus on providing fault-tolerance strategies that aim to minimize fault's effects on applications. By far the most popular technique is the checkpoint–restart strategy. A complement to this classical approach is failure avoidance, by which the occurrence of a fault is predicted and proactive measures are taken. This requires a reliable prediction system to anticipate failures and their locations. One way of offering prediction is by the analysis of system logs generated during …",https://journals.sagepub.com/doi/abs/10.1177/1094342013488258,True,,40.0,,,"['new-method', 'discussion']",['failure-prediction'],['hpc'],,True,['failure-management'],,,['system-failure-prediction'],,,,,,
1053,Predicting Failures with Hidden Markov Models,"['Salfner, Felix']",2005,"A key challenge for proactive handling of faults is the prediction of system failures. The main principle of the approach presented here is to identify and recognize patterns of errors that lead to failures. I propose the use of hidden Markov models (HMMs) as they have been successfully used in other pattern recognition tasks. The paper further motivates their use, explains how HMMs can be used to predict failures and describes the training procedure. An outlook to a more sophisticated treatment of time between events is also presented.",https://www.semanticscholar.org/paper/Predicting-Failures-with-Hidden-Markov-Models-Salfner/a45b45f008c4a30fd9c20c6b6c39d38db5b34199,True,,32.0,['markov-model'],['events'],['novel-use'],['failure-prediction'],,Proceedings of 5th European Dependable Computing Conference (EDCC-5),True,['failure-management'],,,['system-failure-prediction'],,,,,,
1054,Bayesian Neural Networks for Internet Traffic Classification,"['Tom Auld', 'Andrew W. Moore', 'Stephen F. Gull']",2007,"Internet traffic identification is an important tool for network management. It allows operators to better predict future traffic matrices and demands, security personnel to detect anomalous behavior, and researchers to develop more realistic traffic models. We present here a traffic classifier that can achieve a high accuracy across a range of application types without any source or destination host-address or port information. We use supervised machine learning based on a Bayesian trained neural network. Though our technique uses training data with categories derived from packet content, training and testing were done using features derived from packet streams consisting of one or more packet headers. By providing classification without access to the contents of packets, our technique offers wider application than methods that require full packet/payloads for classification. This is a powerful advantage, using samples of classified traffic to permit the categorization of traffic based only upon commonly available information",https://ieeexplore.ieee.org/document/4049810,True,,545.0,['multilayer-perceptron'],"['packet-content', 'network-metrics']",['new-method'],['failure-detection'],['network'],,True,['failure-management'],True,,['traffic-classification'],True,65.0,,,,
1055,A Fuzzy Controller for Self-Adaptive Lightweight Edge Container Orchestration,"['Gand, Fabian', 'Fronza, Ilenia', 'El Ioini, Nabil', 'Barzegar, Hamid R', 'Azimi, Shelernaz', 'Pahl, Claus']",2020,"Edge clusters consisting of small and affordable single-board devices are used in a range of different applications such as microcontrollers regulating an industrial process or controllers monitoring and managing traffic roadside. We call this wider context of computational infrastructure between the sensor and Internet-of-Things world and centralised cloud data centres the edge or edge computing. Despite the growing hardware capabilities of edge devices, resources are often still limited and need to be used intelligently. This can achieved …",https://www.researchgate.net/publication/339747802_A_Fuzzy_Controller_for_Self-Adaptive_Lightweight_Edge_Container_Orchestration,True,,1.0,['fuzzy-logic'],,,['resource-consolidation'],['container'],International Conference on Cloud Computing and Services Science CLOSER,True,['resource-provisioning'],False,,,,,,,,
1056,Anomaly Detection and Analysis for Clustered Cloud Computing Reliability,"['Samir, Areeg', 'Pahl, Claus']",2019,"Cloud and edge computing allow applications to be deployed and managed through third-party provided services that typically make virtualised resources available. However, often there is no direct insight into execution parameters at resource level, and only some quality factors can be directly observed while others remain hidden from the consumer. We investigate a framework for autonomous anomaly analysis for clustered cloud or edge resources. The framework determines possible causes of consumer-observed anomalies in …",https://www.researchgate.net/publication/333235823_Anomaly_Detection_and_Analysis_for_Clustered_Cloud_Computing_Reliability,True,,7.0,"['clustering', 'markov-model']",,['new-method'],"['root-cause-analysis', 'failure-detection']",,,True,['failure-management'],,,"['anomaly-detection', 'root-cause-diagnosis']",,,,,,
1057,DLA: Detecting and Localizing Anomalies in Containerized Microservice Architectures Using Markov Models,"['Samir, Areeg', 'Pahl, Claus']",2019,"Container-based microservice architectures are emerging as a new approach for building distributed applications as a collection of independent services that works together. As a result, with microservices, we are able to scale and update their applications based on the load attributed to each service. Monitoring and managing the load in a distributed system is a complex task as the degradation of performance within a single service will cascade reducing the performance of other dependent services. Such performance degradations …",https://www.researchgate.net/publication/333994728_DLA_Detecting_and_Localizing_Anomalies_in_Containerized_Microservice_Architectures_Using_Markov_Models,True,,0.0,['markov-model'],"['sla', 'host-metrics']",['new-method'],"['root-cause-analysis', 'failure-detection']","['kubernetes', 'container', 'docker']",2019 7th International Conference on Future Internet of Things and Cloud (FiCloud),True,['failure-management'],,,['anomaly-detection'],,,,,,
1058,Flow Clustering Using Machine Learning Techniques,"['McGregor, Anthony', 'Hall, Mark', 'Lorier, Perry', 'Brunskill, James']",2004,"Packet header traces are widely used in network analysis. Header traces are the aggregate of traffic from many concurrent applications. We present a methodology, based on machine learning, that can break the trace down into clusters of traffic where each cluster has different traffic characteristics. Typical clusters include bulk transfer, single and multiple transactions and interactive traffic, amongst others. The paper includes a description of the methodology, a visualisation of the attribute statistics that aids in recognising cluster types and a discussion of the stability and effectiveness of the methodology.",https://link.springer.com/chapter/10.1007/978-3-540-24668-8_21,True,,665.0,['clustering'],['packet-content'],['novel-use'],['failure-detection'],['network'],International workshop on passive and active network measurement,True,['failure-management'],,,['traffic-classification'],,,10.1007/978-3-540-24668-8_21,,,
1059,Availability modeling and analysis on high performance cluster computing systems,"['Song, Hertong', 'Leangsuksun, Chokchai', 'Nassar, Raja', 'Gottumukkala, NR', 'Scott, S']",2006,"Cluster computing has been attracting more and more attention from both the industry and the academia for its enormous computing power, cost effectiveness, and scalability. Availability is a key system attribute that needs to be considered both at system design stage and must reflect the actuality. System monitoring and logging enables identifying unplanned events to reflect the actual system's availability. This paper proposes a single framework that coordinates event monitoring, filtering, data analysis and dynamic availability modeling. The …",https://ieeexplore.ieee.org/document/1625325,True,,33.0,['markov-model'],['logs'],,"['root-cause-analysis', 'failure-detection', 'root-cause-diagnosis', 'anomaly-detection']",,"First International Conference on Availability, Reliability and Security (ARES'06)",True,"['resource-provisioning', 'failure-management']",,,,,,,,,
1060,An Explainable Deep Model for Defect Prediction,"['Humphreys, Jack', 'Dam, Hoa Khanh']",2019,"Self attention transformer encoders represent an effective method for sequence to class prediction tasks as they can disentangle long distance dependencies and have many regularising effects. We achieve results substantially better than state of the art in one such task, namely, defect prediction and with many added benefits. Existing techniques do not normalise for correlations that are inversely proportional to the usefulness of the prediction but do, in fact, go further, specifically exploiting these features which is tantamount to data …",https://ieeexplore.ieee.org/document/8823688,True,,1.0,['rnn'],['code-metrics'],"['comparison', 'novel-use']",['failure-prevention'],['source-code'],2019 IEEE/ACM 7th International Workshop on Realizing Artificial Intelligence Synergies in Software Engineering (RAISE),True,['failure-management'],,,['software-defect-prediction'],,,,,,
1061,What happened in my network: mining network events from router syslogs,"['Qiu, Tongqing', 'Ge, Zihui', 'Pei, Dan', 'Wang, Jia', 'Xu, Jun']",2010,"Router syslogs are messages that a router logs to describe a wide range of events observed by it. They are considered one of the most valuable data sources for monitoring network health and for trou-bleshooting network faults and performance anomalies. However, router syslog messages are essentially free-form text with only a minimal structure, and their formats vary among different vendors and router OSes. Furthermore, since router syslogs are aimed for tracking and debugging router software/hardware problems, they are often too …",https://dl.acm.org/doi/10.1145/1879141.1879202,True,,78.0,,['logs'],,['root-cause-analysis'],"['network', 'router']",Proceedings of the 10th ACM SIGCOMM conference on Internet measurement,True,['failure-management'],,,"['root-cause-diagnosis', 'log-enhancement']",,,10.1145/1879141.1879202,,,
1062,Task Failure Prediction in Cloud Data Centers Using Deep Learning,"['Gao, Jiechao', 'Wang, Haoyu', 'Shen, Haiying']",2019,"A large-scale cloud data center needs to provide high service reliability and availability with low failure occurrence probability. However, current large-scale cloud data centers still face high failure rates due to many reasons such as hardware and software failures, which often result in task and job failures. Such failures can severely reduce the reliability of cloud services and also occupy huge amount of resources to recover the service from failures. Therefore, it is important to predict task or job failures before occurrence with high accuracy to avoid unexpected wastage. Many machine learning and deep learning based methods have been proposed for the task or job failure prediction by analyzing past system message logs and identifying the relationship between the data and the failures. In order to further improve the failure prediction accuracy of the previous machine learning and deep learning based methods, in this paper, we propose a failure prediction algorithm based on multi-layer Bidirectional Long Short Term Memory (Bi-LSTM) to identify task and job failures in the cloud. The goal of Bi-LSTM prediction algorithm is to predict whether the tasks and jobs are failed or completed. The trace-driven experiments show that our algorithm outperforms other state-of-art prediction methods with 93% accuracy and 87% for task failure and job failures respectively.",https://ieeexplore.ieee.org/abstract/document/9006011,True,,0.0,['rnn'],,['novel-use'],['failure-prediction'],['job'],2019 IEEE International Conference on Big Data (Big Data),True,['failure-management'],,,['system-failure-prediction'],,,,,,
1063,Predicting Disk Failures with HMM- and HSMM-Based Approaches,"['Zhao, Ying', 'Liu, Xiang', 'Gan, Siqing', 'Zheng, Weimin']",2010,"Understanding and predicting disk failures are essential for both disk vendors and users to manufacture more reliable disk drives and build more reliable storage systems, in order to avoid service downtime and possible data loss. Predicting disk failure from observable disk attributes, such as those provided by the Self-Monitoring and Reporting Technology (SMART) system, has been shown to be effective. In the paper, we treat SMART data as time series, and explore the prediction power by using HMM- and HSMM-based approaches. Our experimental results show that our prediction models outperform other models that do not capture the temporal relationship among attribute values over time. Using the best single attribute, our approach can achieve a detection rate of 46% at 0% false alarm. Combining the two best attributes, our approach can achieve a detection rate of 52% at 0% false alarm.",https://link.springer.com/chapter/10.1007/978-3-642-14400-4_30,True,,48.0,['markov-model'],['host-metrics'],['novel-use'],['failure-prediction'],['hard-drive'],Industrial Conference on Data Mining,True,['failure-management'],True,,['hardware-failure-prediction'],True,27.0,,,,
1064,Mining system logs to learn error predictors: a case study of a telemetry system,"['Russo, Barbara', 'Succi, Giancarlo', 'Pedrycz, Witold']",2014,"Predicting system failures can be of great benefit to managers that get a better command over system performance. Data that systems generate in the form of logs is a valuable source of information to predict system reliability. As such, there is an increasing demand of tools to mine logs and provide accurate predictions. However, interpreting information in logs poses some challenges. This study discusses how to effectively mining sequences of logs and provide correct predictions. The approach integrates different machine learning techniques to control for data brittleness, provide accuracy of model selection and validation, and increase robustness of classification results. We apply the proposed approach to log sequences of 25 different applications of a software system for telemetry and performance of cars. On this system, we discuss the ability of three well-known support vector machines - multilayer perceptron, radial basis function and linear kernels - to fit and predict defective log sequences. Our results show that a good analysis strategy provides stable, accurate predictions. Such strategy must at least require high fitting ability of models used for prediction. We demonstrate that such models give excellent predictions both on individual applications - e.g., 1 % false positive rate, 94 % true positive rate, and 95 % precision - and across system applications - on average, 9 % false positive rate, 78 % true positive rate, and 95 % precision. We also show that these results are similarly achieved for different degree of sequence defectiveness. To show how good are our results, we compare them with recent studies in system log analysis. We finally provide some recommendations that we draw reflecting on our study.",https://link.springer.com/article/10.1007/s10664-014-9303-2,True,,17.0,"['multilayer-perceptron', 'support-vector-machine']",['logs'],"['novel-use', 'comparison']",['failure-prediction'],,,True,['failure-management'],,,['system-failure-prediction'],,,,,,
1065,Fast Dimensional Analysis for Root Cause Investigation in Large-Scale Service Environment,"['Lin, Fred', 'Muzumdar, Keyur', 'Laptev, Nikolay Pavlovich', 'Curelea, Mihai-Valentin', 'Lee, Seunghak', 'Sankar, Sriram']",2019,"Root cause analysis in a large-scale production environment is challenging due to the complexity of services running across global data centers. Due to the distributed nature of a large-scale system, the various hardware, software, and tooling logs are often maintained separately, making it difficult to review the logs jointly for detecting issues. Another challenge in reviewing the logs for identifying issues is the scale-there could easily be millions of entities, each with hundreds of features. In this paper we present a fast dimensional analysis …",https://arxiv.org/abs/1911.01225,True,,0.0,['rule-mining'],['logs'],['new-method'],['root-cause-analysis'],[],,True,['failure-management'],,,['root-cause-diagnosis'],,,,,,
1066,State-Driven Testing of Distributed Systems,,2013,,https://link.springer.com/chapter/10.1007/978-3-319-03850-6_9,True,,9.0,"['automaton', 'clustering']",,,['failure-prevention'],,,True,['failure-management'],,,"['fault-injection', 'anomaly-detection']",,,,,,
1067,On Fault Representativeness of Software Fault Injection,,2012,,https://ieeexplore.ieee.org/document/6122035,True,,173.0,"['decision-tree', 'clustering']",['source-code'],['new-method'],['failure-prevention'],['software'],,True,['failure-management'],True,True,"['fault-injection', 'software-defect-prediction']",True,16.0,,,,
1068,A survey of AI in operations management from 2005 to 2009,,2011,,https://www.emerald.com/insight/content/doi/10.1108/17410381111149602/full/html,True,,16.0,,,['survey'],,,,True,['aiops-general'],True,,,,,,,,
1069,A Survey of Artificial Intelligence for Prognostics,,2007,,https://www.aaai.org/Library/Symposia/Fall/2007/fs07-02-016.php,True,,245.0,,,"['survey', 'discussion']",,,,True,['aiops-general'],True,,,,,,,,
1070,AI and OR in management of operations: history and trends,,2005,,https://orsociety.tandfonline.com/doi/abs/10.1057/palgrave.jors.2602132,True,,55.0,,,['survey'],,,,True,['aiops-general'],True,,,,,,,,
1071,A survey of techniques for internet traffic classification using machine learning,,2008,,https://ieeexplore.ieee.org/document/4738466,True,,1430.0,,,['survey'],['failure-detection'],['network'],,True,['failure-management'],True,,['traffic-classification'],,,,,,
1072,Transfer defect learning,"['J.Nam', 'S.J.Pam', 'S.Kim']",2013,,https://ieeexplore.ieee.org/document/6606584,True,,317.0,"['logistic-regression', 'dimensionality-reduction']","['code-metrics', 'code-history']",['new-method'],['failure-prevention'],['source-code'],,True,['failure-management'],True,,['software-defect-prediction'],True,12.0,,,,
1073,GP-based software quality prediction,"['Evett, Matthew', 'Khoshgoftar, Taghi', 'Chien, Pei-der', 'Allen, Edward', 'others']",1998,"Software development managers use software quality prediction methods to determine to which modules expensive reliability techniques should be applied. In this paper we describe a genetic programming (GP) based system for targetting software modules for reliability enhancement. The paper describes the GP system, and provides a case study using software quality data from two actual industrial projects. The system is shown to be robust enough for use in industrial domains.",https://emunix.emich.edu/~evett/Publications/gp98-se.pdf,True,,90.0,['genetic-programming'],['code-metrics'],['novel-use'],['failure-prevention'],['software'],"Proceedings of the Third Annual Conference Genetic Programming, volume",True,['failure-management'],True,,['software-defect-prediction'],False,,,,,
1074,Early prediction of software component reliability,,2008,,https://ieeexplore.ieee.org/document/4814122,True,,187.0,['markov-model'],,['novel-use'],['failure-prevention'],['software'],,True,['failure-management'],True,,['software-defect-prediction'],False,,,,,
1075,Method-level bug prediction,,2012,,https://ieeexplore.ieee.org/document/6475415,True,,90.0,"['random-forest', 'bayesian-network', 'support-vector-machine', 'decision-tree']","['code-metrics', 'code-history']",['novel-use'],['failure-prevention'],['source-code'],,True,['failure-management'],True,,['software-defect-prediction'],True,9.0,,,,
1076,Sufficient mutation operators for measuring test effectiveness,,2008,,https://ieeexplore.ieee.org/document/4814146,True,,189.0,['linear-regression'],"['source-code', 'test-cases']",['novel-use'],['failure-prevention'],['software'],,True,['failure-management'],True,,['fault-injection'],True,15.0,,,,
1077,Rejuvenation and failure detection in partitionable systems,,2001,,https://ieeexplore.ieee.org/document/992692,True,,10.0,,,,['failure-prevention'],['modem'],,True,['failure-management'],,,['rejuvenation'],,,,,,
1078,Modeling and analysis of software rejuvenation in cable modem termination systems,,2002,,https://ieeexplore.ieee.org/document/1173239,True,,53.0,,,,['failure-prevention'],['modem'],,True,['failure-management'],,,['rejuvenation'],,,,,,
1079,"Design, Modeling, and Evaluation of a Scalable Multi-level Checkpointing System",,2010,,https://ieeexplore.ieee.org/document/5645453,True,,575.0,['markov-model'],['configuration'],['new-method'],['failure-prevention'],"['cluster', 'node', 'software']",,True,['failure-management'],True,,['checkpointing'],True,23.0,,,,['hpc']
1080,Effective Cost Reduction for Elastic Clouds under Spot Instance Pricing Through Adaptive Checkpointing,"['Jangjaimon, Itthichok', 'Tzeng, Nian-Feng']",2013,"Cloud computing users are most concerned about the application turnaround time and the monetary cost involved. For lower monetary costs, less expensive services, like spot instances offered by Amazon, are often made available, albeit to their relatively frequent resource unavailability that leads to on-going execution being evicted, thereby undercutting execution performance. Meanwhile, multithreaded applications may take advantage of elastic resource availability and cost fluctuation inherent to the systems. However, their …",https://ieeexplore.ieee.org/document/6678343,True,,33.0,['markov-model'],,['new-method'],['failure-prevention'],"['cluster', 'node', 'software']",,True,['failure-management'],True,,['checkpointing'],True,24.0,,,,
1081,Performance debugging for distributed systems of black boxes,,2003,,https://dl.acm.org/doi/10.1145/945445.945454,True,,800.0,['search'],['traces'],,['root-cause-analysis'],,SOSP '03: Proceedings of the nineteenth ACM symposium on Operating systems principles,True,['failure-management'],True,True,['rca-others'],True,82.0,,,,
1082,iHelp: An Intelligent Online Helpdesk System,,2010,,https://ieeexplore.ieee.org/document/5475278,True,,44.0,,,,['remediation'],,"IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics)",True,['failure-management'],,,['solution-recommendation'],,,,,,
1083,EasyTicket: a ticket routing recommendation engine for enterprise problem resolution,"['Shao, Qihong', 'Chen, Yi', 'Tao, Shu', 'Yan, Xifeng', 'Anerousis, Nikos']",2008,"Managing problem tickets is a key issue in IT service industry. A large service provider may handle thousands of problem tickets from its customers on a daily basis. The efficiency of processing these tickets highly depends on ticket routing---transferring problem tickets among expert groups in search of the right resolver to the ticket. Despite that many ticket management systems are available, ticket routing in these systems is still manually operated by support personnel. In this demo, we introduce EasyTicket, a ticket routing …",https://dl.acm.org/doi/abs/10.14778/1454159.1454193,True,,34.0,['markov-model'],['tickets'],,['remediation'],,,True,['failure-management'],,,['solution-recommendation'],,,,,,
1084,Dependency Analysis in Distributed Systems using Fault Injection: Application to Problem Determination in an e-commerce Environment,,2001,,http://proceedings.utwente.nl/14/,True,,96.0,['graph-mining'],,,"['root-cause-analysis', 'failure-prevention']",,,True,['failure-management'],,False,"['fault-injection', 'root-cause-diagnosis']",,,,,,
1085,An ant colony optimization algorithm to improve software quality prediction models: Case of class stability,"['Azar, Danielle', 'Vybihal, Joseph']",2010,"Context Assessing software quality at the early stages of the design and development process is very difficult since most of the software quality characteristics are not directly measurable. Nonetheless, they can be derived from other measurable attributes. For this purpose, software quality prediction models have been extensively used. However, building accurate prediction models is hard due to the lack of data in the domain of software engineering. As a result, the prediction models built on one data set show a significant …",https://laur.lau.edu.lb:8443/xmlui/bitstream/handle/10725/3407/An%20ant.pdf?sequence=1,True,,52.0,['ant-colony'],,['novel-use'],['failure-prevention'],['software'],,True,['failure-management'],True,,['software-defect-prediction'],False,,,,,
1086,Hora: Architecture-aware online failure prediction,,2017,,https://www.sciencedirect.com/science/article/pii/S0164121217300390?via%3Dihub,True,,23.0,"['autoregression', 'bayesian-network']","['configuration', 'host-metrics']",['new-method'],['failure-prediction'],,,True,['failure-management'],True,,['system-failure-prediction'],True,46.0,,,,
1087,Deep Learning for Just-in-Time Defect Prediction,,2015,,https://ieeexplore.ieee.org/document/7272910,True,,135.0,,,,['failure-prevention'],['source-code'],,False,['failure-management'],,,['software-defect-prediction'],,,,,,