Statistical Reports and Data Analytics with Distributed Computing
Authors/Creators
- 1. CERN openlab Summer Student
- 2. Summer Student Supervisor
Description
Project Specification
This project involves the following:
-
Data Analytics as a Service - Providing a way for users to write their own Data Analytics algorithms in R through the use of Jupyter. This will allow them to easily run and share their code, data and results.
-
Scale out data analytics algorithm using Docker - Parallelizing the algorithm across a cluster on Openstack by dividing the task over multiple nodes to execution time.
Once completed, this project will provide users with the necessary tools and infrastructure to be able to perform Data Analytics themselves on their data as well as view their results in a straightforward manner.
Abstract
The control systems needed to run the Large Hadron Collider (LHC), its injector accelerators and their infrastructure generate massive amounts of data. This data can be used to optimize the control systems, and provide meaningful information to the machines operators and experts. At present, algorithms perform Data Analytics on over 3000 signals each day, and this number is only increasing. Performing such a large amount of computations is time consuming, especially when run on a single machine. Therefore it is through this project that we aim to search for a means of parallelizing the execution of such algorithms. The proposed solution makes use of Docker which allows for straightforward scalability. Results show that scaling up the system does indeed decrease the execution time required.
Files
SummerStudentReport-GabriellaAzzopardi.pdf
Files
(1.4 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:f2e78a6ddc46ec4bdf5eacd7aa75bcb3
|
1.4 MB | Preview Download |