Published April 1, 2020 | Version v1
Journal article Open

Genomic repeats detection using Boyer-Moore algorithm on Apache Spark Streaming

  • 1. Universitas Pendidikan Indonesia
  • 2. Djillali Liabes University

Description

Genomic repeats, i.e., pattern searching in the string processing process to find repeated base pairs in the order of Deoxyribonucleic Acid (DNA), requires a long processing time. This research builds a big-data computational model to look for patterns in strings by modifying and implementing the Boyer-Moore algorithm on Apache Spark Streaming for human DNA sequences from the Ensemble site. Moreover, we perform some experiments on cloud computing by varying different specifications of computer clusters with involving datasets of human DNA sequences. The results obtained show that the proposed computational model on Apache Spark Streaming is f aster than standalone computing and parallel computing with multicore. Therefore, it can be stated that the main contribution in this research, which is to develop a computational model for reducing the computational costs, has been achieved.

Files

24 14883.pdf

Files (602.7 kB)

Name Size Download all
md5:193105db72f6780d521f7d58d4554db8
602.7 kB Preview Download