Published May 22, 2017 | Version v1
Dataset Open

Experimental results of "Managing variant calling datasets the big data way"

  • 1. Wageningen University and Research

Description

Tomatula was demonstrated for retrieving the allele frequencies for a given region in the data from Aflitos et al (2014). We developed scripts to retrieve allele frequencies, either from the VCF file storage or Apache Parquet. We executed a series of experiments, querying for a region of 2000 bases in the file of chromosome 6, that corresponds to the approximate length of a gene. We compared both storage formats (VCF files and Parquet), two input sizes (104 and 1144 individuals), different cluster sizes varying between 2 and 150 executor nodes, and HDFS replication factor was set to 3, 5, 7, and 9, in order to examine four main factors that
can affect the performance of a Big Data cluster: (a) the storage format, (b) the size of the input files,  (c) the number of computing nodes of the cluster, and (d) the replication factor of HDFS. The block size of the HDFS was kept at the default value of 128MB. All experiments were executed five times and the detailed results are provided here, along with a script that produces the corresponding figures.

Notes

This work was carried out on the Dutch national e-infrastructure with the support of SURF Cooperative.

Files

data.csv

Files (584.7 kB)

Name Size Download all
md5:c192c4206108a0eb2087fba386955429
2.7 kB Download
md5:56cf9093c311799a2cbca7f4a8a324b1
8.9 kB Preview Download
md5:5142efb949e3df601eb21ae15742b395
134.7 kB Preview Download
md5:1b0caffe517481c3cb6a56afbe301592
219.4 kB Preview Download
md5:e25f8cd31c495b8e55d6e1e9cb56332a
219.0 kB Preview Download