Published August 21, 2020 | Version v1

DETERMINATION AND CORRECTION OF READING ERRORS IN THE SEQUENCES OBTAINED BY THE NEW GENERATION SEQUENCE METHOD

  • 1. Karadeniz Technical University

Description

Bioinformatics; It is a branch of science, which is the synthesis of mathematics, statistics, computer science, molecular biology and genetics in order to make sense of biological data, store it, visualize it and make maximum use of this huge knowledge. Bioinformatics studies; genetic disease research, disease detection and DNA sequence methods in order to produce solutions to detect diseases are focused on. Accordingly, the main purpose of bioinformatics is to try to understand the mechanism that causes diseases starting from the nucleotide sequence in the DNA where our genetic code is written and contribute to the development of treatment methods accordingly. Today, especially with the “Human Genome Project”, it has become a vital issue to analyze genetic data with faster and more reliable methods. The analysis of genetic data includes sub-studies such as reducing their size, choosing a subset from their properties and classifying the data, clustering, and estimating their new status. One of the purposes of analyzing biological data by computer is to make a preliminary analysis with the computer and analyze the predicted variables (bio-pointers) in the laboratory environment before analyzing these very high-dimensional / variable data with classical laboratory research. The main problems in the analysis of genetic data are the complexities of data sequences based on their size. This complexity brings about the error of reading data encountered during the processing of the data. Since large data sequences cannot be read at once, it is necessary to process the data in pieces (Sequence). With the help of Next Generation Sequencing (YND) devices, it is possible to read large genetic data, but the cost of these operations is high. In addition, YND devices perform erroneous readings between 1% and 3% during reading of genetic data sequences. In this study, a method is proposed to detect and correct common data reading errors for the detection of one of the most common cancer diseases, V-raf Murine Sarcoma Viral Oncogen Homologous B1 (BRAF) gene mutation. In this method, the healthy BRAF gene shared over the National Center for Biotechnology Information (NCBI) was used. Reading errors in proportion to the error ratios that occur by simulated YND devices have been added to this gene. The faulty gene was read at the specified depth size and recorded in a filter environment. Healthy and faulty genes were compared as a result of piecewise reading and correction procedures were applied with the help of the algorithm developed on detected errors. The proposed data reading and correction method is aimed to assist in the detection and resolution of DNA mutations that cause genetic disorders to contribute to the studies in the field of bioinformatics.

Files

YENİ NESİL DİZİLEME YÖNTEMİYLE ELDE EDİLEN SEKANSLARDA OKUMA .pdf

Files (970.7 kB)