Sequencing Reveals Putative Noncanonical SARS-CoV-2 Genomes with Predicted Programmed Ribosomal frameshifting: Datasets and supplementary materials
Authors/Creators
-
Cadena-Caballero, Cristian E.
(Researcher)1
-
Navarro-Corredor, Maria A.
(Researcher)1
-
Vera-Cala, Lina M.
(Supervisor)2
-
Barrios H., Carlos J.
(Supervisor)3, 4, 5
-
Torres-Jiménez, Carolina S.
(Researcher)1
-
Pardo-Diaz, Luis A.
(Researcher)1
-
Salome, Avila Nieves
(Researcher)1
-
Acuña-Carvajal, Cristina I.
(Researcher)1
-
Jimenez-Gutierrez, Laura Rebeca
(Researcher)6, 7
-
Martinez-Perez, Francisco
(Project leader)1, 3, 4
- 1. Laboratorio de Genómica Celular Aplicada del Grupo de Investigación en Computo Avanzado y a Gran Escala - CAGE, Universidad Industrial de Santander, Cl. 9 # Cra 27, 680006, Bucaramanga, Santander, Colombia
- 2. Departamento de Salud Pública, Escuela de Medicina, Universidad Industrial de Santander
- 3. Grupo de Investigación en Computo Avanzado y a Gran Escala - CAGE, Universidad Industrial de Santander, Cl. 9 # Cra 27, 680006, Bucaramanga, Santander, Colombia
- 4. Centro de Supercomputación y Cálculo Científico de la Universidad Industrial de Santander- SC3UIS, Universidad Industrial de Santander, Cl. 9 # Cra 27, 680006, Bucaramanga, Santander, Colombia
- 5. Institut National de recherche en Informatique et Automatique, Laboratoire d'informatique de Grenoble, Institut National des sciences appliquées de Lyon, Centre of Innovation in Telecommunications and Integration of Service
- 6. Facultad de Ciencias del Mar, Universidad Autónoma de Sinaloa, Paseo Claussen s/n, Mazatlán 82000, Culiacan Mexico
- 7. Investigadores por México. Secretaria de Ciencias, Humanidades y Tecnología, 03940, CDMX, México
Contributors
Project leader:
Researcher (7):
Supervisor (2):
- 1. Laboratorio de Genómica Celular Aplicada del Grupo de Investigación en Computo Avanzado y a Gran Escala - CAGE, Universidad Industrial de Santander, Cl. 9 # Cra 27, 680006, Bucaramanga, Santander, Colombia
- 2. Departamento de Salud Pública, Escuela de Medicina, Universidad Industrial de Santander
- 3. Grupo de Investigación en Computo Avanzado y a Gran Escala - CAGE, Universidad Industrial de Santander, Cl. 9 # Cra 27, 680006, Bucaramanga, Santander, Colombia
- 4. Centro de Supercomputación y Cálculo Científico de la Universidad Industrial de Santander- SC3UIS, Universidad Industrial de Santander, Cl. 9 # Cra 27, 680006, Bucaramanga, Santander, Colombia
- 5. Institut National de recherche en Informatique et Automatique, Laboratoire d'informatique de Grenoble, Institut National des sciences appliquées de Lyon, Centre of Innovation in Telecommunications and Integration of Service
- 6. Facultad de Ciencias del Mar, Universidad Autónoma de Sinaloa, Paseo Claussen s/n, Mazatlán 82000, Culiacan Mexico
- 7. Investigadores por México. Secretaria de Ciencias, Humanidades y Tecnología, 03940, CDMX, México
Description
Repository 1. Design Specificity, Structural Mapping, and Synthesis Validation of the 3' Primer RV30AkCOVID19.
This repository contains BLAST alignments against the GenBank database for the two sequence components of the RV30AkCOVID19 primer:
(1) The GC-rich polylinker region.
(2) The consensus nucleotide sequence complementary to the 3′ end of the SARS-CoV-2 genome.
It also includes a restriction enzyme recognition map of the RV30AkCOVID19 primer and MALDI-TOF quality control data confirming the integrity of oligonucleotide synthesis. The red box indicates the GC-rich polylinker sequence, whereas the green box highlights the region complementary to the SARS-CoV-2 3′ untranslated region (3′UTR).
Repository 2. Sequencing reads, quality assessment, and genome assemblies of SARS-CoV-2 isolates.
This repository contains sequencing outputs, quality control analyses, and genome assemblies generated from SARS-CoV-2–positive samples. The dataset is organised into four folders and one supplementary Excel file: “Percentage AT and GC SARS-CoV-2 genomes.xlsx” — which provides nucleotide composition data used for genome compositional analysis.
Folder structure
(1) Reads – Ion Torrent. Contains raw sequencing reads generated using Ion Torrent technology.
(2) FastQC. Contains quality control reports produced with FastQC. Reports are provided as compressed (.zip) files.
(3) IRMA. Includes genome assemblies generated using the IRMA pipeline. FASTA files contain the reference genome accession identifier followed by the corresponding sequenced SARS-CoV-2 genome.
(4) Bowtie2. Contains subfolders named according to each sequenced genome and the corresponding Bowtie2 analyses. The consensus subfolder includes consensus genome sequences and comparative alignments in FASTA format. This folder contains alignment results organised by sample, including read alignments (BAM/SAM), sorted alignments with index files, and consensus genome sequences in FASTA format.
For samples 3, 7, 10, 36, and 44, subfolders include results generated under multiple cDNA synthesis and reaction conditions, enabling comparison between denaturation treatment, dNTP optimization, commercial protocols, and direct RNA processing.
The Consensus folder contains final consensus genome sequences and reference-derived variants used for comparative and phylogenetic analyses.
The file reference_sequence.fasta.fai is an index file generated during alignment processing to enable rapid access to the reference genome sequence.
(5) The file Table 1 Supplementary, provides the results, which supports Figure 2. It includes sample metadata, sequencing metrics, genomic cDNA concentrations, coverage percentages, Pango lineage assignments (via IRMA and Bowtie 2), GISAID accession numbers, and the classification of canonical and putative nc-sgRNA SARS-CoV-2 genomes.
Repository 3. BLAST alignment of assembled SARS-CoV-2 genomes.
This repository contains BLAST alignment results for assembled SARS-CoV-2 genomes generated using two assembly approaches. Two folders are provided: BLAST – Bowtie2 and BLAST – IRMA, each containing plain text files with BLAST alignment outputs corresponding to genomes assembled using the respective pipeline. The files include results for both standard sequencing workflows and experimental conditions, enabling verification of genome identity and consistency across assembly strategies.
Repository 4. Pango lineage classification and Nextclade analysis of SARS-CoV-2 genomes.
This repository contains lineage assignments and clade analyses for SARS-CoV-2 genomes generated using Pangolin v1.16 and Nextclade v2.9.1.
The dataset is organised into two directories corresponding to genome assemblies produced with the Bowtie2 and IRMA pipelines. Within each directory:
- An Excel file provides PANGO lineage classifications for genomes assembled using the corresponding pipeline.
- A compressed file (nextclade.zip) contains Nextclade analysis outputs, including clade assignments, mutation profiles, quality metrics, and translated protein sequences.
- A CSV file contains Pangolin lineage outputs.
The same file structure and documentation are provided for both assembly approaches to enable direct comparison of lineage and clade assignments.
Repository 5. Reference genome alignments and assembled SARS-CoV-2 genomes.
This repository contains reference-based alignments and genome sequences used for comparative and structural analyses. The dataset is organised into three folders:
(1) IRMA genomes: Contains ClustalW alignments between genomes assembled using the IRMA pipeline and the SARS-CoV-2 reference genome, together with the corresponding genome sequences in FASTA format used to generate these alignments.
(2) Bowtie2 genomes: Contains ClustalW alignments between genomes assembled using the Bowtie2 pipeline and the SARS-CoV-2 reference genome, along with the corresponding FASTA sequences used for alignment.
(3) Genomes 07dN120320 and 27sT122620: Contains reference-based alignments of the putative non-canonical subgenomic RNAs (nc-sgRNAs) 07dN120320 and 27sT122620 relative to the SARS-CoV-2 reference genome. File names indicate the corresponding genome analysed.
Repository 6. RNA secondary and tertiary structures associated with predicted programmed ribosomal frameshifting.
This repository contains predicted RNA secondary and tertiary structures associated with predicted ribosomal frameshifting regions identified in non-canonical SARS-CoV-2 sgRNAs. The dataset is organised into three folders:
(1) Gibbs free energy 2D modelling: Contains secondary structure predictions for each frameshifting region in Vienna RNA format (dot-bracket notation), including minimum free energy configurations.
(2) Data modelling 3D: Contains structural data required to generate three-dimensional models of the corresponding RNA secondary structures.
(3) 3D structure images: Contains visual representations of the tertiary structures for each predicted frameshifting region. File names indicate the putative non-canonical genome identifier, gene, and nucleotide position associated with each structure. The file version2_5120-5121_structure_co_007dN.png illustrates nucleotide interactions using an arc diagram representation, whereas version2_7556-7582_structure_co_007dN.png corresponds to the secondary structure of the indicated region in Vienna RNA format (dot-bracket notation).
Repository 7. Curatorial workflow, consensus variant construction, and global comparative alignment of SARS-CoV-2 genomes.
This repository contains the curated SARS-CoV-2 genome dataset and comparative analyses used to evaluate deletion patterns within a global evolutionary framework. The workflow includes sequence depuration, consensus construction, and comparative alignments designed to identify genomic regions with elevated variability across circulating variants.
Folder structure
(1) SARS-CoV-2 complete genomes: Contains individual SARS-CoV-2 genome sequences downloaded from the GISAID database and organised by variant. Each FASTA file includes complete genome sequences and preserves the original GISAID identifiers and metadata (geographical origin and collection date). These sequences represent the primary dataset used for subsequent quality filtering and comparative analyses.
(2) Depuration of sequences_GISAID: This folder documents the quality filtering workflow applied to the initial SARS-CoV-2 genome collection.
Subfolders
- SARS-CoV-2 complete genome. Contains genome sequences retained after quality filtering. These sequences passed the removal criteria for consecutive undetermined nucleotides and did not exhibit atypical divergence relative to the main population.
- SARS-CoV-2 eliminates genome. Contains genome sequences excluded during the filtering process due to the presence of consecutive undetermined nucleotides or anomalous behaviour relative to the main dataset.
(3) SARS-CoV-2 consensus variants: Contains FASTA files representing the consensus genome sequence for each SARS-CoV-2 variant derived from the curated genome dataset. These consensus sequences summarize the predominant nucleotide composition within each variant lineage and were used for comparative genomic analyses and alignment-based evaluation of variable regions.
The numeric designation included in each consensus FASTA file indicates the threshold criterion used for consensus construction:
- 20 - majority nucleotide at each position (20% frequency threshold).
- 100 - positions with multiple dominant nucleotides represented using IUPAC nucleotide codes.
(4) SARS-CoV-2 alignment consensus variants: This folder contains multiple sequence alignments generated from SARS-CoV-2 consensus variant genomes to support comparative genomic analysis across variants.
Subfolders
- Alignment 20: Contains alignments generated using the 20% consensus threshold, representing the majority nucleotide at each genomic position.
- Alignment 100: Contains alignments generated using the 100% consensus criterion, where positions with multiple dominant nucleotides are represented using IUPAC ambiguity codes.
File contents: Each subfolder includes:
- Multiple sequence alignments in CLUSTAL format.
- Corresponding sequences in FASTA format.
- Versions generated with and without undetermined nucleotides (Ns).
(5) SARS-CoV-2 codons alignment consensus variants and putative nc-sgRNA
This folder contains codon-level alignments and graphical genome representations used to examine deletion-associated open reading frame (ORF) remodelling and putative nc-sgRNA structure.
- SARS-CoV-2 codons nc-sgRNA: Contains alignment documents generated from genomes obtained by reverse transcription using the following procedures: (1) Standard protocol, (2) Modified dNTP condition, (3) Denaturant condition.
Files include: codon-level alignments highlighting nucleotide deletions, short alignment versions focusing on key genomic regions, and comparative alignments used to evaluate reading frame preservation and ORF remodelling. These alignments were generated from genomes derived from the same clinical sample, enabling evaluation of intra-host genomic variability. Accordingly, the name of each Word document corresponds to the clinical sample used to determine the SARS-CoV-2 genomes.
Alignment ORFs.pdf: Provides an open reading frame (ORF)-focused alignment reference summarizing reading frame effects associated with deletion events that give rise to SARS-CoV-2 putative non-canonical subgenomic RNAs identified in this study.
- SARS-CoV-2 Geneious Prime: Contains graphical genome representations generated using Geneious Prime.
Files include: individual genome structural maps, multi-scale visualizations (25%, 50%, and 100% identity thresholds), graphical representations of ORF organization and deletion-associated structural variation. These visualizations support structural interpretation of genomic rearrangements and nc-sgRNA-associated variation.
(6) Variant Alignment – Ns
Contains ClustalW alignments of SARS-CoV-2 variant genomes that include undetermined nucleotides (Ns), enabling evaluation of alignment behaviour in regions with ambiguous base calls. The folder also includes reference FASTA sequences used to generate alignments, allowing verification of positional consistency and comparison and assessment of sequence regions containing unresolved nucleotides.
Repository 8. Maximum likelihood phylogenetic analyses of SARS-CoV-2 genomes.
This repository contains the results of maximum likelihood phylogenetic analyses of SARS-CoV-2 genomes. Two folders are provided:
- Phylogeny with Ns: Results obtained from analyses performed using genome sequences that include undetermined nucleotides (Ns).
- Phylogeny without Ns: Results obtained after removal of undetermined nucleotide positions to evaluate phylogenetic robustness.
Each folder contains consensus trees, final tree files, and associated analysis logs generated during phylogenetic inference.
Notes (English)
Notes (English)
Files
Data Zenodo genomic-UIS-SARS-CoV-2 (2026-08-26).zip
Additional details
Funding
- Ministerio de Ciencia, Tecnología e Innovación
- Ministry of Science, Technology, and Innovation of Colombia (Minciencias) [contract No. 369-2020, code 1102101576900]
- Industrial University of Santander
- Vice-Rectory for Research and Extension of the Universidad Industrial de Santander project No. 76900
Software
- Repository URL
- https://github.com/GenomicUIS
- Programming language
- Python , Shell , R
- Development Status
- Active