There is a newer version of the record available.

Published October 18, 2019 | Version v9
Dataset Restricted

Genomic sequences and annotations for Solanum lycopersicum, Solanum pennellii and Solanum habrochaites

  • 1. University of Amsterdam

Description

 

=== Genome sequences ===

These are the different genome references (fasta formats) available for:

The two genome assemblies of S. habrochaites LA1777 and PI127826 were obtained through a combination of 10X Linked-Reads and BioNano Optical Mapping. This sequencing has been funded by the DTL Technology Hotel 2018 funding scheme.

 

=== Transcriptomes and proteomes ===

Solanum lycopersicum (assembly 4.0):

  • Transcriptome: ITAG4.0_cDNA.fasta 
  • Proteome: ITAG4.0_proteins.fasta

Solanum pennellii (one version only from Bolger et al., 2014):

Solanum lycopersicoides (version 1.0)

 

=== Genome annotations files ===

Solanum lycopersicum

Solanum lycopersicoides:

Solanum habrochaites

  • PI127826: a GFF file was produced using RepeatMasker and funannotate (see below). The file is named Solanum_habrochaites_PI127826.
    RepeatMasker -qq -e rmblast -small -xsmall -pa 10 -lib mipsREdat_9.3p_Eudicot_TEs.fasta -dir repeat_masking_run/ PI127826.fasta
    
    funannotate sort -i PI127826.fasta -b scaffold -o PI127826.sorted.fasta 
    
    funannotate mask -i PI127826.sorted.fasta \
                     -o PI127826.sorted.masked.fasta \
                     -m tantan \
                     -s tomato \
                     --cpus 12
    
    funannotate train -i PI127826.sorted.masked.fasta -o 01_train_step/ \
                      --left fastq/PI127826_01_R1.fastq.gz fastq/PI127826_06_R1.fastq.gz \
                      --right fastq/PI127826_01_R2.fastq.gz fastq/PI127826_06_R2.fastq.gz \
                      --single fastq/PI127826_02_R1.fastq.gz fastq/PI127826_03_R1.fastq.gz fastq/PI127826_04_R1.fastq.gz fastq/PI127826_05_R1.fastq.gz \
                     --species "Solanum lycopersicum" \
                     --cpus 16
                     --max_intronlen 3000
    
    funannotate predict -i 00_mask_step/PI127826.sorted.masked.fasta \
                        --out 01_train_step/ \
                        -s "Solanum lycopersicum" \
                        --cpus 16 \
                        --organism other \
                        --min_intronlen 10 \
                        --max_intronlen 10000 \
                        --repeats2evm 
    
    funannotate update -i 01_train_step/ \
                         --fasta 00_mask_step/PI127826.sorted.masked.fasta \
                          --cpus 16 \
                          --species "Solanum lycopersicum" \
                          --max_intronlen 10000 
    
    # [09:07 AM]: Previous annotation consists of: 75,886 protein coding gene models and 1,617 non-coding gene models.
    
    
    # After update:
    # 45,235 contigs containing 78,629 protein coding genes and 1,583 tRNA genes
    # ~ 45000 contigs = super-scaffolds + contigs
    
    
    funannotate fix -i 01_train_step/update_results/Solanum_lycopersicum.gbk -t 01_train_step/update_results/Solanum_lycopersicum.tbl
    
    
    # used Docker image https://github.com/blaxterlab/interproscan-docker
    # singularity pull docker://blaxterlab/interproscan-docker
    # singularity run interproscan.simg (create and start a container + enters inside)
    # then typed:
    
    
    # run eggnog to annotate proteins
    ./emapper.py --cpu 20 \
                 -i ../01_train_step/update_results/Solanum_lycopersicum.proteins.fa \
                 -m diamond \
                 -o Solanum_habrochaites_PI127826_eggnog  
    
    # add functional preduction from eggnog
    funannotate annotate -i 01_train_step/ \
    --gff Solanum_habrochaites_PI127826.gff3 \
    --out 03_annotate \
    --species "Solanum habrochaites" \
    --eggnog eggnog-mapper/Solanum_habrochaites_PI127826_eggnog.emapper.annotations \
    --busco_db embryophyta \
    --cpus 10
    
    
    # Annotation consists of: 77,297 gene models
    
    

     

  • LA1777

 

Reference:

Tomato Genome Sequencing Consortium. 2012. The tomato genome sequence provides insights into fleshy fruit evolution. Nature volume 485, pages 635–641.

Bolger et al. 2014. The genome of the stress-tolerant wild tomato species Solanum pennellii http://www.nature.com/ng/journal/v46/n9/full/ng.3046.html 

Hosmani et al. 2019. An improved de novo assembly and annotation of the tomato reference genome using single-molecule sequencing, Hi-C proximity ligation and optical maps. https://www.biorxiv.org/content/10.1101/767764v1

Aflitos et al. 2014. Exploring genetic variation in the tomato (Solanum section Lycopersicon) clade by whole‐genome sequencing. https://onlinelibrary.wiley.com/doi/full/10.1111/tpj.12616

Stam et al. 2019. The de Novo Reference Genome and Transcriptome Assemblies of the Wild Tomato Species Solanum chilense Highlights Birth and Death of NLR Genes Between Tomato Species. G3: Genes, Genomes, Genetics December 1, 2019 vol. 9 no. 12 3933-3941; https://doi.org/10.1534/g3.119.400529

 

 

Files

Restricted

The record is publicly accessible, but files are restricted. Log in to check if you have access.

Request access

If you would like to request access to these files, please fill out the form below.

You need to satisfy these conditions in order for this request to be accepted:

To get access to these files, please contact Petra Bleeker (P.M.Bleeker@uva.nl) or Marc Galland (m.galland@uva.nl).

You are currently not logged in. Do you have an account? Log in here

Additional details

Funding

Dutch Research Council
Defence in the wild; from trichome transcriptomes and metabolomes to breeding tools for defence markers in tomato 2300178970