Published September 17, 2025 | Version v1

Deep-learning-based annotation of 346 genomes in the Solanaceae family yields a harmonized dataset of 15 million proteins

  • 1. ROR icon Sainsbury Laboratory

Description

Accurate and consistent genome annotation is essential for reliable phylogenomic analyses. However, in the Solanaceae, publicly available genome annotations often differ in quality and structure due to inconsistencies in annotation pipelines and inclusion of alternative transcripts. These discrepancies hinder reliable evolutionary analyses. Here, we reannotated 346 publicly available genomes—including haplotype-phased assemblies of the same samples—using Helixer, a deep-learning-based gene prediction tool (Holst et al., 2023, https://doi.org/pzhq). Helixer predicted 15,079,126 protein-coding genes across 85 species spanning eight solanaceous genera. We provide a comprehensive annotation resource including predicted proteomes, coding sequences (CDS), GFF files, and domain annotations derived from InterProScan. Furthermore, we identified 197,834 nucleotide-binding and leucine-rich repeat receptors (NLRs)—a major class of intracellular immune receptors in plants—using NLRtracker (Kourelis et al., 2021, https://doi.org/g8vnsx), comprising approximately 1.3% of the predicted proteome. This harmonized dataset offers a valuable foundation for large-scale phylogenomic studies and gene family analyses in Solanaceae.

Files

25-09-17_SolDB_v2.pdf

Files (1.8 MB)

Name Size Download all
md5:f9f5466a637ef7eba314acf5a6e911b9
1.8 MB Preview Download

Additional details

References

  • Deanna et al., 2025, DOI: 10.1101/2025.07.10.663745
  • Holst et al., 2023, DOI: 10.1101/2023.02.06.527280
  • Pertea et al., 2020, PMID: 32489650
  • Jones et al., 2014, PMID: 24451626
  • Kourelis et al., 2021, PMID: 34669691