PanTEon Database: A Cross-Kingdom, Automatically Curated Reference for Transposable Elements
Authors/Creators
-
Orozco Arias, Simon
(Contact person)1
- Ferrer-Pomer, Iamil (Data collector)2
-
Rodrigues de Góes, Fabiana
(Data curator)3
-
Gaviria-Orrego, Simon
(Data curator)4
-
Rossi Paschoal, Alexandre
(Data manager)5, 6
-
guyot, romain
(Data manager)7
-
Gabaldón, Toni
(Data manager)1, 8
- Gómiz, Juan (Data collector)2
- Llatser Torres, Jordi (Data collector)2
-
1.
Barcelona Supercomputing Center
-
2.
Universitat Oberta de Catalunya
- 3. The Rosalind Franklin Institute
-
4.
Universidad Autonoma de Manizales
-
5.
Rosalind Franklin Institute
-
6.
Universidade Tecnológica Federal do Paraná
- 7. Institut de recherche pour le développement France-Sud
-
8.
Institute for Research in Biomedicine
Description
The PanTEon Database is a freely available collection of more than 126,000 automatically curated transposable element (TE) sequences, spanning animals, plants, and fungi and covering all major TE orders. The database was designed to maximize sequence fidelity, taxonomic diversity, and methodological consistency, making it suitable for both training and benchmarking state-of-the-art TE classification tools.
Data sources and integration
PanTEon integrates TE sequences from multiple complementary resources:
-
Curated sequences from Dfam (version 3.9)
-
Automatically curated sequences from APTEdb (Pedro et al., 2021)
-
Uncurated sequences from Dfam
-
TE sequences from the Ensembl 2023 release (Martin et al., 2023)
All uncurated sequences were originally generated using RepeatModeler2 (Flynn et al., 2020) and subsequently automatically curated with MCHelper (Orozco-Arias et al., 2024). Only TEs showing clear evidence of structural completeness and expected length profiles were retained, resulting in a high-confidence dataset suitable for training and benchmarking machine learning models.
Sequence identification
Each TE sequence in the PanTEon Database is assigned a standardized identifier composed of:
-
A sequence name (either provided by Dfam or systematically assigned by the PanTEon framework),
-
A three-level classification (Class / Order / Superfamily),
-
The species of origin.
For example, a TE sequence derived from curated Dfam data is identified as:
>PumCon-1.141#CLASSII/TIR/TC1MARINER @Puma concolor
In this case, PumCon-1.141 is the original family name, the element belongs to the TIR/Tc1–Mariner superfamily, and it was obtained from Puma concolor.
In contrast, a TE sequence that was originally uncurated and subsequently processed and integrated into the PanTEon Database follows the systematic naming scheme:
>PDB00000038#CLASSI/LTR/LARD @Certhia brachydactyla
Metadata and taxonomic context
Additional taxonomic information—such as order, family, phylum, and higher ranks—as well as details about the origin of each sequence, are provided in the accompanying metadata file:
PanTEon_Database_metadata_v.1.5.1.csv
Benchmark edition
To enable fair and reproducible benchmarking of state-of-the-art TE classification tools, a benchmark edition of the PanTEon Database was generated. This version includes only TE superfamilies represented by more than 10 sequences and merges rare or uncommon superfamilies with their closest relatives (see the PanTEon paper for full details).
The benchmark dataset is available as:
PanTEon_Database_v1.5.1_benchmark_edition.fasta
Trained models for PanTEon Inference
This repository also contains pre-trained models corresponding to the different in-built architectures of PanTEon (classification task). These models can be downloaded and used with the PanTEon inference module by specifying their paths via the -d parameter. This release includes models trained on the following datasets: all (entire PanTEon Database v1.5.1), Animalia, Chordata, Arthropods, Plantae, Angiosperms, and Fungi. PanTEon can be downloaded from the following GitHub repository: https://github.com/simonorozcoarias/PanTEon
Why PanTEon?
By combining broad taxonomic coverage, automated curation, and a standardized nomenclature, the PanTEon Database provides a robust reference resource for:
-
developing and training machine learning and deep learning models,
-
benchmarking TE classification tools,
-
and exploring TE diversity across kingdoms.
Funding
Simon Orozco-Arias is supported by a fellowship within the “Generación D” initiative, Red.es, Ministerio para la Transformación Digital y de la Función Pública, for talent attraction (C005/24-ED CV1). Funded by the European Union NextGenerationEU funds, through PRTR.
Files
PanTEon_Database_metadata_v1.5.1.csv
Files
(13.8 GB)
| Name | Size | |
|---|---|---|
|
md5:357f2bca34286806f80b88620691f516
|
110.3 kB | Preview Download |
|
md5:3e50976edf3fb2bf7ac5c4bd7153cafc
|
410.7 MB | Download |
|
md5:23c9326d23d9b61e46b313a56faec575
|
401.4 MB | Download |
|
md5:83692f1c8984835535d49208ffe1b8f9
|
1.8 GB | Preview Download |
|
md5:10507ec668b6c91ec3334d43a82d32dc
|
1.7 GB | Preview Download |
|
md5:ab0f1c0e50f8e3188af943ba05bccbf7
|
1.9 GB | Preview Download |
|
md5:b34c85ae7f5916c2cf57003250d34250
|
1.8 GB | Preview Download |
|
md5:cfdf05a454f35b536508a9d0df113053
|
2.0 GB | Preview Download |
|
md5:e8ed1de47caf64e6b581ed36e9feea93
|
1.9 GB | Preview Download |
|
md5:b8da567e23d374cc5845749f84824d02
|
1.8 GB | Preview Download |
Additional details
Funding
- Barcelona Supercomputing Center
- AI4Science Fellowship