There is a newer version of the record available.

Published December 23, 2025 | Version 1.5.1

PanTEon Database: A Cross-Kingdom, Automatically Curated Reference for Transposable Elements

  • 1. ROR icon Barcelona Supercomputing Center
  • 2. ROR icon Universitat Oberta de Catalunya
  • 3. The Rosalind Franklin Institute
  • 4. ROR icon Universidad Autonoma de Manizales
  • 5. ROR icon Rosalind Franklin Institute
  • 6. ROR icon Universidade Tecnológica Federal do Paraná
  • 7. Institut de recherche pour le développement France-Sud
  • 8. ROR icon Institute for Research in Biomedicine

Description

 

The PanTEon Database is a freely available collection of more than 126,000 automatically curated transposable element (TE) sequences, spanning animals, plants, and fungi and covering all major TE orders. The database was designed to maximize sequence fidelity, taxonomic diversity, and methodological consistency, making it suitable for both training and benchmarking state-of-the-art TE classification tools.

Data sources and integration

PanTEon integrates TE sequences from multiple complementary resources:

  • Curated sequences from Dfam (version 3.9)

  • Automatically curated sequences from APTEdb (Pedro et al., 2021)

  • Uncurated sequences from Dfam

  • TE sequences from the Ensembl 2023 release (Martin et al., 2023)

All uncurated sequences were originally generated using RepeatModeler2 (Flynn et al., 2020) and subsequently automatically curated with MCHelper (Orozco-Arias et al., 2024). Only TEs showing clear evidence of structural completeness and expected length profiles were retained, resulting in a high-confidence dataset suitable for training and benchmarking machine learning models.

Sequence identification

Each TE sequence in the PanTEon Database is assigned a standardized identifier composed of:

  1. A sequence name (either provided by Dfam or systematically assigned by the PanTEon framework),

  2. A three-level classification (Class / Order / Superfamily),

  3. The species of origin.

For example, a TE sequence derived from curated Dfam data is identified as:

>PumCon-1.141#CLASSII/TIR/TC1MARINER @Puma concolor

In this case, PumCon-1.141 is the original family name, the element belongs to the TIR/Tc1–Mariner superfamily, and it was obtained from Puma concolor.

In contrast, a TE sequence that was originally uncurated and subsequently processed and integrated into the PanTEon Database follows the systematic naming scheme:

>PDB00000038#CLASSI/LTR/LARD @Certhia brachydactyla

Metadata and taxonomic context

Additional taxonomic information—such as order, family, phylum, and higher ranks—as well as details about the origin of each sequence, are provided in the accompanying metadata file:

PanTEon_Database_metadata_v.1.5.1.csv

Benchmark edition

To enable fair and reproducible benchmarking of state-of-the-art TE classification tools, a benchmark edition of the PanTEon Database was generated. This version includes only TE superfamilies represented by more than 10 sequences and merges rare or uncommon superfamilies with their closest relatives (see the PanTEon paper for full details).

The benchmark dataset is available as:

PanTEon_Database_v1.5.1_benchmark_edition.fasta

Trained models for PanTEon Inference

This repository also contains pre-trained models corresponding to the different in-built architectures of PanTEon (classification task). These models can be downloaded and used with the PanTEon inference module by specifying their paths via the -d parameter. This release includes models trained on the following datasets: all (entire PanTEon Database v1.5.1), Animalia, Chordata, Arthropods, Plantae, Angiosperms, and Fungi. PanTEon can be downloaded from the following GitHub repository: https://github.com/simonorozcoarias/PanTEon

Why PanTEon?

By combining broad taxonomic coverage, automated curation, and a standardized nomenclature, the PanTEon Database provides a robust reference resource for:

  • developing and training machine learning and deep learning models,

  • benchmarking TE classification tools,

  • and exploring TE diversity across kingdoms.

Funding

Simon Orozco-Arias is supported by a fellowship within the “Generación D” initiative, Red.es, Ministerio para la Transformación Digital y de la Función Pública, for talent attraction (C005/24-ED CV1). Funded by the European Union NextGenerationEU funds, through PRTR.

 

Files

PanTEon_Database_metadata_v1.5.1.csv

Files (13.8 GB)

Name Size
md5:357f2bca34286806f80b88620691f516
110.3 kB Preview Download
md5:3e50976edf3fb2bf7ac5c4bd7153cafc
410.7 MB Download
md5:23c9326d23d9b61e46b313a56faec575
401.4 MB Download
md5:83692f1c8984835535d49208ffe1b8f9
1.8 GB Preview Download
md5:10507ec668b6c91ec3334d43a82d32dc
1.7 GB Preview Download
md5:ab0f1c0e50f8e3188af943ba05bccbf7
1.9 GB Preview Download
md5:b34c85ae7f5916c2cf57003250d34250
1.8 GB Preview Download
md5:cfdf05a454f35b536508a9d0df113053
2.0 GB Preview Download
md5:e8ed1de47caf64e6b581ed36e9feea93
1.9 GB Preview Download
md5:b8da567e23d374cc5845749f84824d02
1.8 GB Preview Download

Additional details

Funding

Barcelona Supercomputing Center
AI4Science Fellowship