Published April 1, 2026 | Version 1.6.1

PanTEon Database: A Cross-Kingdom, Automatically Curated Reference for Transposable Elements

  • 1. ROR icon Barcelona Supercomputing Center
  • 2. ROR icon Universitat Oberta de Catalunya
  • 3. The Rosalind Franklin Institute
  • 4. ROR icon Universidad Autonoma de Manizales
  • 5. ROR icon Rosalind Franklin Institute
  • 6. ROR icon Universidade Tecnológica Federal do Paraná
  • 7. Institut de recherche pour le développement France-Sud
  • 8. ROR icon Institute for Research in Biomedicine

Description

 

The PanTEon Database is a freely available collection of almost 240,000 automatically curated transposable element (TE) sequences, spanning animals, plants, and fungi and covering all major TE orders. The database was designed to maximize sequence fidelity, taxonomic diversity, and methodological consistency, making it suitable for both training and benchmarking state-of-the-art TE classification tools.

Data sources and integration

PanTEon integrates TE sequences from multiple complementary resources:

  • Curated sequences from Dfam (version 3.9)

  • Automatically curated sequences from APTEdb (Pedro et al., 2021)

  • Uncurated sequences from Dfam

  • TE sequences from the Ensembl 2025 release (Dyer et al., 2025)

All uncurated sequences were originally generated using RepeatModeler2 (Flynn et al., 2020) and subsequently automatically curated with MCHelper (Orozco-Arias et al., 2024). Only TEs showing clear evidence of structural completeness and expected length profiles were retained, resulting in a high-confidence dataset suitable for training and benchmarking machine learning models.

Sequence identification

Each TE sequence in the PanTEon Database is assigned a standardized identifier composed of:

  1. A sequence name (either provided by Dfam or systematically assigned by the PanTEon framework),

  2. A three-level classification (Class / Order / Superfamily),

  3. The species of origin.

For example, a TE sequence derived from curated Dfam data is identified as:

>PumCon-1.141#CLASSII/TIR/TC1MARINER @Puma concolor

In this case, PumCon-1.141 is the original family name, the element belongs to the TIR/Tc1–Mariner superfamily, and it was obtained from Puma concolor.

In contrast, a TE sequence that was originally uncurated and subsequently processed and integrated into the PanTEon Database follows the systematic naming scheme:

>PDB00000038#CLASSI/LTR/LARD @Certhia brachydactyla

Metadata and taxonomic context

Additional taxonomic information—such as order, family, phylum, and higher ranks—as well as details about the origin of each sequence, are provided in the accompanying metadata file:

PanTEon_Database_metadata_v.1.6.1.csv

Benchmark edition

To enable fair and reproducible benchmarking of state-of-the-art TE classification tools, a benchmark edition of the PanTEon Database was generated. This version includes only TE superfamilies represented by more than 10 sequences and merges rare or uncommon superfamilies with their closest relatives (see the PanTEon paper for full details).

The benchmark dataset is available as:

PanTEon_Database_v1.6.1_benchmark_edition.fasta

Trained models for PanTEon Inference

This repository also contains pre-trained models corresponding to the different in-built architectures of PanTEon Platform (classification task). These models can be downloaded and used with the PanTEon inference module by specifying their paths via the -d parameter. This release includes models trained on the following datasets: all (entire PanTEon Database v1.6.1), Animalia, Chordata, Arthropods, Plantae, Angiosperms, Fungi, Ascomycota, and Basidiomycota, as well as models for discriminating between TEs and non-TE sequences. PanTEon Platform can be downloaded from the following GitHub repository: https://github.com/simonorozcoarias/PanTEon

Why PanTEon?

By combining broad taxonomic coverage, automated curation, and a standardized nomenclature, the PanTEon Database provides a robust reference resource for:

  • developing and training machine learning and deep learning models,

  • benchmarking TE classification tools,

  • exploring TE diversity across kingdoms.

Funding

Simon Orozco-Arias is supported by a fellowship within the “Generación D” initiative, Red.es, Ministerio para la Transformación Digital y de la Función Pública, for talent attraction (C005/24-ED CV1). Funded by the European Union NextGenerationEU funds, through PRTR.

Toni Galbadón group acknowledges support from the Spanish Ministry of Science and Innovation (grant numbers  PID2021-126067NB-I00 and PLEC2023-010225) cofounded by ERDF “A way of making Europe”, as well as support from the Gordon and Betty Moore Foundation (grant number GBMF9742); the Catalan Research Agency (AGAUR) (grant number 2022 INNOV 00065, 2024 PROD 00175 and 2024 PROD 00043); “La Caixa” foundation (grant number LCF/PR/HR21/00737 and CI23-20260); Fundació La Marató de TV3 (202328-31); AECC (PRYGN234923GABA and 290059); Instituto de Salud Carlos III (CIBERINFEC CB21/13/00061- ISCIII-SGEFI/ERDF and DTS25/00141); European Commission, Horizon Europe-HORIZON-MSCA-2023-DN-01-01 (grant number 101168618) and European Union’s Horizon 2020 research and innovation programme under the Marie Skłodowska-Curie grant agreement Nº 101226544 (grant number 101227078).

Alexandre R. Paschoal is supported by Fundação Araucária with NAPI Bioinformática (grant number 66.2021) and Brazilian National Research Council (CNPq - grant number 440412/2022-6).

 

Files

PanTEon_Database_metadata_v1.6.1.csv

Files (18.1 GB)

Name Size
md5:7f4ea0ef9c242d2ccf4c0ee723c9f08c
371.4 kB Preview Download
md5:222bfd9899d866ca06cda062fc3c52ff
992.2 MB Download
md5:4788d6697875d4e31f9a09cc679f6f4a
973.8 MB Download
md5:c952ee9db3d3a48b21b4ad83c6c7ab5b
1.4 GB Preview Download
md5:17fccd0b4d26ad3423005d6a0570d7b1
1.4 GB Preview Download
md5:de502caeb596ff6e39e09eea5f9991f7
1.5 GB Preview Download
md5:3f783d0bf0d789126c76caedc4af57f7
1.7 GB Preview Download
md5:e9f9b3a3f24b101aa9e8b06603fdddbc
1.8 GB Preview Download
md5:da248c94a25951b7690f7279ec89e5e5
1.8 GB Preview Download
md5:43d3c7d2c4fa7b3761939f2a77059b70
1.8 GB Preview Download
md5:459e076b55a931375aee13f843130215
1.6 GB Preview Download
md5:ba04ba10bc35d0f4f7b4299b2cae3920
1.4 GB Preview Download
md5:26c549973ffebbe58f6f3a53c6bf477f
1.8 GB Preview Download

Additional details

Funding

Barcelona Supercomputing Center
AI4Science Fellowship