Published October 2, 2026 | Version v1.1.0

TFBindFormer Dataset v1.1.0

Authors/Creators

Description

This dataset contains the training, validation, and test data used for transcription factor (TF)–DNA binding prediction in the revised TFBindFormer framework. It integrates genomic DNA sequence data, transcription factor protein sequence and structural information, processed TF representations, and metadata to support reproducible model training, evaluation, and zero-shot testing on unseen transcription factors.

DNA Sequence Data (dna_data/)

The dna_data directory contains one-hot-encoded genomic DNA sequence data and corresponding TF-binding labels, organized into three mutually exclusive dataset splits:

  • train/ – training data

    • train_data.npy
    • train_labels.npy
  • val/ – validation data

    • val_data.npy
    • val_labels.npy
  • test/ – held-out test data

    • test_data.npy
    • test_labels.npy

The DNA inputs correspond to genomic sequence windows used by TFBindFormer for TF-binding prediction. The label arrays contain the corresponding TF-binding annotations for each DNA sequence.

The genomic data are partitioned using chromosome-based splits. The training set contains chromosomes other than chr4, chr7, chr8, chr9, and chrY. Chromosomes chr4 and chr7 are used for validation, while chr8 and chr9 are reserved for testing.

Each input sequence represents a 1,000-bp genomic window centered on a 200-bp genomic bin. A TF-binding label is considered positive when at least 50% of the corresponding 200-bp bin overlaps a ChIP-seq narrowPeak region.

Metadata (metadata/)

The metadata directory contains metadata describing the TF–cell-type prediction tasks and the cell-type identifiers used by the model.

  • seen_tf_metadata.tsv – metadata for TF–cell-type tasks associated with transcription factors represented during model training.
  • unseen_tf_metadata.tsv – metadata for TF–cell-type tasks associated with transcription factors excluded from training and reserved for zero-shot evaluation.
  • seen_cell_type_ids.npy – encoded cell-type identifiers corresponding to the seen-TF tasks.
  • unseen_cell_type_ids.npy – encoded cell-type identifiers corresponding to the unseen-TF tasks.

The metadata files provide the mapping between transcription factors, cell types or experimental conditions, and the corresponding prediction tasks.

The revised TFBindFormer dataset contains:

  • 100 seen TFs corresponding to 422 TF–cell-type tasks.
  • 8 unseen TFs corresponding to 35 TF–cell-type tasks.

The unseen TFs are completely excluded from model training and are used to evaluate zero-shot generalization to previously unseen transcription factors.

Transcription Factor Data (tf_data/)

The tf_data directory contains the transcription factor protein sequence, structural, and processed representation data used by the TFBindFormer protein encoder.

Amino-Acid Sequences (aa_sequences/)

The aa_sequences directory contains amino-acid FASTA sequences for the transcription factors(seen+unseen) included in the dataset.

These protein sequences provide the sequence-based information used to generate pretrained TF representations.

Protein Structures (pdb_structures/)

The pdb_structures directory contains protein structure files in PDB format for the transcription factors(seen+unseen).

These structures are used to derive Foldseek 3Di structural representations.

Foldseek 3Di Structural Sequences (foldseek_3Di_ss.fasta)

foldseek_3Di_ss.fasta contains precomputed Foldseek 3Di structural token sequences derived from the TF protein structures.

The 3Di representation converts local three-dimensional protein environments into discrete structural tokens, providing structure-aware information that complements the amino-acid sequence representation.

ProstT5 Embeddings (prostt5_embeddings/)

The prostt5_embeddings directory contains precomputed TF protein embeddings(seen+unseen) generated using ProstT5-based sequence and structural representations.

These embeddings are stored before conversion to the fixed-length protein representation used directly by TFBindFormer.

Fixed-Length TF Representations (fixed_length_200/)

TF proteins vary in sequence length. Before being provided to the TFBindFormer model, their protein representations are standardized to a fixed length of 200 tokens.

The fixed_length_200 directory contains the final model-ready TF representations and is divided into seen and unseen TF sets:

Seen TFs (fixed_length_200/seen_tf/)

  • fixed_tf_embs.pt
  • fixed_tf_masks.pt
  • tf_names_in_label_order.tsv

fixed_tf_embs.pt contains the fixed-length protein embeddings for TFs represented during model training.

fixed_tf_masks.pt contains the corresponding masks used to distinguish valid protein-token positions from padded positions.

tf_names_in_label_order.tsv records the TF ordering corresponding to the protein tensor and TF-binding label ordering. This file ensures that each protein representation is correctly matched to the corresponding prediction task.

Unseen TFs (fixed_length_200/unseen_tf/)

  • fixed_tf_embs.pt
  • fixed_tf_masks.pt
  • tf_names_in_label_order.tsv

These files contain the corresponding fixed-length representations, masks, and TF ordering information for the transcription factors held out from training.

The unseen-TF representations are used only for zero-shot evaluation.

TF Protein Processing Workflow

The TF protein data follow the general processing workflow:

Amino-acid sequences
        ↓
     ProstT5
        ↓
sequence-based representations

Protein structures
        ↓
     Foldseek
        ↓
  3Di structural sequences
        ↓
structure-aware representations

Sequence + structural representations
        ↓
fixed-length processing
        ↓
200-token TF representations
        ↓
seen TF / unseen TF sets

Seen and unseen TFs are processed together during the initial protein preprocessing stages. They are separated only at the final fixed-length representation stage used by the model.

Intended Use

This dataset is intended for:

  • Training and evaluating TF–DNA binding prediction models
  • Reproducing the experiments reported in the revised TFBindFormer study
  • Evaluating zero-shot generalization to unseen transcription factors
  • Studying protein-conditioned DNA binding specificity
  • Investigating sequence- and structure-aware TF representations
  • Developing multimodal models integrating genomic DNA sequence, TF protein information, and cell-type context

The dataset is provided for research and academic use and is intended to support reproducible model training and evaluation in the TFBindFormer framework.

Files

Files (1.2 GB)

Name Size
md5:af7235ae175241174534748c1e0c473b
1.2 GB Download

Additional details

Related works

Is part of
Model: https://github.com/BioinfoMachineLearning/TFBindFormer (Other)

Dates

Updated
2026-10-02
TFBindFormer dataset v1.1.0