TFBindFormer is a hybrid cross-attention Transformer model for transcription factor (TF)–DNA binding prediction. The model explicitly integrates transcription factor protein representations derived from amino-acid sequence and protein structural context with genomic DNA sequence bins, enabling position-specific modeling of protein–DNA interactions beyond sequence-only approaches.
The revised TFBindFormer framework additionally incorporates cell-type embeddings and supports zero-shot evaluation on unseen transcription factors.
Overview of TFBindFormer architecture.
Hybrid cross-attention module illustrating residue–nucleotide interactions.
TFBindFormer contains the following major components:
- A DNA encoder for extracting contextual representations from genomic DNA
- TF protein representations derived from amino-acid sequences and structure-derived 3Di information
- A hybrid cross-attention module for modeling interactions between TF residues and DNA positions
- Cell-type embeddings for incorporating cell-type-specific information
- Content-aware pooling and a prediction head for TF–DNA binding classification
TFBindFormer v1.1.0 introduces several major updates to the dataset construction and model framework.
In the previous dataset, forward and reverse-complement DNA sequences were treated as separate samples.
In the revised dataset, both orientations are incorporated within the same sample for each genomic window. As a result, the total number of genomic-window samples is approximately half that of the previous dataset.
In the original dataset, all 108 TFs were included in the training, validation, and evaluation sets.
In the revised dataset:
- 100 TFs are used for model training, validation, and seen-TF evaluation
- 8 TFs are completely excluded from training and validation
- The 8 held-out TFs are used only for zero-shot evaluation
This setting allows TFBindFormer to be evaluated on transcription factors that are never observed during model training.
Cell-type information is incorporated into TFBindFormer as an additional input representation, enabling the model to capture cell-type-specific differences in TF–DNA binding patterns.
- Hybrid cross-attention architecture for explicit TF–DNA interactions
- Integration of TF amino-acid sequence and protein structure information
- Cell-type-aware TF–DNA binding prediction
- Zero-shot evaluation on unseen transcription factors
- Genome-wide TF binding prediction under severe class imbalance
- Reproducible training and evaluation pipeline
The GitHub repository contains the TFBindFormer source code, model architectures, evaluation utilities, and documentation. The dataset is distributed separately through Zenodo.
TFBindFormer/
├── figures/
├── scripts/
│ ├── eval.py
│ ├── extract_tf_embeddings.py
│ ├── generate_3di_tokens.sh
│ ├── per_task_metrics.py
│ ├── reduce_or_pad_tf_embeddings.py
│ └── train.py
├── src/
│ ├── architectures/
│ │ ├── __init__.py
│ │ ├── binding_predictor.py
│ │ ├── cross_attention_encoder.py
│ │ └── tbinet_dna_encoder.py
│ ├── __init__.py
│ ├── model.py
│ └── utils.py
├── README.md
├── environment.yml
└── LICENSE
- scripts/train.py: Trains TFBindFormer using the revised dataset, cached DNA–TF pairs, fixed TF protein representations, and cell-type embeddings.
- scripts/eval.py: Evaluates trained models on both seen-TF and unseen-TF datasets.
- scripts/extract_tf_embeddings.py: Generates TF protein embeddings from amino-acid sequences and 3Di structural representations.
- scripts/reduce_or_pad_tf_embeddings.py: Converts TF protein embeddings to the fixed-length representation used by TFBindFormer.
- scripts/generate_3di_tokens.sh: Generates Foldseek-derived 3Di structural tokens from TF protein structures.
- scripts/per_task_metrics.py: Computes task-level evaluation metrics from saved predictions.
- src/architectures/: Core TFBindFormer neural-network components and cross-attention modules.
- src/model.py: PyTorch Lightning model wrapper for TFBindFormer.
- src/utils.py: Dataset loading, cached-pair handling, and DataModule utilities.
- figures/: Model architecture and framework figures.
- environment.yml: Conda environment and software dependencies.
The TFBindFormer v1.1.0 dataset contains the data required for model training, validation, seen-TF testing, and unseen-TF zero-shot evaluation.
The revised dataset includes:
- Genomic DNA sequence data
- Forward sequence representations within each sample
- TF–DNA binding labels
- TF amino-acid sequence data
- TF protein structural information
- Structure-derived 3Di representations
- Precomputed TF embeddings
- Cell-type identifiers
- Metadata describing TF/cell-type prediction tasks
- Seen-TF and unseen-TF dataset partitions
The dataset archive is:
TFBindFormer_dataset_v1.1.0.tar
The dataset is available through Zenodo:
DOI: 10.5281/zenodo.23050978
URL:
https://doi.org/10.5281/zenodo.23050978
The seen-TF dataset contains:
- 100 transcription factors, corresponding to 422 TF/cell-type prediction tasks
The genomic chromosome split is:
- Training: all chromosomes except chr4, chr7, chr8, chr9, and chrY
- Validation: chr4 and chr7
- Test: chr8 and chr9
The zero-shot evaluation dataset contains:
- 8 unseen transcription factors, corresponding to 35 TF/cell-type prediction tasks
These TFs are completely excluded from model training and validation and are used only for zero-shot evaluation.
All positive DNA–TF pairs are retained.
Negative pairs are sampled using predefined sampling fractions:
| Dataset | Negative sampling fraction |
|---|---|
| Training | 0.03 |
| Validation | 0.5 |
| Seen-TF test | 1.0 |
| Unseen-TF test | 1.0 |
Pair sampling can be independently generated using the specified random seed.
Clone the repository:
git clone https://github.com/BioinfoMachineLearning/TFBindFormer.git
cd TFBindFormerCreate the Conda environment:
conda env create -f environment.yml
conda activate tfbindformerTFBindFormer uses Foldseek-derived 3Di structural tokens to encode protein structural information.
The 3Di tokens used in this work are included in the released dataset. Users interested in recomputing 3Di representations from raw protein structures or applying the method to additional transcription factors may install Foldseek following the official documentation:
https://github.com/steineggerlab/foldseek
Ensure that the foldseek executable is available in your $PATH.
All DNA sequence data and TF-related data used by TFBindFormer v1.1.0 are available on Zenodo.
Dataset archive:
TFBindFormer_dataset_v1.1.0.tar
Zenodo record:
https://doi.org/10.5281/zenodo.23050978
After downloading the archive, place it under the TFBindFormer repository directory and extract it:
pwd
# .../TFBindFormer
tar -xf TFBindFormer_dataset_v1.1.0.tarThe provided files include the preprocessed DNA inputs, binding labels, cell-type identifiers, metadata, and TF-related data required for training and evaluation.
The 3Di tokens used in this study are already included in the released dataset.
To recompute 3Di tokens from protein structure files or generate 3Di representations for additional transcription factors, use:
pwd
# .../TFBindFormer
chmod +x scripts/generate_3di_tokens.sh
./scripts/generate_3di_tokens.sh \
<pdb_dir> \
<output_dir>Arguments:
<pdb_dir>:
Directory containing TF protein structure files in PDB format
<output_dir>:
Directory where generated 3Di token FASTA files will be saved
Example:
./scripts/generate_3di_tokens.sh \
data/tf_data/tf_structure \
data/tf_data/3di_outTFBindFormer represents transcription factors using embeddings derived from amino-acid sequences and 3Di structural tokens.
The TF protein embeddings used in this study are included in the released dataset.
To recompute TF protein embeddings from the provided amino-acid sequences and 3Di tokens, run:
pwd
# .../TFBindFormer
nohup python scripts/extract_tf_embeddings.py \
--aa_dir data/tf_data/tf_sequence \
--di_fasta data/tf_data/3di_out/pdb_3Di_ss.fasta \
--out_dir data/tf_data/tf_embeddings \
> extract_tf_embeddings.log 2>&1 &This command loads TF amino-acid sequences and the corresponding structure-derived 3Di token sequences and generates TF protein embeddings.
The TF protein embeddings generated in the previous step may have different sequence lengths. TFBindFormer uses a fixed protein representation length of 200 tokens for model training and evaluation.
The model-ready fixed-length TF protein embeddings used in this study are included in the released dataset.
To reproduce this preprocessing step from the variable-length TF embeddings,
reduce or pad the embeddings to a fixed length of 200 using
reduce_or_pad_tf_embeddings.py.
pwd
# .../TFBindFormer/scripts
python reduce_or_pad_tf_embeddings.py \
--metadata_tsv ../data/metadata/seen_tf_metadata.tsv \
--embedding_dir ../data/tf_data/prostt5_embeddings \
--out_dir ../data/tf_data/fixed_length_200/seen_tf \
--target_len 200 \
--method avgmax \
--dtype float16pwd
# .../TFBindFormer/scripts
python reduce_or_pad_tf_embeddings.py \
--metadata_tsv ../data/metadata/unseen_tf_metadata.tsv \
--embedding_dir ../data/tf_data/prostt5_embeddings \
--out_dir ../data/tf_data/fixed_length_200/unseen_tf \
--target_len 200 \
--method avgmax \
--dtype float16This preprocessing step:
- Converts variable-length TF embeddings to a fixed length of 200 tokens
- Uses the
avgmaxreduction strategy for embeddings longer than 200 tokens - Pads shorter embeddings to the target length
- Saves the corresponding TF masks
- Produces model-ready TF representations separately for seen and unseen TFs
The processed files are stored as:
data/tf_data/fixed_length_200/
├── seen_tf/
│ ├── fixed_tf_embs.pt
│ ├── fixed_tf_masks.pt
│ └── tf_names_in_label_order.tsv
└── unseen_tf/
├── fixed_tf_embs.pt
├── fixed_tf_masks.pt
└── tf_names_in_label_order.tsv
These fixed-length embeddings and masks are then used directly by
train.py and eval.py.
TFBindFormer v1.1.0 supports the revised dataset format and cell-type embeddings.
Run training from the scripts/ directory:
pwd
# .../TFBindFormer/scripts
mkdir -p ....../results/seed42
mkdir -p ....../results/seed42/ckpts
CUDA_VISIBLE_DEVICES=0 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
nohup python train.py \
--train_dna_npy ....../data/dna_data/train/train_data.npy \
--train_labels_npy ....../data/dna_data/train/train_labels.npy \
--train_metadata_tsv ....../data/metadata/seen_tf_metadata.tsv \
--val_dna_npy ....../data/dna_data/val/val_data.npy \
--val_labels_npy ....../data/dna_data/val/val_labels.npy \
--val_metadata_tsv ....../data/metadata/seen_tf_metadata.tsv \
--fixed_tf_embs_pt ....../data/tf_data/fixed_length_200/seen_tf/fixed_tf_embs.pt \
--fixed_tf_masks_pt ....../data/tf_data/fixed_length_200/seen_tf/fixed_tf_masks.pt \
--train_pairs_file ....../data/cached_pairs/seed42/seentf/train_pairs.npy \
--val_pairs_file ....../data/cached_pairs/seed42/seentf/val_pairs.npy \
--use_cell_type \
--cell_type_dim 16 \
--cell_type_ids_npy ....../data/metadata/seen_cell_type_ids.npy \
--protein_in_dim 1024 \
--epochs 20 \
--batch_size 1024 \
--num_workers 6 \
--lr 1e-4 \
--wandb_project tfbind \
--run_name tfbind_seed42 \
--output_dir ....../results/seed42/ckpts \
> ....../results/seed42/tfbind_train_seed42.log 2>&1 &This command trains TFBindFormer using preprocessed genomic DNA inputs, TF–DNA binding labels, fixed-length TF protein representations, cached DNA–TF pair files, and cell-type information.
- DNA sequence inputs are loaded from NumPy arrays
- Binding labels are loaded from the corresponding label matrices
- Fixed-length TF protein embeddings are loaded from
--fixed_tf_embs_pt - TF masks are loaded from
--fixed_tf_masks_pt - Pre-generated training and validation DNA–TF pairs are loaded from
--train_pairs_fileand--val_pairs_file - Cell-type embeddings are enabled with
--use_cell_type - Cell-type IDs are loaded from
--cell_type_ids_npy - Model checkpoints are selected based on validation AUPRC
- Early stopping is applied with a patience of 5 epochs
- Model checkpoints are saved to the directory specified by
--output_dir - Training metrics can be tracked using Weights & Biases
The seen-TF test set evaluates TF–DNA binding prediction on held-out genomic chromosomes for TFs represented during model development.
Run:
pwd
# .../TFBindFormer/scripts
mkdir -p ....../results/seed_42/seenTF_eval
CUDA_VISIBLE_DEVICES=1 PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
nohup python eval.py \
--ckpt_path ....../results/ckpts/epoch=16-val/pr_auc=0.5694-val/loss=0.1102.ckpt \
--test_dna_npy ....../data/dna_data/test/seen/test_data.npy \
--test_labels_npy ....../data/dna_data/test/seen/test_labels.npy \
--test_metadata_tsv ....../data/metadata/seen_tf_metadata.tsv \
--fixed_tf_embs_pt ....../data/tf_data/fixed_length_200/seen_tf/fixed_tf_embs.pt \
--fixed_tf_masks_pt ....../data/tf_data/fixed_length_200/seen_tf/fixed_tf_masks.pt \
--test_pairs_file ....../data/cached_pairs/seed42/seentf/test_pairs.npy \
--use_cell_type \
--cell_type_dim 16 \
--cell_type_ids_npy ....../data/metadata/seen_cell_type_ids.npy \
--protein_in_dim 1024 \
--batch_size 1024 \
--num_workers 6 \
--precision 16-mixed \
--wandb_project tfbind_eval \
--run_name eval_seed42_seenTF \
--output_dir ....../results/seed_42/seenTF_eval \
> ....../results/seed_42/seenTF_eval/eval_seenTF.log 2>&1 &This command loads the specified TFBindFormer checkpoint and evaluates performance on the held-out seen-TF test set.
TFBindFormer v1.1.0 includes zero-shot evaluation on 8 transcription factors that are completely excluded from model training and validation.
The trained model is not fine-tuned on these TFs. Instead, their protein representations are provided during evaluation, allowing TFBindFormer's ability to generalize to unseen transcription factors to be assessed.
The unseen-TF dataset can be evaluated using the same evaluation pipeline:
mkdir -p ....../results/seed_42/unseenTF_eval
CUDA_VISIBLE_DEVICES=1 \
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
nohup python eval.py \
--ckpt_path ....../results/ckpts/epoch=16-val/pr_auc=0.5694-val/loss=0.1102.ckpt \
--test_dna_npy ....../data/dna_data/test/unseen/test_data.npy \
--test_labels_npy ....../data/dna_data/test/unseen/test_labels.npy \
--test_metadata_tsv ....../data/metadata/unseen_tf_metadata.tsv \
--fixed_tf_embs_pt ....../data/tf_data/fixed_length_200/unseen_tf/fixed_tf_embs.pt \
--fixed_tf_masks_pt ....../data/tf_data/fixed_length_200/unseen_tf/fixed_tf_masks.pt \
--test_pairs_file ....../data/cached_pairs/seed42/unseentf/test_pairs.npy \
--use_cell_type \
--cell_type_dim 16 \
--cell_type_ids_npy ....../data/metadata/unseen_cell_type_ids.npy \
--protein_in_dim 1024 \
--batch_size 1024 \
--num_workers 6 \
--precision 16-mixed \
--wandb_project tfbind_eval \
--run_name eval_seed42_unseenTF \
--output_dir ....../results/seed_42/unseenTF_eval \
> ....../results/seed_42/unseenTF_eval/eval_unseenTF.log 2>&1 &Replace the placeholder paths with the corresponding unseen-TF data files from the released dataset.
The stacked cross-attention blocks illustrated above are implemented in:
src/architectures/cross_attention_encoder.py
src/architectures/binding_predictor.py
The number of cross-attention blocks, as well as their internal configuration including hidden dimension, number of heads, and dropout, can be adjusted by modifying the corresponding initialization parameters and module definitions.
The depth of the hybrid cross-attention module controls how many cross-attention blocks are stacked sequentially.
Each block models TF-DNA interactions through cross-attention followed by feed-forward transformations.
Advanced users may modify these settings to explore alternative model capacities or architectures.
- Combined forward and reverse-complement DNA sequences within each genomic-window sample
- Held out 8 transcription factors for zero-shot evaluation
- Added cell-type embeddings
- Added zero-shot evaluation on unseen TFs
- Updated the training and evaluation pipeline for the revised dataset
Initial public release of TFBindFormer.
TFBindFormer v1.1.0 contains updates beyond the version described in the original bioRxiv preprint.
If you use the TFBindFormer framework, please cite the original manuscript:
@article{liu2026tfbindformer,
title = {TFBindFormer: A Cross-Attention Transformer for Transcription Factor-DNA Binding Prediction},
author = {Liu, Ping and Wang, Lyuwei and Basnet, Shreya and Cheng, Jianlin},
journal = {bioRxiv},
year = {2026},
doi = {10.64898/2026.04.09.717563}
}
---
## Contact
For questions or issues related to the code or dataset, please open an
issue in this repository.
Additional inquiries may be directed to:
Ping Liu
Email: pl5vw@missouri.edu
