There is a newer version of the record available.

Published February 27, 2026 | Version V1.0

ProteinAligner

Authors/Creators

  • 1. EDMO icon University of California, San Diego

Description

ProteinAligner is an innovative multimodal protein representation framework that integrates structural, sequential, and textual data into a unified embedding space. It employs three distinct encoding pathways for the amino acid sequence, the 3D structure, and the textual description of a protein, using the sequence as the anchor modality for alignment. The sequence pathway processes the one-dimensional amino acid sequence, generating representations that capture the molecular characteristics of each amino acid using the ESM-2 protein language model. The structure pathway, utilizing the ESM-IF1 model, processes the three-dimensional structure, capturing the protein's molecular interactions and dynamics. The textual pathway employs a vanilla 8-layer Transformer encoder to process textual descriptions derived from experimentally verified publications, creating unique representations for each protein. To ensure compatibility across modalities, each encoded input is projected to the same dimension.

The framework is trained on a large-scale dataset of 150,000 (structure, sequence, description) triples, sourced from the UniProtKB/Swiss-Prot and RCSB PDB databases. The dataset was assembled by mapping PDB IDs to UniProt IDs to retrieve triples for the same protein. During the pretraining stage, ProteinAligner aims to minimize the contrastive loss between sequence-structure and sequence-text pairs, aligning the embeddings of structures and textual descriptions with sequence embeddings to form a unified representation. This pretraining stage leverages the sequence as the anchor modality to facilitate this alignment process.

The training of ProteinAligner follows a two-stage pretrain-finetune pipeline. Initially, the pretraining stage involves computing the contrastive loss on sequence-paired data, specifically focusing on sequence-structure and sequence-text pairs. This contrastive learning approach ensures that the embeddings of protein structures and textual descriptions align with the sequence embeddings of the same protein. Following this, the fine-tuning stage integrates the pretrained encoder weights with task-specific layers, enabling the application of ProteinAligner to a variety of domain-specific tasks. This structured approach allows ProteinAligner to leverage the strengths of each modality, creating a robust and versatile framework for protein representation.

Files

ProteinAligner-1.0.zip

Files (17.5 MB)

Name Size Download all
md5:ec254ed589afd8e79ec6883e352b97d8
17.5 MB Preview Download

Additional details

Funding

U.S. National Science Foundation
SCH: Develop Clinical Time Series Foundation Models for Sepsis Early Detection 2405974
U.S. National Science Foundation
CAREER: Mitigating the Lack of Labeled Training Data in Machine Learning Based on Multi-level Optimization 2339216
National Institutes of Health
Develop Multi-modal Foundation Models for Sepsis Early Detection 1R35GM157217-01
National Institutes of Health
Safe Continual Learning for Sepsis Early Detection 1R21GM154171-01