CLSS: Mapping Protein Sequence and Structure into a Shared AI Space
bravo_abad · x · 2026-08-07
Guy Yanai and coauthors introduce CLSS, a contrastive protein language model that learns a shared representation of protein sequence, structure, and even subsequences.
Architecture & Training
- Employs a two-tower architecture to process sequence and structure information separately.
- Trained contrastively to pull matching sequence-structure pairs together in latent space while pushing unrelated proteins apart.
- Self-supervised on 1 million protein domains without using ECOD or CATH labels.
Results & Significance
CLSS goes beyond multimodal alignment by learning a unified "protein space." In this space, sequence and structure yield nearly the same global organization, successfully recovering expert-curated evolutionary hierarchies that the model never saw during training.
More from Research
- Google Open-Sources WeatherNext 2: Generates 15-Day Forecasts in Under a Minute — aigclink · 2026-08-07
- Google Open-Sources WeatherNext 2: Generates 15-Day Forecasts on a Single TPU in Under a Minute — aigclink · 2026-08-07
- UT Nuremberg Opens PhD Position on VLA Models and 3D Geometry — y_m_asano · 2026-08-07
- Stanford and Arc Institute Use AI to Design Working Bacteria-Killing Viruses — The Decoder · 2026-08-07
- RL Training Collapses at Step 8: Dev Implements LoRA to Fix GPU Sync — mervenoyann · 2026-08-07
- Noether: Open-Source Tool to Automatically Prove Math Properties of JAX Code in Lean — jonkhler · 2026-08-07