Mol-JEPA: multimodal joint-embedding predictive architecture hits SOTA on molecule benchmarks

randall_balestr · x · 2026-08-26

A new arXiv paper, "Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules" (Rottach, Balestriero, et al.), proposes a scalable framework for learning molecular world models. It targets three persistent limitations of molecular foundation models: chemically invalid augmentations, modality collapse, and incomplete representation of biochemical environments.

Instead of suboptimal molecular perturbations, Mol-JEPA uses modality masking to jointly exploit molecular structures, cellular phenotypes, binding affinities, ADMET profiles, quantum chemistry simulations, and other drug discovery data via latent-space prediction. The learned representations deliver SOTA performance across multiple benchmarks, and the paper provides numerous practical guidelines for multi-modality (>2) pretraining from scratch.

Related event: Mol-JEPA: Molecular Foundation Model Learning Across 14 Modalities(2 posts)→

Original post →

More from Research

Research channel →