OpenCLIP's new native ModernText encoder: more customizable than Transformers ModernBERT, decoder-capable

wightmanr · x · 2026-09-04

After OpenCLIP added a native ModernText encoder, users asked how it differs from the pretrained ModernBERT models available via Hugging Face Transformers. The author explains both share similar goals, but the native implementation offers more config knobs and easier customization, while the Transformers versions provide solid pretrained weights with more pooling constraints.

Two additional notes: ModernText extras are easy to customize (suggestions welcome), and it's not just an encoder — the additions were designed to work as encoder and/or decoder for models like MaMMUT.

Related event: OpenCLIP Maintainers Explain Native ModernText vs Transformers ModernBERT(2 posts)→

Original post →

More from Research

Research channel →