Major OpenCLIP Update Integrates Audio Models and Variable Resolution

wightmanr · x · 2026-08-14

The author has made extensive modifications to the OpenCLIP library, focusing on supporting variable resolution/aspect ratio NaFlexViT encoders and matching WebDataset pipelines, alongside integrating existing audio CLAP models.

Furthermore, the author experimented with combining NaFlex and CLAP by adding a mel patch embedder to the timm NaFlexViT base, achieving variable time capability for audio processing without major architectural changes. A new 'modern-text' encoder, utilizing recent ideas similar to ModernBERT, was also introduced. Preliminary training of modest-sized models has successfully validated the architecture and code changes.

Related event: OpenCLIP Update Introduces NaFlexCLAP for Audio-Text Multimodality(2 posts)→

Original post →

More from Multimodal

Multimodal channel →