Major OpenCLIP Update Integrates Audio Models and Variable Resolution
wightmanr · x · 2026-08-14
The author has made extensive modifications to the OpenCLIP library, focusing on supporting variable resolution/aspect ratio NaFlexViT encoders and matching WebDataset pipelines, alongside integrating existing audio CLAP models.
Furthermore, the author experimented with combining NaFlex and CLAP by adding a mel patch embedder to the timm NaFlexViT base, achieving variable time capability for audio processing without major architectural changes. A new 'modern-text' encoder, utilizing recent ideas similar to ModernBERT, was also introduced. Preliminary training of modest-sized models has successfully validated the architecture and code changes.
Related event: OpenCLIP Update Introduces NaFlexCLAP for Audio-Text Multimodality(2 posts)→
More from Multimodal
- Customuse Launches MCP to Automate 3D Asset Generation in Claude — JaynitMakwana · 2026-08-14
- AI Avatar Music Video: Psychedelic Cyberwave Anime — VIV-AF-D · 2026-08-14
- Creating a 10-Minune Short Film for Kids Using MiniMax H3 and Multi-Tools — davekilljoy · 2026-08-14
- MiniMax Music Wins Praise: User Cancels Suno Subscription — cointalkz · 2026-08-14
- MiniMax H3 Test: Giant Turtle Climbing with 4x Upscaling — LionLikeMan · 2026-08-14
- ElevenMusic Launches Free AI-Powered Sounds and Loops Library — lukeharries · 2026-08-14