Open-Sourcing NaFlexCLAP: Variable-Resolution Audio-Text Models
wightmanr · x · 2026-08-14
Prominent vision model researcher Ross Wightman has shared his latest experimental model collection, NaFlexCLAP, now open-sourced on Hugging Face.
- Architecture: The project combines the timm NaFlexViT audio encoder with a modern text encoder to build CLAP (Contrastive Language-Audio Pretraining) models supporting variable time and variable length inputs.
- Engineering: The author made extensive modifications to the OpenCLIP main branch to support variable resolution/aspect ratio encoders and matching webdataset pipelines, integrating existing audio CLAP weights.
- Availability: Multiple model versions (ranging from 160M to 240M parameters) are currently available and can be used directly via the OpenCLIP main branch.
Related event: OpenCLIP Update Introduces NaFlexCLAP for Audio-Text Multimodality(2 posts)→
More from Multimodal
- Gemini 3.7 Generates a 3D Rolex in Pure Three.js for Just $0.038 — rohanpaul_ai · 2026-08-14
- NVIDIA x Runway: Gen-4.5 Integrated into Vera Rubin Platform in One Day — nvidia · 2026-08-14
- MiniMax Music 3 Model Now Available Locally in ComfyUI — PurzBeats · 2026-08-14
- ComfyUI Node Highlight: A More Flexible Resize Image/Mask Alt — altoiddealer · 2026-08-14
- Testing MiniMax Music 3.0: Crisp Banjos and Stunning Country Style — Nowawes · 2026-08-14
- LTX-2.5 Video Generation: How to Structure Effective Multishot Prompts — Interesting_Room2820 · 2026-08-14