OpenCLIP Next Release to Support CLAP, CoCa, and Multimodal Architectures
wightmanr · x · 2026-08-21
OpenCLIP is preparing for its next major release with support for multiple model architectures and objectives:
- CLIP / SigLIP: Image+text encoders with contrastive learning.
- CLAP: Audio+text encoders with contrastive learning.
- CoCa(2): Img encoder + text encoder + cross-attn decoder, combining contrastive and caption CE loss.
- MaMMUT(2): Img encoder + single text decoder run twice (contrastive pass + caption pass).
- GenLIP / GenLAP: No separate encoder; single trunk over [image|audio ; text] with prefix-LM caption CE loss.
All are composable with NaFlex native-aspect ViT, modern text towers, and fused caption CE + z-loss.
Related event: OpenCLIP's Next Release Adds Multi-Architecture Support(2 posts)→
More from Research
- Harvey Details Post-Training Gains for Specialized Legal Intelligence — HamelHusain · 2026-08-21
- Scholar Criticizes ARR Review Quality, Notes Lack of LLM Disclosure — TuhinChakr · 2026-08-21
- SineKAN Replaces B-Splines with Sine Functions for Faster Inference — burkov · 2026-08-21
- Tech Optimist joins HDC Labs to explore hyperdimensional computing — rjurney · 2026-08-21
- Meta previews WildArtifactBench to evaluate multimodal agents — AIatMeta · 2026-08-21
- GoodfireAI launches $1M grants for AI interpretability research — niloofar_mire · 2026-08-21