OpenCLIP next release to support multiple architectures
wightmanr · x · 2026-08-21
The next major release of OpenCLIP will support several model architectures:
- CLIP / SigLIP: Image+text encoders with contrastive learning.
- CLAP: Audio+text encoders with contrastive learning.
- CoCa: Image encoder + text encoder + cross-attn text decoder, combining contrastive and caption CE loss.
- MaMMUT: Image encoder + text decoder run twice (contrastive pass + caption pass).
- GenLIP / GenLAP: No separate encoder, single trunk over [image|audio ; text], prefix-LM caption CE loss only.
All composable with NaFlex native-aspect ViT, modern text towers, and fused caption CE + z-loss.
Related event: OpenCLIP's Next Release Adds Multi-Architecture Support(2 posts)→
More from Infra
- Opposition to local data centers in US surges 33 points to 75% — Polymarket · 2026-08-21
- Ramp Router cuts GPT-5.6 Sol inference costs by 50% — KlausCodes · 2026-08-21
- Researcher Rants: Conference Season Blocks GPU Access for Days — ChongZzZhang · 2026-08-21
- AT&T routes 40% of employee AI usage to open models — Hesamation · 2026-08-21
- Memory and silicon production set to 4x; older chips sufficient for future models — teortaxesTex · 2026-08-21
- Alibaba's Qwen Open Models Drive Cloud Growth, $56B AI Spend pays off — TiernanRayTech · 2026-08-21