OpenCLIP Adds MaMMUT Support: Fixes Pooling, Introduces MaMMUT2
wightmanr · x · 2026-08-20
OpenCLIP now supports the MaMMUT architecture (image encoder + text decoder), enabling two-pass contrastive and generative training steps.
Key Updates:
- Integrates existing OpenMaMMUT weight loading without a separate fork.
- Fixes multiple pooling issues from the original fork.
- Introduces a new MaMMUT '2' variant with fixed pooling mechanisms.
- Adds support for NaFlex ViT encoders and 'Modern Text' decoders.
- Includes new configs for fused/chunked linearcrossentropy and z-loss.
Related event: OpenCLIP Adds MaMMUT Support and MaMMUT2 Validation Experiments(4 posts)→
More from coding & agent
- NVIDIA Tutorial: Post-Train Cosmos 3 Edge for On-Device Robot Control on Jetson Thor — MonaJalal_ · 2026-08-20
- PR count or LOC aren't great metrics: Agents can farm bug fix numbers — Daniel_Farinax · 2026-08-20
- New agent prompting method: Feed intuition instead of independent scaffolding — Sauers_ · 2026-08-20
- Sharing layers across workers via a "prelude" Effect in alchemy — samgoodwin89 · 2026-08-20
- Hands-on with Instinct, Grok Bots and ChatGPT Work: personal agents compared — illscience · 2026-08-20
- AI Engineering Updates for Aug 20: Runtime Fixes, Inference Overhead, Agent Research — ZestycloseTie1793 · 2026-08-20