MaMMUT2 Architecture Validation: Multiple CC12M Training Runs Compared
wightmanr · x · 2026-08-20
The author conducted several small-scale (2x GPU) MaMMUT2 weight training runs on the CC12M dataset to validate architectural variations and compare against CLIP baselines.
Configurations and Results:
- Included variants like Modern Text decoder and NaFlex ViT encoders.
- Compared different learning rates (1e-3, 2e-3, 3e-3) and QK normalization settings.
- Some models showed competitive Zero-Shot Image Classification metrics.
- All experiments were completed within 32 epochs.
Related event: OpenCLIP Adds MaMMUT Support and MaMMUT2 Validation Experiments(4 posts)→
More from Research
- RSI is not a perpetual motion machine; the real bottleneck is compute — shuchaobi · 2026-08-21
- New Paper Proposes 'Spectral Neuron' Using Eigenvalues as Non-Linearity — CatAstro_Piyush · 2026-08-21
- Microsoft releases Skala 1.1 DL functional for computational chemistry — vdbergrianne · 2026-08-21
- Open Source A2A Adapter Enables Interoperability Between AI Agent Frameworks — kevinlu310 · 2026-08-21
- Counterintuitive LLM Inference: Batching, Quantization, and Speculative Decoding Pitfalls — techNmak · 2026-08-21
- Paper: AI Scientists Should Be Studied as Human-Agent Systems — rohanpaul_ai · 2026-08-21