Fine-Tuning 35B Multimodal MoE on 4 A100s: 5 Configs Failed
VectorInst · x · 2026-07-30
Vector's AI engineering team attempted to fine-tune a 35.26-billion-parameter multimodal MoE on four A100 GPUs, but all five configurations across two frameworks failed.
The crashes were primarily driven by out-of-memory errors (peaking at 77-78 GB against an 80 GB limit) and optimizer crashes. The team identified the root cause: data-dependent routing in MoE is fundamentally incompatible with gather-on-demand parameter sharding.
Related event: Fine-tuning 35B MoE on 4 A100s(2 posts)→
More from Research
- UMich team gets $750K NSF grant for transparent AI learning tools — QVeraLiao · 2026-07-30
- AI Can De-anonymize Internet Users for Under $4 with 90% Accuracy — luisdans · 2026-07-30
- GRAM Paper: Precisely Removing LLM Dangerous Capabilities via Gradient-Routed Modules — burkov · 2026-07-30
- Yale Scholar Shares Practical Guide on Using AI in Household Finance Research — TaniaBabina · 2026-07-30
- Experiment: Injecting Fictional Lore as Executable System Prompt into RAG Crawlers — BitcoinsOrganizer · 2026-07-30
- llamppl: Integrating Large Language Models into Probabilistic Programming — xuanalogue · 2026-07-30