Fine-Tuning 35B Multimodal MoE on 4 A100s: 5 Configs Failed

VectorInst · x · 2026-07-30

Vector's AI engineering team attempted to fine-tune a 35.26-billion-parameter multimodal MoE on four A100 GPUs, but all five configurations across two frameworks failed.

The crashes were primarily driven by out-of-memory errors (peaking at 77-78 GB against an 80 GB limit) and optimizer crashes. The team identified the root cause: data-dependent routing in MoE is fundamentally incompatible with gather-on-demand parameter sharding.

Related event: Fine-tuning 35B MoE on 4 A100s(2 posts)→

Original post →

More from Research

Research channel →