Vector Institute Shares 35B MoE Fine-Tuning Practices
Vector Institute detailed their successful approach to fine-tuning a 35-billion-parameter multimodal MoE on just four A100 GPUs, overcoming previous out-of-memory crashes by combining 4-bit quantization, Flash Attention 2, and keeping MoE layers unsharded.
2026-07-30 ~ 2026-07-30 · 3 related posts
- Fine-Tuning 35B Multimodal MoE on 4 A100s: 5 Configs Failed — VectorInst · 2026-07-30
- Successfully Fine-Tuning 35B MoE: Insights into Model Routing — VectorInst · 2026-07-30
- Vector Institute Demystifies MoE: Slashes Logit Memory from 23.3GB to 0.3GB — VectorInst · 2026-07-30