BigMac Accelerates Multimodal Training with Nested Pipelines

The new BigMac parallel training paradigm uses a dependency-safe nested pipeline to optimize native multimodal training. By intelligently inserting encoder and generator computations into the main LLM pipeline, it significantly reduces memory bubbles and accelerates training by up to 1.9x.

2026-07-22 ~ 2026-07-24 · 2 related posts