Peking University & Xiaohongshu Introduce BigMac: 1.9x Faster Multimodal Training with Stable Memory

jiqizhixin · x · 2026-08-04

Researchers from Peking University, independent researchers, and Xiaohongshu introduce BigMac, a new training pipeline for multimodal large language models. BigMac cleverly nests the encoder and generator work inside the standard LLM training pipeline, executing them in a safe, dependency-friendly order.

This design slashes the memory required during training to a constant amount, allowing the system to run at full speed without actually requiring unlimited memory. Experiments show that BigMac delivers up to 1.9x faster training compared to existing systems while keeping memory usage stable even as batch sizes scale up, breaking the long-standing Pareto frontier trade-off between compute and memory in multimodal AI training.

Original post →

More from Research

Research channel →