BigMac: Breaking the Pareto Frontier of Compute‑Memory Trade‑offs in Multimodal LLM Training
BigMac introduces a dependency‑safe nested pipeline that preserves LLM compute efficiency while reducing encoder and generator activation memory to O(1), delivering 1.08‑1.9× speedup and stable memory usage for large‑scale multimodal model training.
