Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72
Industry
Source: NVIDIA 技术博客Publish time unverified
Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token... Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token falls, communication increasingly determines how efficiently models scale across thousands of GPUs.