Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72
产业动态
来源:NVIDIA 技术博客发布时间待核实
Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token... Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token falls, communication increasingly determines how efficiently models scale across thousands of GPUs.