Embodied Intelligence Observer

Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72

Industry

Source: NVIDIA 技术博客Publish time unverified

Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token... Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token falls, communication increasingly determines how efficiently models scale across thousands of GPUs.

Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72 | Embodied Intelligence Observer