← Founder Notes
Archive

Dropless moe training just went 10x faster on gpus. nvidia's transformer engine work, out september…

Yethikrishna ROriginal on Threads

dropless moe training just went 10x faster on gpus. nvidia's transformer engine work, out september 14, processes every token assigned to an expert without dropping under load imbalance, hitting 97% scaling across 1,024 gpus.

the bottleneck is now the network, not the math.

Context

NVIDIA's technical blog of 14 September 2026, Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine, reports 10.4x throughput on GB200 for DeepSeek-V3 (103 to 1,068 TFLOPS per GPU) and 97 percent scaling efficiency at 1,024 GPUs on GB300 NVL72 training DeepSeek-V3 671B. Dropless MoE processes every token without dropping or padding. The post says communication took 84 percent of accumulated kernel time in the unoptimized baseline.

How it compares

Both figures are first-party and vendor-reported. They are different configurations: the 10x is on GB200 against an unoptimized baseline of the blog's own, and the 97 percent is on GB300 NVL72 at 1,024 GPUs; the note combines them. The bottleneck is now the network is the author's take; the 84 percent describes the baseline and not a current bottleneck.

Watch next

  • Third-party reproduction.

Sources

  1. Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine (NVIDIA, 14 Sep 2026)developer.nvidia.com

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 20 September 2026 at 22:04 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/dropless-moe-training-just-went-10x-faster-on-DdhBm0SEV3x" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Dropless moe training just went 10x faster on gpus. nvidia's transformer engine work, out september…"></iframe>

More notes