Skip to content
worth noting Open-source

Olmo-core 3 introduces redesigned open infrastructure for training large MoE models

clearly official source

Olmo-core 3 has been released, providing infrastructure that keeps experts in GPU memory and routes data to them when training MoE models. According to the authors, it achieved approximately 2.7 times the throughput of the previous implementation in a preliminary test on eight NVIDIA B300 GPUs.

Olmo-core 3 has been released as open infrastructure for training large MoE models. It replaces the previous FSDP-based implementation with a DDP-based system: experts remain in GPU memory, and the system routes the relevant data to them. This eliminates repeated gathering and redistribution of weights. The infrastructure also distributes experts, model layers and optimizer state across GPUs.

According to the authors, increasing the number of experts from 8 to 128 reduced training throughput by less than 5 %. Each token continued to use four experts and approximately 3.2 billion active parameters, while the total parameter count grew from 4.6 to 47 billion. In a preliminary test of a model with 47 billion parameters on eight NVIDIA B300 GPUs, the new implementation achieved 52 000 tokens per second per GPU, compared with the previous 19 400, or approximately 2.7 times the throughput.

The infrastructure also supports the MXFP8 format. According to the authors, in a controlled test on four NVIDIA B300 GPUs with work evenly distributed among experts, it increased throughput by approximately 21 % compared with BF16 and reduced peak active memory usage from 103 GiB to 95 GiB. The benefit also depends on the overhead of conversion between numerical formats.

The authors report testing a model with 1.2 trillion parameters, of which 58.36 billion were active per token, on 512 NVIDIA B300 GPUs. The highest observed throughput reached 858 TFLOP/s per GPU. These tests used random routing and measured infrastructure performance, rather than the quality of the trained model.

What changed

Why it matters

For researchers and teams training MoE models, the ability to increase total model capacity while keeping the number of active parameters per token nearly unchanged is significant. Higher throughput may shorten training and reduce the GPU time required. However, the reported results come from specific configurations using NVIDIA B300 GPUs and do not, on their own, demonstrate savings in every operational setting or the quality of the resulting model.

Two audiences, two different impacts

What this means

01

For individuals

Researchers and developers training MoE models gain open infrastructure that lets them control how the model is distributed across GPUs and choose numerical precision.

What to do When testing the Olmo-core 3 infrastructure, compare throughput and memory usage with BF16 and MXFP8 on your own training configuration.
More practical updates →
02

For a business

For companies training their own MoE models, higher throughput may change the GPU time required and compute capacity planning. However, the documented tests do not quantify financial savings in a specific company's operations.

Development
What to decide Before changing the training budget, verify throughput on your company's configuration with routing suited to the intended model.
More business impacts →
DDP FSDP MoE NVIDIA B300 Olmo-core 3

Check the original

Event sources

clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
Hugging Face Blog primary source · first detected Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs