Hugging Face Releases Olmo-core 3 for Trillion-Parameter MoE Training
Hugging Face has introduced Olmo-core 3, an open training infrastructure designed to scale mixture-of-experts models into the trillion-parameter range while maintaining efficiency.

Training large artificial intelligence models demands substantial computing resources, which increases costs and energy consumption. This dynamic often places advanced model development out of reach for academic researchers and smaller laboratories. Mixture-of-experts (MoE) architectures offer a more efficient alternative by containing many parameter components without requiring every input to utilize all of them simultaneously. However, scaling MoEs introduces distinct communication and coordination challenges across GPU clusters, which can reduce the computational advantages of activating only a portion of the model per input.
To address these hurdles, Hugging Face has released Olmo-core 3, a redesigned open training framework for large language models. The updated system shifts from fully sharded data parallelism to a distributed data parallelism approach. This modification keeps experts resident on the GPUs and routes relevant data directly to them, thereby eliminating the need for repeated weight gathering. In preliminary tests on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE achieved a throughput of 52,000 tokens per second per GPU, representing roughly 2.7 times the throughput of the earlier implementation.
Scaling and Optimization Techniques
Olmo-core 3 combines several methods to distribute large MoE architectures across hardware clusters efficiently. Expert parallelism spreads experts across GPUs, pipeline parallelism divides model layers across groups of devices, and a distributed optimizer shares optimizer states instead of duplicating them fully on every card. Furthermore, the framework incorporates rowwise expert parallelism and GPU-resident routing to minimize data rearrangement overhead, alongside grouped GEMM to execute small expert computations together.
The infrastructure also supports MXFP8, a lower-precision number format that reduces computational load and data transfer between GPUs. In controlled benchmarks on four NVIDIA B300 GPUs, enabling MXFP8 across critical system components boosted end-to-end training throughput by approximately 21% compared to a BF16 baseline, while peak active memory decreased from 103 GiB to 95 GiB.
When exploring cutting-edge architectures, developers can try top AI models cheaply through one API at https://apixoai.online.
What it means for developers
For developers and researchers working on open-source AI, Olmo-core 3 provides an adaptable infrastructure capable of scaling into the trillion-parameter range. Hugging Face benchmarked the framework on 512 NVIDIA B300 GPUs with a 1.2-trillion-parameter configuration, achieving a peak throughput of 858 TFLOP/s/GPU, and tested configurations reaching up to 2.38 trillion total parameters using alternative communication layers like DeepEP v2.
Beyond raw performance metrics, the technical report details specific engineering lessons, such as identifying potential metric distortions like token gerrymandering, evaluating learning rate adjustments, and analyzing the trade-offs of overlapping communication and computation on separate GPU streams. As the foundation for the next generation of Olmo models, the fully open-source framework allows developers to experiment with routing, parallelism, and hardware adaptations for their own MoE projects.
Source: Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs — Hugging Face blog. Written by the Apixo team from that report.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key

