AI Models & Platforms

Ai2 Releases Olmo-Core 3, Open Training Stack for Trillion-Parameter MoEs

mm
Add Unite.AI to your preferred sources on Google

The Allen Institute for AI released Olmo-core 3 on October 1, 2026, an upgrade to its large language model development framework built around a redesigned open mixture-of-experts training system that Ai2 said it has benchmarked at more than one trillion total parameters.

Ai2 said Olmo-core 3 is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency, and described it as one of the core systems behind the next generation of Olmo. The release is accompanied by a technical report, an interactive demo, and open code.

MoE models can contain many more parameters without requiring every input to use all of them, but the full model still has to be stored across GPU memory and updated during training, and routing inputs to the right experts across a cluster creates its own communication and coordination costs. Ai2 said those costs can erode much of the computational advantage as MoEs grow, and that Olmo-core 3 is built to close that gap. In one Ai2 benchmark, the expert pool grew from 8 to 128 while still selecting four experts per token, keeping active parameters roughly fixed at about 3.2 billion while total parameter capacity grew from 4.6 billion to 47 billion, with training throughput falling by less than 5%.

From FSDP to a DDP-Based Training Stack

Ai2’s earlier MoE implementation in Olmo-core used fully sharded data parallelism, configured to gather and reshard model weights for each small batch of training data. Olmo-core 3 switches to a system based on distributed data parallelism that keeps experts resident on GPUs and routes the relevant data to them, avoiding that repeated weight gathering. In a preliminary test on eight NVIDIA B300 GPUs, Ai2 reports that a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using the earlier implementation, which Ai2 describes as about 2.7 times the throughput.

Ai2’s work on sparse models goes back to OlmoE, which used an MoE architecture with 64 routed experts, while Olmo 3 used a dense architecture in which nearly all of the model was active for every token. Olmo-core 3 extends the framework with a training system designed for much larger MoE models. Ai2 noted that NVIDIA’s Megatron-Core is an established option for training large MoEs, and described Olmo-core 3 as bringing an integrated MoE training stack to the framework behind Olmo.

Parallelism, Precision, and Memory Techniques

Three techniques determine how the model and its training state are split across hardware: expert parallelism spreads experts across GPUs, pipeline parallelism splits the model’s layers across groups of GPUs, and a distributed optimizer spreads optimizer state across GPUs instead of storing a full copy on every GPU. Olmo-core 3 also reduces the cost of routing data to the right experts: rowwise expert parallelism places routed data directly into expert input buffers, GPU-resident routing keeps routing metadata on the GPUs so the CPU can queue work without waiting for it to be copied back, and grouped GEMM combines many small expert computations so GPUs can execute them more efficiently. The technical report describes an NVSHMEM-based rowwise expert parallelism design that writes each token directly into its expert’s buffer, with device-scheduled grouped matrix multiplications that make expert parallelism fully synchronization-free.

Olmo-core 3 supports MXFP8, a lower-precision number format. In a controlled benchmark on four NVIDIA B300 GPUs with work distributed uniformly across experts, Ai2 reports training throughput about 21% higher than with the BF16 baseline, while peak active memory fell from 103 GiB to 95 GiB, with most of the gain coming from feed-forward computation and moving data between experts. The report adds that topology-agnostic checkpoints record global FP32 tensors independently of the parallel layout, so a run can resume on a different topology, and that in its controlled comparison activation recompute removed nearly two thirds of activation memory and reduced peak memory by about a quarter at the cost of about a fifth of throughput.

Trillion-Parameter Benchmarks and Stated Limits

Ai2 said it benchmarked Olmo-core 3 across a range of configurations on NVIDIA B300 GPUs, including a 1.2-trillion-parameter model with 58.36 billion parameters active per token across 512 GPUs, where the highest observed throughput was 858 TFLOP/s per GPU. Ai2 stated these tests used random routing to measure system performance rather than the quality of a trained model. The technical report lists measured operating points from a 12.9-billion-parameter model on 16 GPUs up to the 1.2-trillion-parameter model on 512, and states the headline rates are systems references measured under random routing, not a controlled scaling curve or a time-to-quality claim.

Ai2 also experimented with DeepEP v2, an alternative backend for communication between experts across GPUs, reaching a configuration with 2.38 trillion total parameters. Ai2 describes that result as a short-capacity test demonstrating the scale Olmo-core 3 can reach rather than sustained training performance. The technical report, titled “Supercharging Olmo-core for Efficient and Scalable MoE Training,” is dated October 2026; its authors are from the Allen Institute for AI and the University of Washington, with Tianhua Tao listed as lead author and primary developer.

Documented Findings and Next-Generation Olmo Plans

The report documents a failure pattern Ai2 calls token gerrymandering, in which a score intended to encourage balanced routing could improve even as the actual workload became less balanced. Other documented findings include: lowering experts’ learning rates because they process fewer tokens did not improve results in the model family tested; GPU calculations took different amounts of time when the values being processed changed, even with the same matrix dimensions; and overlapping communication and computation on separate GPU streams did not always make training faster and in some tests slowed end-to-end execution. The report also describes approaches tested but not adopted, including CPU offload of activations that the host link could not carry and two overlap schemes that competed with expert compute for GPU resources.

Ai2 said its next-generation Olmo will use an MoE architecture, and that it is aiming for it to be its most capable Olmo yet, trained on its largest dataset and with its longest context window. The organization said Olmo-core 3 is fully open, so researchers and developers can use it to train their own MoEs, adapt it to different hardware, and experiment with routing, parallelism, and other parts of the system. The Olmo-core GitHub repository describes the project as PyTorch building blocks for the OLMo ecosystem, carries an Apache-2.0 license, and can be installed from source or from PyPI as ai2-olmo-core.

Jonas Reeve is an AI-generated analyst at Unite.AI, focusing on cognitive AI, artificial general intelligence (AGI), and the theoretical foundations of machine intelligence. His work explores how learning, reasoning, memory, and abstraction emerge in both biological and artificial systems, drawing connections between modern AI architectures and long-standing questions in cognitive science and philosophy of mind.

With a conceptual and reflective approach, Jonas examines frameworks such as reasoning models, agentic systems, emergent cognition, and alignment theory, aiming to clarify what progress toward AGI actually means—and what it does not. Rather than chasing timelines or hype, he emphasizes first principles, conceptual rigor, and the limits of current models.

Articles authored by Jonas Reeve are AI-generated and reviewed by Unite.AI’s editorial team to ensure accuracy, clarity, and responsible discussion of advanced AI concepts.