Ai2 releases Olmo-core 3 for open mixture-of-experts training
The open framework adds a redesigned MoE training stack; Ai2's trillion-parameter figures describe system tests, not a trained model launch.
What Ai2 released
Ai2 announced Olmo-core 3 on October 1 as an open training framework upgrade for large mixture-of-experts (MoE) language models. MoE systems route each token through only some of a model’s specialized components, but still have to store and coordinate the full set of experts during training. Ai2 says its redesigned stack keeps expert weights resident on GPUs and moves the relevant data to them, replacing a more costly weight-gathering approach in its earlier MoE implementation. The public Olmo-core repository provides the code; this announcement is about training infrastructure, not a new finished Olmo model.
Ai2 reports a preliminary comparison on eight NVIDIA B300 GPUs: a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, versus 19,400 with the earlier one. The company also describes a 1.2-trillion-parameter configuration tested across 512 GPUs. That larger benchmark used random routing to measure systems performance rather than the quality of a trained model. A separate 2.38-trillion-parameter configuration was a short-capacity test, not a sustained training run.
Why the distinction matters
The release gives researchers access to more of the machinery behind Ai2’s future open models, including expert parallelism, pipeline parallelism and a distributed optimizer. Ai2 says the next Olmo generation will use an MoE architecture. Its performance numbers are company-reported benchmarks under specific hardware and test conditions; they do not establish that the future model is available or that the stack will deliver the same gains on every cluster.