Back to News
A

Ai2 Olmo-core 3 MoE training stack

News Oct 1, 2026

Ai2 released an open MoE training-stack upgrade and reported scaling benchmarks, while its next Olmo model remains in development.

2 sources attached

What happened lately

Ai2 Olmo-core 3 MoE training stack timeline

An open stack for larger sparse models

Ai2 released Olmo-core 3 on October 1, describing a redesign of its open training framework for mixture-of-experts models. The framework keeps expert weights on GPUs and routes data to them instead of repeatedly gathering and redistributing those weights. Ai2 also describes expert and pipeline parallelism, a distributed optimizer and lower-precision computation. Its code is available in the public Olmo-core repository.

Benchmarks are not a model release

Ai2 reports that a preliminary test of a 47-billion-parameter model on eight NVIDIA B300 GPUs processed about 2.7 times as many tokens per second per GPU as its earlier implementation. It also tested a 1.2-trillion-parameter configuration across 512 GPUs using random routing to measure system throughput. Those are Ai2’s reported engineering benchmarks, not independent measurements of model quality or evidence that a trillion-parameter Olmo model has been fully trained. Ai2 says its next Olmo generation will use this MoE architecture; it has not announced that model as released.