Ai2 Olmo-core 3 MoE training stack
Ai2 released an open MoE training-stack upgrade and reported scaling benchmarks, while its next Olmo model remains in development.
What happened lately
Ai2 Olmo-core 3 MoE training stack timeline
An open stack for larger sparse models
Ai2 released Olmo-core 3 on October 1, describing a redesign of its open training framework for mixture-of-experts models. The framework keeps expert weights on GPUs and routes data to them instead of repeatedly gathering and redistributing those weights. Ai2 also describes expert and pipeline parallelism, a distributed optimizer and lower-precision computation. Its code is available in the public Olmo-core repository.
Benchmarks are not a model release
Ai2 reports that a preliminary test of a 47-billion-parameter model on eight NVIDIA B300 GPUs processed about 2.7 times as many tokens per second per GPU as its earlier implementation. It also tested a 1.2-trillion-parameter configuration across 512 GPUs using random routing to measure system throughput. Those are Ai2’s reported engineering benchmarks, not independent measurements of model quality or evidence that a trillion-parameter Olmo model has been fully trained. Ai2 says its next Olmo generation will use this MoE architecture; it has not announced that model as released.