Ai2 has released Olmo-core 3, an upgrade to its open framework for training large language models that targets mixture-of-experts (MoE) architectures. The new system is designed to scale into the trillion-parameter range while avoiding the communication and coordination overhead that often erodes MoE efficiency. The announcement positions it as part of Ai2's effort to make advanced model development more accessible to academic researchers and smaller labs.
The core change is a move from fully sharded data parallelism (FSDP) to a distributed data parallelism (DDP) approach. Instead of gathering and resharding model weights for each small batch, Olmo-core 3 keeps experts resident on GPUs and routes the relevant data to them. It combines expert parallelism, pipeline parallelism, and a distributed optimizer to split the model and training state across hardware. Additional optimizations include rowwise expert parallelism, GPU-resident routing, grouped GEMM, and support for MXFP8 lower-precision arithmetic.
In one benchmark, increasing the expert pool from 8 to 128 while keeping four active experts per token raised total parameters from 4.6B to 47B, with training throughput falling less than 5%. A separate test on eight NVIDIA B300 GPUs showed a 47B-parameter MoE processing 52,000 tokens per second per GPU, about 2.7 times the throughput of the earlier implementation. With MXFP8 enabled, end-to-end training throughput was roughly 21% higher than with BF16, and peak active memory dropped from 103 GiB to 95 GiB.
The source notes that these techniques involve trade-offs: speeding up one part of training can increase data movement, and lower-precision formats only help when conversion costs are outweighed by savings. Olmo-core 3 is built to give researchers control over those trade-offs across the full training process, with code and an interactive walkthrough released alongside the technical report.