The Allen Institute for AI (AllenAI) has introduced Olmo-core 3, an open-source training infrastructure built to support the development of exceptionally large Mixture-of-Experts (MoE) language models. The framework is engineered to scale training processes into the trillion-parameter range, addressing a key challenge in the efficient development of these complex AI systems. Olmo-core 3 represents a significant advancement over its predecessor, Olmo-core 2, by shifting its training approach from fully sharded data parallelism (FSDP) to a distributed data parallelism (DDP) system.

This architectural change allows experts, which are specialized components within an MoE model, to remain resident on GPUs. Instead of gathering and resharding model weights for each training data batch, Olmo-core 3 routes relevant data directly to these resident experts. AllenAI stated this method avoids repeated weight gathering, a process that can diminish the computational advantages of MoE models as they grow in size.

In benchmarks conducted on NVIDIA B300 GPUs, Olmo-core 3 demonstrated increased throughput. A preliminary test on eight GPUs showed a 47-billion-parameter MoE model processed 52,000 tokens per second per GPU with the new Olmo-core 3 stack, compared to 19,400 tokens per second per GPU with the earlier implementation. This represents approximately a 2.7 times improvement in training throughput. The infrastructure has also been benchmarked with models exceeding one trillion total parameters, including a 1.2-trillion-parameter model utilizing 58.36 billion active parameters per token across 512 GPUs, achieving a throughput of 858 TFLOP/s per GPU.

Olmo-core 3 incorporates several techniques to enhance efficiency. Rowwise expert parallelism places routed data directly into expert input buffers, minimizing data rearrangement. GPU-resident routing keeps routing metadata on the GPUs, allowing the CPU to queue work without waiting for data transfer. Additionally, grouped GEMM combines multiple small expert computations for more efficient GPU execution.

The framework is designed to be a core component for AllenAI's next generation of Olmo models. These future iterations are expected to feature MoE architectures, larger datasets, and extended context windows, aiming to be the most capable Olmo models to date. AllenAI emphasizes that Olmo-core 3 is fully open, enabling researchers and developers to train their own MoEs, adapt the system to different hardware, and experiment with various aspects of MoE training, such as routing and parallelism.

The release of Olmo-core 3 aligns with AllenAI's commitment to open research and development in artificial intelligence. By providing transparent access to the tools and infrastructure behind advanced model development, the institute aims to democratize access to large-scale AI training for academic researchers and smaller laboratories.