A new arXiv preprint, DivMoE, targets the high cost of building fine-grained Mixture-of-Experts (MoE) models. While fine-grained expert designs are known to improve scaling, training such models from scratch is expensive. The paper proposes an 'upcycling' approach that reuses existing dense models instead.
The method, called DivMoE, composes experts across domains to create a fine-grained MoE structure. This cross-domain composition is the key mechanism for turning a dense checkpoint into a sparse MoE model without starting from zero. The abstract is truncated, so specific experimental results are not available in the source.
Overall, DivMoE adds to a line of work making MoE more accessible by reducing the computational barrier to entry.