Mixture-of-experts (MoE) models are an established way to scale neural networks: they add parameters while keeping the active computation per token roughly constant. But whether that advantage carries over to particle-physics transformers has been an open question. A new arXiv preprint, posted as arXiv:2610.02701, directly investigates this trade-off.

The authors study both dense and MoE particle transformers, aiming to see how conditional capacity and routing behave in this domain. The abstract notes that MoE models can increase parameter capacity without proportionally increasing active computation, but it flags that the trade-off in particle-physics transformers is unclear.

As released, the abstract is truncated, so the specific findings are not fully visible. The paper's motivation, however, is clear: determining whether MoE routing helps or hurts when processing particle-physics data, where the structure differs from typical language or vision inputs.