Sparse mixture-of-experts (MoE) models offer a way to scale neural networks without activating all parameters for every token. Instead, a router sends each token to a subset of expert networks. The new preprint introduces RouterInterp, a method for interpreting this routing behaviour, specifically the phenomenon of "superposed specialisation" — where an expert may hold multiple distinct specialisations.

The abstract notes that a leading hypothesis for MoE performance is related to this routing, but the text cuts off before stating the hypothesis in full. Based on the title, the paper appears to investigate how router decisions reflect overlapping or superposed expert roles.

As an arXiv announcement, the work has not yet been peer-reviewed. The provided abstract is partial, so the specific methods and results are not available from the source.