Cross-lingual contrastive learning has long been a standard technique for training multilingual encoders, but it has not translated cleanly to decoder-only large language models. The reason, according to a new paper on arXiv, is that varying multilingual tokenization prevents explicit alignment of representations across languages.
The paper, identified as arXiv:2610.01921, proposes using mixture-of-experts (MoE) routers to perform this alignment. Rather than relying on token-level contrastive objectives, the authors suggest that routers—typically used to route tokens to specialized expert modules—can be repurposed to bring representations from different languages into a shared space.
The abstract is brief and does not include experimental results or implementation specifics. Still, the proposal addresses a real limitation in current multilingual LLM training, and the use of MoE routers as an alignment mechanism is a novel direction worth watching.