Alibaba's Qwen team has released Qwen3.8-LiveTranslate, a real-time simultaneous interpretation model. According to MarkTechPost, the model is built on a new Interleave architecture and cuts average lag from 2.8 seconds to 2.3 seconds.
The model covers 60 languages and introduces real-time speaker diarization with stable voice cloning, along with synchronized output. This points to an emphasis on preserving speaker identity and timing during live interpretation.
The lag reduction is modest but meaningful for real-time use cases. The source does not provide details on benchmark methodology or comparisons with other interpretation systems.