NVIDIA has introduced Nemotron 3 Diarization, a model designed for real-time speaker diarization in multi-speaker audio. Diarization is the process of identifying which speaker is talking at each point in a recording, and doing so in real time makes it suitable for live applications.

The model is available via Hugging Face, where the announcement is posted. According to the announcement, it is positioned as a way to build real-time, multi-speaker AI systems, which could be useful for live meeting notes, captions, and similar tools.

This article is based solely on the Hugging Face post. The model comes from NVIDIA and is intended for developers who need to track speakers during live conversations, with access provided through the Hugging Face platform.