Reka has introduced Rho-1, a 19-billion-parameter omni-reasoning model trained from scratch. The model is designed to handle multiple modalities within a single network, covering text, images, video, and robot actions.
Rho-1 processes these inputs and outputs over a shared KV cache, which allows the model to reason across modalities without separate components. According to Reka, the model can both read and generate content in all these forms.
A distilled version of Rho-1 can produce a 5.3-second video clip in roughly one second. Reka is positioning Rho-1 as a research preview, meaning it is not yet a fully supported product.