LiquidAI has released two open decision models on Hugging Face: d1-3B and d1-omni-600M (experimental). Unlike generative models that produce tokens, these decision models answer in a single forward pass. d1-3B is built on the LFM2.5-VL-3B vision-language backbone and accepts text and images; d1-omni-600M, built on a bidirectional encoder with vision and audio encoders, handles text plus image or text plus audio.

On the Decision Index 0.2.1, d1-3B scores 48.57, making it the best decision model under 10B parameters and ahead of Decider 35B-A3B (47.11). Across seven public benchmarks, d1-3B averages 82.9, while d1-omni-600M scores 78.4, surpassing Decider 2B (77.1) with a quarter of the parameters. The authors note that vision and audio benchmarks are not reported because the Decision Index v0.3 includes only a private vision split and audio decision benchmarks remain an open problem.

Working with NVIDIA, the team measured d1-3B inference speed on edge hardware: 16 ms on a Jetson AGX Thor, 26 ms on a Jetson AGX Orin, and 50 ms on a Jetson Orin Nano. On an RTX 4090, a single question takes 8 ms. No speed figures are given for d1-omni-600M, as it is an early research release. Both models are open-weight and available on Hugging Face, with usage examples in the System One Arcade Space.