Gimlet Labs and Cerebras Systems have announced a collaboration aimed at ultrafast AI inference. The deployment combines Cerebras wafer-scale compute with Gimlet Cloud, the company's inference platform, and is expected to deliver up to 3,000 tokens per second.

The announcement, reported by HPCwire, comes from San Francisco and Sunnyvale, California. Gimlet Cloud is the service layer in the arrangement, while Cerebras supplies the wafer-scale hardware. The companies describe the result as a new class of high-speed AI inference.

Because this is a single vendor announcement, the 3,000-token figure is a stated maximum rather than an independently measured result. No third-party benchmarks or comparisons with other inference systems are included in the source.