Prime Intellect has launched Prime Inference, an OpenAI-compatible platform for serving frontier open models on NVIDIA Blackwell. The service offers both serverless and reserved serving options, giving users a choice between on-demand scaling and pre-provisioned capacity.

The initial GLM-5.3 deployment is built on Dynamo, vLLM, and NVFP4 KV compression. According to the announcement, this setup serves 66 sessions per prefill group and delivers 101 tokens per second per user.

With only one source, there are no independent figures to corroborate or contradict the performance claims. The reported numbers are vendor-provided and should be read as such.