StepFun has introduced Step 5 Preview, a sparse mixture-of-experts model that scales to 600 billion total parameters while activating only 27 billion per token. This design keeps inference costs lower than a dense model of similar size, according to the announcement. The model also features a 1M-token context window, which is unusually large for a model of this class.
Beyond its parameter count, Step 5 Preview is multimodal, accepting text, image, and video inputs. The company positions the model for long-horizon agentic work—tasks that require sustained reasoning and multiple steps—with software engineering as a primary use case. The combination of a large context and multimodal input suggests an intent to handle complex, real-world coding and debugging scenarios.
No benchmark numbers or comparisons to other models were provided in the source, so the practical performance of Step 5 Preview remains unverified. The release is a preview, meaning StepFun may refine the model before a full launch. For now, the key differentiators are the sparse MoE architecture, the 1M-token context, and the explicit focus on agentic software engineering workflows.