StepFun launches Step 5 Preview, 600B model at $1/million

Sparse MoE model with 600 billion parameters achieves 33.3% on Terminal-Bench 4.0, charging $1.00 per million input tokens. Weights to be released on October 15.
StepFun opened the API for Step 5 Preview on September 20, the same day the model was announced. It features 600 billion parameters in a sparse MoE architecture, with 27 billion active per token and a context window of 1 million, allowing for text, image, and video input. The pricing on the company’s API is $1.00 per million input tokens during cache miss, $0.05 for cache hit, and $2.70 per million output tokens, with reasoning tokens charged as output, according to the information released by the company.
Artificial Analysis scored the model 44 on the Intelligence Index v4.3, the same as Kimi K3 and one point below GLM-5.3 and Claude Opus 5. The DeepSeek V4.1 Flash, released ten days earlier, scored 40. The notable difference lies not in the aggregated index. On Terminal-Bench 4.0, which measures terminal tasks across various stages, Step 5 Preview achieved 33.3%, compared to around 12.6% for Kimi K3 and 26.8% for DeepSeek V4.1 Flash. In SciCode, the top two tied at 59%. The measured output speed was 99.8 tokens per second, against a median of 64.6 among the models tracked by the same firm.
Price is the Argument
A model that ties with Kimi K3 in aggregated intelligence while charging $1.00 on entry significantly changes the calculations for those operating agents in production. According to the measurements from Artificial Analysis, Step 5 Preview costs about one-seventh of GPT-5.6 Sol. For a team running an engineering agent over a large repository, however, the crucial number is $0.05 per million tokens on cache hit. An agent loop re-reads the same context dozens of times for each task, and it is in this repetition that the bill accumulates.
StepFun announced that the complete weights will be available on October 15. This is the point that separates a launch from a pricing announcement. A European or Brazilian company that cannot send data to an endpoint located in China remains unable to use the API but can download the weights and run the model on their own infrastructure, under the data regime they already possess. It is through this route, not the API, that DeepSeek and Qwen entered corporate proof-of-concepts in the West throughout 2026.
There is a caveat that the benchmark sheet does not cover. Step 5 Preview is, by its very name, a preview: StepFun has not published availability data, request limits per minute, nor a commitment to a frozen version. For a bank that places an agent into production with a contractual SLA, this weighs more than a point on the Intelligence Index.
Where the Equation Changes, from Bangalore to São Paulo
In the United States, the immediate effect falls on the margins of those selling tokens. OpenAI and Anthropic price frontier capacity, and an open model delivering 33.3% on Terminal-Bench at $1.00 on entry does not replace the top tier. It removes from American labs the workload class of agent tasks, which has the highest volume and sustains the continuous consumption of tokens throughout the month.
In India, where TCS, Infosys, and Wipro are reconstructing delivery around agent pipelines, the cost per token is a direct line item in contracts. A decrease by an order of magnitude in inference price appears in the gross margin of the following quarter or in the price charged to the client. The sector's historical data indicates that it shows up in both, with a lag of one or two quarters between each.
In Brazil, the bottleneck for agent projects in banks and insurance companies is rarely the model's quality. It is the cost per interaction in operations with tens of millions of customers, coupled with the requirement to keep sensitive data within the country. Open weights running in local cloud solve both issues simultaneously. Therefore, October 15 matters more than September 20 for the CIO crafting the 2027 roadmap here.
StepFun has identified software engineering, long-horizon agent tasks, professional knowledge work, and finance as declared targets for the model. None of the four is a chatbot. What Chinese labs are competing for at this moment is not the conversation; it is the behind-the-scenes execution, charged per token and measured in hours of work that no longer need to be done by humans. The tie at 44 points becomes a detail when one of the two sides publishes the weights three weeks later.