The number of Chinese AI labs that are competing close to the frontier continues to grow.
Chinese AI lab StepFun has released Step 5 Preview, a new flagship model built for agentic work. On the Artificial Analysis Intelligence Index, it beats Google’s Gemini 3.8 Flash on both intelligence and cost. Step 5 Preview scores 44 to Gemini 3.8 Flash (high)’s 41. It also costs $0.72 per Intelligence Index task against $1.24 for Google’s model, about 42% less.

StepFun is pitching the model as a move along the Pareto frontier, meaning comparable intelligence at a much lower price. It uses a sparse Mixture-of-Experts architecture with 600 billion total parameters and 27 billion activated per token, and supports a 1 million-token context window. StepFun’s documentation says it takes text, image and video input. The company is targeting software engineering, professional knowledge work and finance. API access opened on September 20, and StepFun says the weights will be released on October 15, 2026.
API pricing is $1.00 per million input tokens (or $0.05 on a cache hit) and $2.70 per million output tokens, including reasoning.
How it stacks up
A score of 44 puts Step 5 Preview level with Kimi K3 and Grok 4.6 (high), and one point behind GLM-5.3. It is ahead of DeepSeek V4.1 Flash (39), GPT-5.6 Luna (37) and DeepSeek V4 Pro 0813 (36). The top of the chart is still held by Claude Fable 5.1 and GPT-6 Astra, which are tied at 53, followed by Claude Opus 5 at 51 and Muse Spark 1.3 at 48.

The cost gap to the frontier is large. Fable 5.1 costs $7.63 per task, Opus 5 costs $5.86, and GPT-6 Astra costs $3.26. Step 5 Preview delivers about 83% of Fable 5.1’s score at roughly a tenth of the price. A few models are cheaper, but they score lower. GPT-5.6 Luna costs $0.18 per task, DeepSeek V4.1 Flash costs $0.27, and DeepSeek V4 Pro 0813 costs $0.67.

There are trade-offs. Gemini 3.8 Flash is far faster, at 331 output tokens per second against 100 for Step 5 Preview. StepFun’s model is still quicker than GPT-6 Astra (71), Claude Fable 5.1 (70), Claude Opus 5 (62) and Kimi K3 (44). It is also verbose. It generated 160 million output tokens on the index run against a median of 92 million, which eats into some of the per-token savings.
It’s also worth noting that Gemini 3.8 Flash had scored 59 on an earlier version of the index. The newer version leans on a harder Terminal-Bench 4.0, where Semi Analysis has argued Google’s model drops sharply compared to older benchmarks.
StepFun’s own numbers
StepFun’s self-reported results are more modest. Step 5 Preview ran at high effort while its rivals ran at max. It scored 66.4 on FrontierFinance against 69.7 for Claude Opus 5, and 83.3 on DRACO against 87.6. GPT-6 Astra scored 55 and 76.8 on the same two tests. On coding, it posted 67.7 on DeepSWE v1.1, 49.0 on StepCodeBench and 80.5 on ProgramBench, and GPT-6 Astra and Claude Opus 5 stay ahead on all three.
The company also ran two 24-hour agent experiments. In one, the model tuned an H100 kernel to 508 TFLOPS against 493 for Claude Opus 5. In the other, it lifted Qwen3-30B-A3B’s AIME24 score from 53.3% to 60% through automated post-training. StepFun says that in a single agent action on a research task, the model coordinated 950 web fetches. On the architecture side, the model uses a 92-layer narrow-deep Transformer, which StepFun argues helps with multi-hop reasoning during long prefill.
About StepFun
StepFun (阶跃星辰) is one of China’s “AI Tigers”, the group of well-funded startups racing to build homegrown frontier models. It was founded on April 6, 2023 in Shanghai. Its listed founders are Jiang Daxin, Zhu Yibo and Jiao Binxing, all former Microsoft employees. Jiang spent 16 years at Microsoft before leaving, and ChatGPT’s release convinced him he could build something comparable. He was involved in products such as Bing and Microsoft 365. CTO Zhu is a former Microsoft and ByteDance AI veteran. Chief scientist Zhang Xiangyu is a core author of ResNet. In January 2026, Megvii co-founder Yin Qi joined as chairman.
StepFun’s models have grown steadily in scale. Its Step-2, launched in July 2024, was the first trillion-parameter MoE language model from a Chinese startup. Step 3.5 Flash, its most capable open-source model at the time, activates 11 billion of 196 billion parameters per token. By one count, StepFun has released 38 foundation models, 31 of them multimodal.
The funding has scaled up with it. In January 2026 it closed a Series B+ of more than RMB 5 billion (about $717 million). Investors include Tencent, Qiming Venture Partners and 5Y Capital. It then reportedly completed a further $2.5 billion round, with new backers including Huaqin Technology, OmniVision and ZTE, and Hong Kong’s sovereign-backed HKIC on the shareholder list. The company hasn’t officially confirmed the round. The Wall Street Journal has reported that StepFun could raise as much as $500 million in a Hong Kong IPO at a valuation of up to $12 billion. It has already converted from a limited liability company to a joint stock company and dismantled its red-chip structure, steps usually taken ahead of a Hong Kong listing. It would follow Z.ai and MiniMax, which listed in January.
StepFun’s strategy leans on putting models into hardware. Reports say its models are on more than 42 million devices and it has deals with about 60% of China’s leading smartphone brands. It reportedly made around RMB 500 million in revenue in 2025 and is targeting RMB 1 billion in 2026. Jiang has said the team numbers over 500, nearly 80% of them algorithm and technical staff.
For now Step 5 Preview is API-only, but the October 15 weights release will show whether it can hold this cost-to-intelligence position when developers can run it themselves. Developers can try the model on the StepFun Open Platform.