Qwen 3.8 Max Scores 56 On Artificial Analysis Intelligence Index, Ahead Of All US Companies Except Anthropic And OpenAI

It’s not just Kimi K3 — another Chinese model is now knocking at the doors of the AI frontier.

Alibaba’s Qwen3.8 Max has landed a score of 56 on the Artificial Analysis Intelligence Index, the benchmark aggregator’s composite measure spanning nine evaluations including GDPval-AA, Terminal-Bench, SciCode and Humanity’s Last Exam. The score places Qwen3.8 Max level with Claude Opus 4.8 (max), and ahead of every model out of Google, Meta and xAI. Only Anthropic and OpenAI’s top-tier releases, along with Moonshot AI’s Kimi K3, currently sit above it.

Artificial Analysis had earlier published a score of 53 for the model, but said those runs were hit by intermittent issues on the endpoint being tested. The re-run, conducted on Alibaba’s public API, pushed the number up to 56.

The bigger surprise in the release is Alibaba’s decision to open source the weights. The company has said it will release them next week, a departure from its usual approach with the Max line, which has stayed closed while smaller Qwen models shipped openly. Once out, Qwen3.8 Max, at 2.4 trillion total parameters with 95 billion active per forward pass, would be roughly six times the size of Alibaba’s biggest open release to date, Qwen3.5 397B, and second in scale only to Kimi K3’s 2.8 trillion parameters.

Qwen 3.8 Max Benchmarks

Qwen3.8 Max’s jump is steep by any recent standard. It scores 10 points higher than its predecessor Qwen3.7 Max, which sat at 46. On GDPval-AA specifically, the model posted a 468 Elo gain over Qwen3.7 Max, landing at 1739 Elo. That’s ahead of Kimi K3 (1685), roughly tied with Claude Fable 5 (1743) and GPT-5.6 Sol max (1730), and behind only Claude Opus 5 max at 1852.

Much of that gain traces back to a change in how the model works rather than raw capability alone. Qwen3.8 Max is taking far more turns to complete agentic tasks, averaging 64 turns on GDPval-AA compared to 14 for Qwen3.7 Max. Input token usage on that same evaluation rose roughly 15x, and output tokens climbed 45% to 145 million. The model is essentially doing more work per task, and that extra effort is showing up directly in the score.

The improvements aren’t limited to agentic evaluations either. Terminal-Bench v2.1 is up 6 points, CritPt gains 7, SciCode adds 4, and HLE improves by 3. GPQA Diamond stayed flat. Two evaluations moved the other way. AA-LCR dropped 2 points, and AA-Omniscience fell 10, driven by the model’s hallucination rate climbing from 23% to 40%. Accuracy on that evaluation stayed roughly where it was, near 31%, which means Qwen3.8 Max is now attempting questions it previously would have declined to answer, and getting a lot of them wrong instead.

There’s also an odd result on 𝜏³-Bench Banking, where the model scored 42%, a 32-point jump from Qwen3.7 Max. Artificial Analysis flagged this as an outlier, noting it puts Qwen3.8 Max ahead of models that beat it comfortably everywhere else.

Among Chinese labs, Qwen3.8 Max now ranks second, ahead of GLM-5.2 max at 51 and behind Kimi K3 max at 57.

Qwen 3.8 Max Pricing

The model runs at $2.00 per million input tokens and $6.00 per million output tokens on Alibaba Cloud’s first-party API, with a cache hit rate of $0.25 per million tokens. That’s actually cheaper across the board than Qwen3.7 Max, which was priced at $2.50/$7.50 with a $0.50 cache hit rate.

But per-task cost tells a different story. Because the model takes so many more turns to complete agentic work, Qwen3.8 Max ends up costing $1.14 per Intelligence Index task, more than double Qwen3.7 Max’s $0.53. That puts it at roughly 1.3x the cost of Kimi K3 max ($0.86) and around 2x GLM-5.2 max ($0.57). The lower per-token pricing gets eaten up entirely by the sheer volume of tokens the model now consumes per task.

For context, Claude Opus 5 max sits at $2.03 per task, GPT-5.6 Sol max at $1.23, and Claude Fable 5 with fallback tops the chart at $3.15. Qwen3.8 Max lands in the middle of that pack on cost while trailing most of them on intelligence, which puts it in an odd spot on the cost-versus-capability curve, expensive for a model in its performance bracket, but still cheap relative to the absolute frontier.

Once the weights are out next week, that pricing math changes entirely for anyone running the model on their own infrastructure or through a third-party host, which is likely to matter more than Alibaba’s own API pricing.

Is China Catching Up?

Line up the recent releases and a pattern is hard to miss. Kimi K3 beat Claude Fable 5 and GPT-5.6 Sol outright on Arena’s Frontend Code leaderboard. GLM-5.2 landed within striking distance of Claude Opus 4.8 on FrontierSWE. Now Qwen3.8 Max is matching Claude Opus 4.8 on the Intelligence Index while undercutting it on price, and doing so as an open-weight release rather than a closed one.

What makes this moment different from earlier rounds of “China catching up” narratives is the combination of factors involved. These aren’t just capable models, they’re also cheap, and increasingly, they’re models anyone can download and run themselves rather than rent through an API. Google, Meta and xAI, three of the best-funded AI labs in the US, are currently being outscored by a lab that only entered the frontier conversation a few releases ago. Meta’s Muse Spark 1.2 sits at 54 on the Index, and Grok 4.5 at the same mark, both below where Qwen3.8 Max now sits.

The gap that remains is concentrated at the very top, with Anthropic and OpenAI still holding a slight lead over everyone else, Chinese or American. But the layer just below that top tier is filling up fast with labs out of Hangzhou, Beijing and Shanghai. Alibaba’s own framing of this release leans into that gap closing, and its decision to open the weights rather than keep them proprietary suggests confidence that the model can hold its own even once anyone can pick it apart.

The trend line over the past year has been fairly consistent. Each new Chinese release arrives closer to the frontier than the one before it, and increasingly, those releases come with either open weights or pricing aggressive enough to make the closed-source alternative a harder sell. Whether that trend continues at the same pace is the real question for the rest of 2026, but for now, the distance between “American frontier lab” and “Chinese frontier lab” keeps shrinking, one release at a time.

Posted in AI