DeepSeek had first burst on to the scene in late 2024, and it’s still coming up with releases that — at least from a pricing perspective — are leaving the competition in the dust.
Artificial Analysis has published its independent evaluation of DeepSeek V4 Flash 0731, and the model scores 50 on the Artificial Analysis Intelligence Index, a jump of 10 points over the original DeepSeek V4 Flash from April and 6 points clear of DeepSeek V4 Pro. What makes the result stand out isn’t the score in isolation, it’s where that score lands on Artificial Analysis’s cost chart. V4 Flash 0731 shares identical architecture and pricing with its predecessor, so the entire 10-point gain shows up as a near-vertical jump on the Pareto frontier, with the new model sitting almost directly above the old one at roughly the same cost per task.

The model remains a 284 billion parameter mixture-of-experts system with 13 billion active parameters at inference, unchanged from the original V4 Flash, and keeps the same 1 million token context window. Pricing hasn’t moved either, at $0.14 per million input tokens and $0.28 per million output tokens, with a cache hit price of $0.0028 per million tokens, a 98% discount that Artificial Analysis notes is considerably steeper than the roughly 90% cache discount most of the industry offers.
That cache pricing turns out to matter a lot against the week’s other big story. OpenAI cut the price of GPT-5.6 Luna by 80% a day before this data came out, and Luna at max effort still edges V4 Flash 0731 by a single point on the Intelligence Index, 51 to 50. But even after that cut, Artificial Analysis puts V4 Flash 0731’s cost per task at roughly 60% lower than Luna’s for essentially the same intelligence, almost entirely because of how aggressively DeepSeek discounts cached tokens on its first-party API. For agentic workloads that repeat large chunks of context across many calls, that gap compounds fast.
The new model sits within a point of Z.AI’s GLM-5.2 at 51 and Meta’s Muse Spark 1.1 at 51, and lines up almost exactly with Google’s Gemini 3.6 Flash at 50. It trails Kimi K3 at 57, still the open-weights leader, by 7 points. DeepSeek hasn’t released the full weights for 0731 yet, but Artificial Analysis expects that to happen in the coming weeks, at which point the model would slot in as the second-highest scoring open-weights system on GDPval-AA v2, behind only Kimi K3’s 1687 Elo and ahead of GLM-5.2’s 1510.
The agentic numbers are where the upgrade shows up hardest. GDPval-AA v2, Artificial Analysis’s Elo-based benchmark for real-world work tasks, jumps from 1189 for the original V4 Flash to 1559 for 0731, a 370-point swing on a scale anchored to a human baseline of 1,000. Terminal-Bench 2.1 rises 17 points to 79%, and τ³-Bench Banking climbs 8 points to 31%. Every single evaluation in the Intelligence Index improved over the predecessor, including CritPt, up 9 points to 17%, SciCode, up 5 to 50%, and Humanity’s Last Exam, up 5 to 37%. None of this came from throwing more compute at the problem either — the model used about 206 million output tokens to run the full Index, down 12% from the 234 million the original V4 Flash needed.

Reliability improved too, though the source of that improvement is worth separating out. DeepSeek V4 Flash has been criticized for its hallucination rate in the past, and V4 Flash 0731 brings the AA-Omniscience Index up from -23 to -16. That gain, however, comes entirely from the model refusing to answer more often rather than getting more answers right — accuracy holds flat at 37% while the hallucination rate drops to 84%, putting it roughly in line with GPT-5.6 Terra’s 85% and Mistral Medium 3.5’s 82%. It’s a meaningful shift in behavior, just not the kind that shows up as the model knowing more.

Anthropic’s Claude Opus 5 still tops the overall board at 61, with Fable 5 and GPT-5.6 Sol filling out the rest of the top three. DeepSeek isn’t competing for that top spot with this release, and it isn’t really trying to. What it’s demonstrating instead is that a mid-tier open model, refreshed without a single architectural change, can close most of a 10-point gap through training alone, and do it while staying attached to the same rock-bottom API pricing that’s made DeepSeek’s models a fixture in high-volume production pipelines since R1.