DeepSeek V4 Flash Outperforms Fable 5 On Terminal Bench While Being 99% Cheaper

DeepSeek V4 Flash appears to be causing the same price disruption that DeepSeek R1 had done a year ago.

New numbers from Terminal-Bench 2.1, a benchmark that tests how well models handle real terminal and command-line tasks, show DeepSeek V4 Flash scoring 82.7, ahead of Anthropic’s Claude Fable 5 at 80.5. GPT-5.6 Sol still tops the chart at 85.8, and Claude Sonnet 5 trails the group at 74.5. What makes the DeepSeek result stand out isn’t the score itself, it’s what sits next to it. V4 Flash is priced at $0.14 per million input tokens and $0.28 per million output tokens. Fable 5, Anthropic’s Mythos-class flagship, charges $10 and $50 for the same. That’s a model beating a frontier lab’s most expensive offering while costing roughly 99% less to run.

The gap gets starker when you move from cost-per-token to cost-per-task. Artificial Analysis’s Intelligence Index tracks average dollar cost per task rather than raw token pricing, since reasoning models can burn through wildly different amounts of tokens to arrive at an answer. On that measure, DeepSeek V4 Flash comes in at $0.03 per task. Kimi K3 sits at $0.86. GPT-5.6 Sol is at $1.86. Opus 5 comes in at $2.34. Fable 5 sits at the top of the cost chart at $3.15 per task, making DeepSeek roughly 105 times cheaper on a like-for-like basis.

A familiar pattern from DeepSeek

This isn’t the first time DeepSeek has pulled this move. When R1 launched in January 2025, it matched OpenAI’s o1 on reasoning benchmarks while charging a fraction of what o1’s API cost, and the release briefly wiped hundreds of billions off Nvidia’s market cap as investors reassessed how much compute frontier AI actually needed. The company followed that up through 2025 and into 2026 with a steady drip of releases, including DeepSeek Math V2, which won IMO gold as an open model while OpenAI and Google kept their equivalents closed, and V3.2-Speciale, which the company pitched as a rival to Gemini 3 Pro on reasoning tasks.

V4 Flash fits the same playbook. It isn’t trying to be the smartest model on the leaderboard. GPT-5.6 Sol still edges it out on Terminal-Bench, and Opus 5 and other frontier systems will likely stay ahead on the hardest reasoning and agentic benchmarks. What DeepSeek is doing instead is making the price-to-performance conversation uncomfortable for everyone else, and it’s doing it with weights that anyone can download, inspect, and run themselves.

Why the open-weight angle matters more this time

DeepSeek has released its models under permissive licenses since V3 and R1, letting developers fine-tune and self-host rather than route every query through a metered API. That was a novelty in early 2025. By mid-2026, it’s a genuine competitive lever. A lab or startup that doesn’t want to pay Anthropic or OpenAI’s per-token rates, or that has data residency concerns about sending traffic to a US frontier lab, now has an open alternative that beats a $10/$50 model on a real coding benchmark.

That puts pressure on Anthropic and OpenAI from two directions at once. On price, neither company has shown much appetite to compete anywhere near DeepSeek’s range. Anthropic priced Fable 5 access at $10 and $50 per million tokens when it moved off the subscription-included allowance, among the highest rates the company has ever charged for a public model, and it has repeatedly pushed back the date that promotional access runs out, a sign that the commercial pricing isn’t an easy sell even to its own subscriber base. OpenAI has leaned the opposite way, and GPT-5.6 Sol has been marketed heavily on cost-efficiency, with Sam Altman publicly calling out Fable 5’s pricing on X after Sol landed within a point of it on the Intelligence Index. DeepSeek’s task cost undercuts both of them by an order of magnitude or more.

On openness, the contrast is sharper still. Anthropic has spent the past two months fighting a US government export control order that pulled Fable 5 and Mythos 5 offline worldwide for foreign users, a fight that, according to David Sacks, started because Anthropic refused to patch a jailbreak or pull the model when asked. Whatever the merits of that standoff, it left Fable 5 unavailable to a meaningful chunk of the world for nineteen days, and it underlined how much control a government can exert over a closed, hosted model. DeepSeek’s weights don’t carry that risk for downstream users. Once they’re published, a directive against the company in Hangzhou doesn’t stop a developer in Berlin or Bangalore from running the model on their own hardware.

The bigger picture

Terminal-Bench 2.1 is a narrow benchmark, and one result doesn’t rewrite the hierarchy of frontier AI. Fable 5 and Opus 5 still lead on some of the harder agentic and reasoning tests, and Anthropic’s own numbers show Opus 5 tripling the next-best model’s score on ARC-AGI-3, a benchmark built specifically to resist pattern-matching. But Terminal-Bench measures something enterprises actually care about, running real commands and completing real developer workflows, and DeepSeek just showed it can match or beat a $10/$50 model there for a fraction of a cent per token.

For Anthropic and OpenAI, the lesson from R1 was supposed to have already landed: efficiency gains from Chinese labs aren’t a one-off. V4 Flash is a reminder that they keep coming, and that the gap between “close enough on capability” and “orders of magnitude cheaper” is exactly the kind of gap that changes buying decisions long before it changes leaderboard rankings.

Posted in AI