OpenAI seems to have retaken the AI crown from Anthropic, if Astra’s leaked benchmarks are to be believed.
OpenAI officially began rolling out GPT-6 Astra on Thursday, and a benchmark comparison table circulating alongside the launch shows the model pulling ahead of Anthropic’s Claude Fable 5.1 on a majority of evaluations, including several where Anthropic’s own model had only recently set records. OpenAI co-founder and president Greg Brockman went as far as to call Astra a “generational leap” during a press briefing, adding that he personally believes the company may have reached AGI with this release.
The two models were launched just days apart, making for one of the more direct head-to-head comparisons the AI industry has seen this year, and Astra appears to have the edge on paper.
OpenAI GPT-6 Astra Benchmarks

On FrontierMath Tier 4 (v2), a notoriously difficult mathematics benchmark, Astra scored 97.6%, comfortably ahead of Fable 5.1’s 87.8% and OpenAI’s own previous model, GPT-5.6 Sol, at 83.0%. The gap holds on GPQA Diamond too, where Astra hit 96.0% against Fable 5.1’s 93.7%.
Astra’s biggest jumps show up on agentic and infrastructure-heavy benchmarks. On AutomationBench, a test of business workflow automation, Astra scored 41.4% compared to Fable 5.1’s 31.4%. On Terminal-Bench Science 0.1, an agentic scientific research benchmark that Anthropic had specifically highlighted as a strength for Fable 5.1 (more than doubling Fable 5’s score), Astra still came out ahead at 64.6% versus Fable 5.1’s 52.6%. Astra also edged out Fable 5.1 on BenchCAD (95.9% vs 84.3%) and DeepSWE v1.1 (74.1% vs 67.4%), though Google’s Gemini 3.8 Flash actually posted a marginally higher DeepSWE score than Fable 5.1 at 73.7%.
On the ARC-AGI-3 benchmark, often cited in AGI debates, OpenAI reported a 98.6% score for Astra. Anthropic hasn’t published an ARC-AGI-3 number for either Fable 5.1 or Fable 5, though Claude Opus 5 scored 30.2% at high reasoning effort. It’s worth noting OpenAI’s own evaluation notes flag that Astra’s result uses its Responses API harness, and comparisons across different agent harnesses aren’t strictly apples to apples — a point underlined by Nvidia’s AVO system separately hitting 100% on ARC-AGI-3’s public set using Claude Opus 5 as its underlying model, with scaffolding doing much of the heavy lifting.
Where Fable 5.1 Doesn’t Show Up At All
Several rows in the table simply have no Fable 5.1 score, including Agents’ Last Exam, GeneBench Pro, MedChemBench, and ExploitBench, areas Anthropic hasn’t published comparable numbers for. On HealthBench Professional (length-adjusted), interestingly, Fable 5.1 actually trails its own predecessor Fable 5 (56.6% vs 60.9%), and both fall behind Astra’s 63.4%.
Astra’s cybersecurity numbers stand out in particular. OpenAI designated Astra as the first model to cross its “Critical” cybersecurity capability threshold under its Preparedness Framework, reporting a perfect 100% on ExploitBench and a 99.2% score on SRE-Bench across four attempts. The company also highlighted an internal alignment test in which Astra never attempted to circumvent its own auto-review safety checks, compared to a 0.29% attempt rate for GPT-5.6 Sol.
Context: Fable 5.1 Was Already A Big Jump For Anthropic
None of this diminishes what Fable 5.1 did on release. Anthropic’s model had already topped the Artificial Analysis Intelligence Index with a score of 66, beating its own Claude Opus 5 by 3 points, and more than doubled Fable 5’s score on Terminal-Bench Science 0.1. That capability jump came at a real cost, though — Fable 5.1 is currently the most expensive model on the Artificial Analysis leaderboard, running 57% pricier per task than Opus 5 at max effort.
On pricing, the two frontier models are now roughly matched. Astra’s standard API pricing is $10 per million input tokens and $50 per million output tokens, identical to what Anthropic charges for Fable 5.1 and Mythos 5.1. OpenAI is also pushing enterprises to think in terms of “price per completed task” rather than per-token cost, pointing to DeepSWE v1.1 results where Astra’s best configuration reportedly beats GPT-5.6 Sol’s best score at roughly 57% lower estimated cost per task.
What This Means
It’s worth flagging, as always with day-one benchmark tables, that these are OpenAI’s own reported numbers rather than independently reproduced results, so they’re best read as directionally informative rather than gospel. Anthropic has not yet responded publicly to the comparison. Astra is rolling out first to enterprise customers in OpenAI’s Daybreak program, with availability for ChatGPT Plus, Pro, Business, and Enterprise users, plus the OpenAI API and cloud platforms including AWS Bedrock and Azure, expected “in the coming days.”