OpenAI fumbled with the public release of its GPT-6 model — an official blog was discovered on its website by users, before it was quickly taken down. OpenAI then shared the link itself, but the link wouldn’t work. At long last, the blogpsot on GPT-6 Astra is accessible, and it has the model’s much-awaited benchmarks.
OpenAI has released GPT-6 Astra, the model company president Greg Brockman is calling a “generational leap,” and in some cases going as far as suggesting it could mark the arrival of AGI. Alongside the launch, OpenAI has published an extensive benchmark comparison table pitting Astra against its own predecessor GPT-5.6 Sol, Google’s Gemini 3.8 Flash, and Anthropic’s Claude lineup, including the recently released Claude Fable 5.1 and Claude Mythos 5.1. Here’s a breakdown of what the numbers actually show.
Computer Use
This is the category OpenAI leaned on hardest during the launch, with Brockman claiming Astra users “won’t ever have to click around a mouse or type on a keyboard ever again.” The scores back up the framing to a degree. Astra scores 92.7% on ScreenSpot-Pro without external tools, well ahead of GPT-5.6 Sol’s 76.9%, and comfortably clear of Claude Fable 5’s 87.3%. On OSWorld 2.0, a benchmark that measures how well an agent can actually get around a desktop, Astra posts 72.6% against Sol’s 65.7% and Claude Opus 5’s 70.2%. Notably, Claude Fable 5.1 doesn’t appear in OpenAI’s table for this benchmark at all — Anthropic has separately claimed a higher score for Fable 5.1 on OSWorld, but on a different release of the test, so it isn’t directly comparable here.

Agents’ Last Exam, a newer and tougher agentic benchmark, tells a similar story: Astra at 59.3%, Sol at 53.6%, Opus 5 at 55.5%, and Fable 5.1 trailing at 48.7%.
Professional
The professional category is where the picture gets messier, and where OpenAI’s table has the most gaps. Astra leads clearly on AutomationBench (41.4% vs Sol’s 18.1%) and BenchCAD (95.9% vs 83.3%), a pair of results that support OpenAI’s pitch that Astra is built for real workplace tasks rather than just benchmarks. But on the Artificial Analysis Intelligence Index v4.1.1, an aggregate score meant to summarise general capability, Claude Fable 5.1 actually posts the highest number of the group at 65.7, ahead of Astra’s 61.2 and Sol’s 60.9. It’s a reminder that no single model is winning across the board, and that which benchmark you highlight depends a lot on what story you want to tell.
BrowseComp, a web-research benchmark, is close to a wash: Astra at 91.5%, Opus 5 at 90.8%, and Sol at 90.4%, all within a point of each other.
Coding
Coding is arguably Astra’s most contested category. On Terminal-Bench 4.0, Astra scores 57.7%, ahead of Fable 5.1’s 55.8% and Opus 5’s 52.3%, but on DeepSWE v1.1, a 113-task agentic coding benchmark, the field bunches up tightly: Astra at 74.1%, Opus 5 at 73.7%, Gemini 3.8 Flash at 73.8%, and Fable 5.1 at 67.4%. That’s a notably small gap for a model OpenAI is calling a generational leap, and it’s worth noting that Meta’s Muse Spark 1.3 has reportedly beaten Astra’s DeepSWE score at its highest reasoning setting, a detail OpenAI’s own table naturally doesn’t include.

On FrontierCode 1.1, the field is essentially tied across every model shown, with scores clustered within a few points of each other on both the Main and Extended scoring tracks.
Academic
This is where Astra’s numbers are the most eye-catching. FrontierMath Tier 4 (v2), widely regarded as one of the hardest math benchmarks in existence, sees Astra hit 97.6%, comfortably ahead of Fable 5.1 and Claude Fable 5, who are tied at 87.8%, and well clear of Opus 5’s 73.2%. Terminal-Bench Science 0.1 shows an even larger gap: Astra at 64.6% against Fable 5.1’s 52.6% and Opus 5’s 30.0%.
GPQA Diamond, on the other hand, has practically every model bunched between 92.6% and 96.0%, suggesting the test may be nearing saturation across frontier labs. Humanity’s Last Exam (with tools) has Fable 5.1 slightly ahead of Astra, 65.0% to 57.2%, one of the few academic rows where Anthropic’s model comes out on top.
Science And Health

OpenAI’s table shows Astra and Sol going head-to-head here, without any Claude or Gemini scores included for GeneBench Pro or MedChemBench, both of which OpenAI leads (37.8% and 49.3% respectively). On HealthBench Professional (length-adjusted), the one row where all five models are represented, the ranking is Astra (63.4%), Sol (60.5%), Fable 5 (60.9%), Fable 5.1 (56.6%), and Opus 5 (54.5%). Interestingly, Fable 5.1 scores lower here than its own predecessor Fable 5, which Anthropic has attributed elsewhere to its safety classifiers intervening on health-related queries and depressing the measured score.
Cybersecurity
This is the category OpenAI has framed as the most significant part of the Astra release. The company says Astra is the first model to cross its “Critical” capability threshold for cybersecurity under its Preparedness Framework, and the benchmark numbers reflect that: a perfect 100.0% on ExploitBench, against Sol’s 78.5% and Opus 5’s 70%. SRE-Bench shows an even starker gap, Astra at 88.0% versus Opus 5’s 12.5%.
Given those capabilities, OpenAI has restricted Astra’s most advanced cyber-offensive functionality to vetted participants in its Daybreak program for cybersecurity defenders, similar in spirit to how Anthropic gates Claude Mythos 5.1’s capabilities behind its own trusted-access program, Project Glasswing.
Alignment

OpenAI has also published internal alignment numbers, and used them to argue Astra is its best-behaved model to date. On an internal computer-use safety benchmark (lower is better), Astra scores 2.4% versus Sol’s 22.0%, and the gap holds on an internal circumvention benchmark, near-zero for Astra against 0.29% for Sol. The standout figure is the ExploitGym honeypot test, where Astra registered 0.0% against Sol’s 48.2% — OpenAI says this reflects Astra not attempting to go outside an authorized target even in scenarios designed to tempt it. Astra’s hallucination rate is also down considerably, 4.2% compared to Sol’s 12.2%.
Long Context
Astra’s long-context numbers are strong but not dramatically different from Sol’s. On the OpenAI MRCR v2 8-needle test at the 256K-512K range, Astra scores a perfect 100.0% against Sol’s 91.5%; stretching to the 512K-1M range, Astra holds up better too, at 96.3% versus 73.8%. Neither Claude nor Gemini numbers appear on this particular test in OpenAI’s table.
The Bigger Picture
Taken together, the table shows a model that’s ahead of the field on several genuinely hard benchmarks, especially FrontierMath, cybersecurity, and long-context retrieval, while being closer to a toss-up with Anthropic’s newest Claude Fable 5.1 on general-purpose coding and aggregate intelligence scores. It’s also worth remembering, as with any day-one release, that these are OpenAI’s own reported numbers rather than independently reproduced results, and several rows simply have no comparison data for one or more rival models. .