OpenAI Announces GPT 6.1 Sol, Says It Has Astra-level Intelligence At 1/5th The Price

OpenAI has announced GPT-6.1 Sol, an upgrade to GPT-6 Sol that the company says nearly matches the intelligence of its flagship GPT-6 Astra on agentic coding, computer use and professional work, while costing a fifth as much. The model is priced at $2 per million input tokens and $10 per million output tokens, against $10 and $50 for GPT-6 Astra.

The release comes barely a week after OpenAI widened its GPT-6 lineup with Sol and Luna, and it pushes the company’s mid-tier model considerably closer to the top of the range.

Cheaper caching for agents

The biggest pricing change is on cached input, which now costs $0.10 per million tokens. OpenAI says that is 95% less than standard input pricing and 50% less than GPT-6 Sol’s cached input rate. The company is pitching this at developers building agents that reuse large amounts of context across requests, where cached tokens make up much of the bill.

Here is how the three GPT-6 models now stack up, per million tokens:

ModelInputOutputCached input
GPT-6 Astra$10$50$1
GPT-6.1 Sol$2$10$0.10
GPT-6 Luna$0.10$0.50$0.01

Coding: matching Astra at a fraction of the cost

On DeepSWE v1.1, a benchmark of long-horizon software engineering tasks in real codebases, OpenAI says GPT-6.1 Sol matches GPT-6 Astra at roughly one-fifth of the cost. It also beats GPT-6 Sol’s best score by 6.4 percentage points while running at a lower reasoning effort and lower cost.

The cost-versus-score chart OpenAI shared shows the gap clearly. GPT-6.1 Sol peaks at a little above 75% at well under $1 per task, while GPT-6 Astra tops out at around 74% at roughly $4.40 per task. GPT-6 Sol’s best result, at around 69%, needed close to $3.

Professional work and computer use

OpenAI is also making comparisons with rival models. On GDP.pdf, which tests how accurately models answer professional questions about complex PDFs with tables, charts and fine print, GPT-6.1 Sol scores higher than Anthropic’s Opus 5.5 with fallbacks at less than half the cost per task, according to OpenAI. It also approaches Astra’s state-of-the-art performance at roughly a fifth of the cost.

On AutomationBench, which measures whether agents can complete multi-step business workflows, GPT-6.1 Sol is 2.2 percentage points ahead of Opus 5.5 at medium reasoning effort, at roughly a third of the cost. It is also up 4.8 points from GPT-6 Sol at the same setting. OpenAI notes that the datapoint for Claude Fable 5.1 understates its real cost, since it leaves out fallbacks that occurred on around 40% of tasks.

For computer use, OpenAI says GPT-6.1 Sol beats GPT-6 Sol by seven percentage points on the offline set of OSWorld 2.0 at maximum reasoning effort, at less than half the cost. It comes within 2.1 points of Astra’s score at roughly one-seventh the cost per task.

Science and factuality

On Terminal-Bench Science 0.1, GPT-6.1 Sol more than doubles GPT-6 Sol’s score at maximum effort. At that setting it costs an average of $5.47 per task, compared with $23.21 for Opus 5.5 and $23.80 for Astra. Astra still posts the highest score of the models tested, at 68.1%, and OpenAI says it should remain the pick for the most difficult research tasks.

On factuality, the model’s biggest gain over GPT-6 Sol comes at low reasoning effort, where the share of responses containing a factual error drops from 11.4% to 7.7%, a reduction of about 32%. Across the tested settings, its error rate stays within 1.9 percentage points of Astra’s. OpenAI cautions that the evaluation uses deliberately difficult prompts drawn from conversations where users flagged errors, and isn’t representative of typical usage.

Safety and alignment

OpenAI says GPT-6.1 Sol shows substantial improvements over GPT-6 Sol on its alignment evaluations, bringing it closer to Astra. It is more transparent about its limitations and more reliable at respecting user intent and safety constraints, the company says.

On a test of whether agents tell users when their search tool is broken rather than guessing, GPT-6.1 Sol failed to disclose the problem in 2.1% of cases, compared with 4.9% for GPT-6 Sol, 1.5% for Astra and 28.7% for GPT-6 Luna. OpenAI also says it observed no attempts by the model to bypass an automated safety reviewer, matching Astra and GPT-6 Sol. The tasks are designed to elicit failures and don’t reflect typical use.

Availability

GPT-6.1 Sol is available starting today to Plus, Pro, Business, Enterprise and Edu users in ChatGPT Work and Codex. It is not yet available in Chat. Developers can access it through the API as gpt-6.1-sol. OpenAI also plans to offer GPT-6.1 Sol Ultrafast in the coming days, with up to 8x faster token generation than standard speed in Codex.

The launch lands in an increasingly crowded race for cost-efficient frontier models. Anthropic released Claude Sonnet 5.5 only a day ago, positioning it as a cheaper alternative to Opus 5.5, and the pressure on pricing is showing on both sides. As always, the benchmark comparisons here come from OpenAI itself, and independent evaluations will show how well the claims hold up.

Posted in AI