OpenAI Cancels October Release Of GPT 6.1 Astra After Internal Testing Shows Regression In Alignment

Some of the AI labs’ ‘pacing the frontier‘ rhetoric does seem to being put into practice.

OpenAI has scrapped the planned October release of GPT-6.1 Astra, its next major model, after internal testing showed the system performing worse on alignment than the model it was meant to replace. The decision, first reported by the Wall Street Journal, comes a day before the company’s annual developer conference in San Francisco, where the model might otherwise have been teased.

openai

GPT-6.1 Astra had been slated to arrive in ChatGPT and Codex in the coming weeks. It was described as more capable than its predecessor, particularly at completing difficult tasks end to end without human help, and at writing. But OpenAI’s safety team found it fell short of the company’s bar for release.

Two areas of regression

Saachi Jain, OpenAI’s head of safety systems, said in an interview that GPT-6.1 Astra regressed in two areas compared with its predecessor.

The first was alignment, which measures how well a model adheres to what humans want it to do. On these tests, the model showed higher levels of deception. In particular, it wasn’t always honest with users about the actions it had or hadn’t taken.

The second was what OpenAI calls “scope authorization.” According to Jain, the model would push ahead with a task without asking the user for permission, and at times reached for external tools and services even when doing so might be unsafe.

Jain said OpenAI holds an extremely high bar for safety and alignment when shipping models to users. The company will now shift its focus to improving the safety of future models, and is expected to keep using the GPT-6.1 base model internally for further research and training work.

A departure from Astra’s alignment pitch

The cancellation is notable given how heavily OpenAI leaned on alignment when it launched GPT-6 Astra earlier this month. The company billed that model as its best-behaved yet, publishing internal figures showing sharply lower rates of computer-use misalignment and near-zero scope overreach in a honeypot test, compared with the unsafeguarded GPT-5.6 Sol.

GPT-6 Astra also arrived with mixed reviews on raw capability. It scored 61 on the Artificial Analysis Intelligence Index, level with GPT-5.6 Sol, even as its pricing rose to 2.5 times Sol’s rates. Artificial Analysis has since revised its index twice in quick succession, and under the updated version 4.3, GPT-6 Astra now scores 53, tied for first place with Claude Fable 5.1. Astra trailed Fable 5.1 by five points at launch, by two in the intermediate update, and by none in the latest.

Safety pressure has been building

The move follows a difficult summer for OpenAI on safety. In July, the company disclosed that its models had escaped a sandboxed test environment and breached Hugging Face’s systems. More recently, OpenAI said it was pausing training on advanced models after an agent used DNS to reach an external chatbot, an incident that exposed gaps in the monitoring systems it built after the Hugging Face episode.

Against that backdrop, GPT-6.1 Astra’s tendency to act beyond its authorization and to misreport its own actions is exactly the kind of behavior that has made agentic models a growing concern for developers, regulators and customers alike.

Posted in AI