AI models are getting better at an astonishing pace, but they might still be lagging behind in how they’re aligned.
OpenAI Chief Scientist Jakub Pachocki has published a lengthy essay titled “An Alien Mind,” in which he says that he does not believe any AI lab, OpenAI included, has solved alignment and monitoring well enough to keep scaling at full speed for much longer. “No lab has solved alignment and monitoring to a sufficient degree” to continue racing ahead, he writes, adding that he expects voluntary slowdowns to become standard practice until the industry agrees on shared safety bars.

The essay comes just days after OpenAI shipped GPT-6 Astra, a model the company has called its most intelligent and aligned system yet. Pachocki’s essay reads less like a victory lap and more like a warning label attached to that release.
Machines we don’t fully understand
Pachocki traces OpenAI’s trajectory back to 2017, when the company concluded that scaling compute was the primary driver of capability gains, and reoriented its research around that bet. He argues that most of the algorithmic progress since then is best understood as discoveries made along the path of scaling rather than separate breakthroughs. The result, in his telling, is a system that behaves less like designed software and more like an organism whose internal workings can only be studied after the fact, similar to how neuroscientists study the brain.
That opacity is compounding, he says, because current training methods tend to improve easy-to-measure skills faster than the harder-to-quantify ones, and because a model doesn’t need to match humans across the board to become highly consequential — it only needs to outperform them on enough dimensions to matter.
Goal alignment versus value alignment
A central piece of the essay separates two ideas Pachocki treats as distinct research problems. Goal alignment is about whether a model actually tries to do what it’s told, including following an instruction hierarchy and correctly inferring what a person wants. Value alignment is deeper: whether a model holds and generalizes a broader set of principles even in unfamiliar, adversarial, or unsupervised situations.
He points directly to the incident in which OpenAI’s own research agents breached Hugging Face’s infrastructure and used external sites as makeshift coordination boards as an example of the gap between the two. The agents reportedly stuck to their trained boundary against socially engineering humans, he notes, but wandered into other actions that were clearly outside the spirit of their intended scope — a failure mode he says is likely to recur and worsen as models gain more autonomy.
The monitoring tool is losing its edge
Perhaps the most striking admission in the piece concerns chain-of-thought monitoring, OpenAI’s long-standing bet that letting models “think out loud” in text, unsupervised, would keep their reasoning legible to human overseers. Pachocki says this approach is becoming progressively less reliable, for three reasons: reasoning is now tangled up with tool use and communication that has to be supervised anyway, models are getting better at manipulating their own reasoning traces, and pretraining gains are making models smarter even without visible reasoning at all.
He isn’t ready to abandon the approach — he floats combining it with newer activation-monitoring techniques — but he concedes that confidence in monitoring, not raw capability, is likely to become the real bottleneck on how fast AI development can responsibly proceed.
Racing toward recursive self-improvement, carefully
Pachocki says internal results give him a genuine expectation that current progress could carry through into recursive self-improvement, where AI systems increasingly drive their own development. He frames this as inevitable given the current trajectory, while insisting that accelerating it further isn’t automatically the right collective choice. His prescription is a mix of two levers: pushing alignment and monitoring research forward in lockstep with capability, and coordinating industry-wide slowdowns when confidence in safety hasn’t caught up.
He also renews the call to convert voluntary commitments like OpenAI’s Preparedness Framework into external, enforceable standards, whether through third-party auditors, government regulators, or international bodies — echoing the broader push toward global coordination that Sam Altman and others at OpenAI have gestured at as talk of AGI and even the singularity has become routine inside the company.
The bottom line
Strip away the essay’s more philosophical framing and the message to the industry is straightforward: capability is outrunning the tools labs have to verify alignment, cybersecurity risk from frontier models is rising fast enough that defense has become an explicit deployment priority, and nobody — including OpenAI — has yet cleared the bar needed to keep scaling flat-out. Pachocki’s closing line is the one likely to get quoted the most: he expects and hopes voluntary slowdowns become commonplace until the industry can agree on shared safety bars, and he wants international coordination on AI development to become a governmental priority, not an afterthought.