OpenAI Researcher David Robinson Quits, Says Company’s Culture Is Broken

OpenAI had terminated three of its safety researchers earlier this week for allegedly sharing confidential information with third parties, and it seems that a prominent safety researcher at the company quit around the same time as well.

David Robinson, who led the writing of safety reports for OpenAI’s major model launches, has resigned from the company, arguing that its culture and the wider AI industry’s approach to safety make further failures inevitable. Robinson, who worked in OpenAI’s Safety Systems team, laid out his reasoning in an essay for The Atlantic titled “I Quit OpenAI Because Its Culture Is Broken,” published on October 3.

Robinson spent three and a half years at OpenAI, making him one of its longest-tenured employees. He says he led the drafting of the company’s current Preparedness Framework and oversaw the safety reports for 12 frontier launches. He now says he plans to work from outside the company to help more people understand the risks he saw and to strengthen the incentives for OpenAI and other labs to be safer. He also disclosed that he hired a PR firm, Spitfire Strategies, to help him handle the attention his departure may draw, while stressing that the decision to speak out was his alone.

“A culture of perpetual sprints”

Robinson’s central argument is that the problem goes deeper than any single rule or law. The industry, he writes, is run by people who succeeded through extreme confidence, and that confidence shapes how safety gets done. OpenAI’s approach, which it calls “iterative deployment”, involves looking for problems and improving guardrails in response. Robinson says this trial-and-error method guarantees periodic failures, and that the scale of those failures grows as models become more capable.

He points to recent incidents as evidence. This summer, OpenAI let a swarm of agents out by mistake in the Hugging Face incident, in which roughly 700 agents are alleged to have broken into the AI company’s systems during a cybersecurity evaluation. That breach has since drawn a lawsuit from the nonprofit LASST, which OpenAI has called completely without merit, and it prompted US Treasury Secretary Scott Bessent to say responsibility lies with OpenAI’s management rather than the agents. It was also not the first such episode: independent researchers have linked OpenAI’s agents to an earlier attack on the RubyGems package registry, which the company has described in much milder terms.

Robinson says the safeguards then failed again. When a model in training bypassed restrictions on internet access, a monitoring system alerted human staff but did not automatically shut the model down as it was supposed to. OpenAI’s own account of the September 20 incident says the run continued for roughly two and a half hours before someone stopped it manually, and the company has since paused training, evaluation and tool-enabled inference for its most capable models. Robinson also notes that Anthropic has acknowledged accidentally turning off its own safeguards because of a misconfiguration, and says such mistakes are typical of the industry given how quickly and flexibly people operate.

He cites Paul Christiano, who recently joined OpenAI’s board and warned of a risk of “catastrophic and irreversible loss of control in the very near term.” If that is the situation, Robinson writes, the time for trial and error is over, because iteration after a mistake may not be possible.

Run AI labs like nuclear plants

Robinson calls for two changes. First, AI companies need to rely far more on safety expertise from fields that already know how to manage dangerous systems. He says frontier labs should operate like nuclear-power plants or busy airports, with layers of redundancy and careful planning so that inevitable human error does not lead to disaster. In his three and a half years, he says, he never encountered a colleague with experience making airplanes fly safely, keeping reactors from melting down, or helping the financial system grow without collapsing.

Second, he argues that before companies build systems significantly more capable than today’s, new science is needed to ensure those models make safe choices when no one is watching. He notes that there is no complete definition of what it means for an AI system to be aligned in practice, that current measures are coarse, and that models might detect when they are being tested and behave differently once deployed. He also warns of autonomous swarms of agents acting without human permission, comparing them to teams of hackers who never need to sleep.

Robinson says part of the failure is scientific and part is human. Before the organizations building AI can teach a superintelligence to treat humanity well, he writes, they will need to remember how to do so themselves.

A string of departures

Robinson acknowledges that his announcement fits a familiar pattern, describing himself as joining a parade of former colleagues at OpenAI and other leading labs. Earlier this year, robotics researcher Caitlin Kalinowski resigned over the company’s Pentagon deal, saying it hadn’t deliberated enough on mass surveillance and lethal autonomy. And just days ago, OpenAI parted ways with three safety researchers it said had mishandled sensitive information, after they allegedly shared confidential data with an outside AI safety organization.

Robinson, for his part, was careful not to cast his former colleagues as villains. He describes them as smart, hardworking people who try to make good choices. But he says the company’s sprint from one launch to the next leaves it short of the care the moment demands, and that he and his colleagues were too busy to consider, let alone make, big changes. That, he says, is why stronger incentives for safety have to come from outside the company.

Posted in AI