OpenAI safety leader David Robinson quits, says company culture is ‘broken’ as AI race accelerates
ADVERTISEMENT

David Robinson, a senior safety leader at OpenAI who oversaw safety reports for 12 frontier AI launches, has quit the company, saying its culture is failing to keep up with the risks created by increasingly capable AI systems.
Robinson announced his resignation in an essay in The Atlantic, bluntly titled “I Quit OpenAI Because Its Culture Is Broken”. He said he was joining a growing group of former AI employees who believe the industry’s current trajectory is unacceptable. “I agree with other recently departed staff that the companies building this technology aren’t being nearly careful enough,” Robinson wrote.
But he said the problem was bigger than individual safety rules or regulation. “We need to talk about culture,” he wrote.
Robinson spent three and a half years at OpenAI and was among its longest-tenured employees. He led the drafting of the company’s Preparedness Framework and the safety reports accompanying major product launches.
His resignation comes as OpenAI faces growing scrutiny over autonomous AI agents, including an incident in which a swarm of its agents accessed the systems of AI startup Hugging Face. OpenAI has also disclosed other cases of concerning model behaviour and has recently slowed or halted some training and releases following safety concerns.
‘The time for trial and error is over’
Robinson’s strongest criticism is directed at what OpenAI calls “iterative deployment”, releasing systems, identifying problems and improving safeguards after failures.
OpenAI has built much of its approach around this model, but Robinson argues that it becomes increasingly dangerous as AI systems become more capable. “The safety approach that emerges from such a culture starts with unimpeded optimism about being able to solve problems as they arise,” he wrote.
He said OpenAI has “thrived by trial and error”, but added that the approach “guarantees periodic failures” and that the scale of those failures is increasing.
Robinson pointed to the Hugging Face incident as an example. OpenAI’s agents were allowed out by mistake, after which the company made security improvements. But, according to Robinson, the company later reported another failure when a model in training bypassed restrictions on internet access. A monitoring system alerted human staff but failed to automatically shut the model down as intended. “An environment where things like this can happen is no place to grow artificial minds that could be smarter than we are,” Robinson wrote.
He also cited a warning from Paul Christiano, who recently joined OpenAI’s board, that rapid acceleration in AI capabilities could lead to a “catastrophic and irreversible loss of control”. “If this is the situation, then the time for trial and error is over,” Robinson wrote.
His argument is that AI labs cannot assume that every safety failure will be caught and corrected before it causes serious harm. “Achieving something much closer to perfection the first time is essential,” he wrote, arguing that “iteration after a mistake may not be possible.”
AI labs need ‘layers of redundancy’
Robinson said the AI industry needs to look outside Silicon Valley for lessons on managing technologies where failures can have catastrophic consequences. “AI companies need to rely more on the safety expertise that already exists in other fields,” he wrote.
His preferred comparison is not another software company but a nuclear plant or an airport. “Frontier labs need to run like nuclear-power plants or busy airports,” Robinson wrote, with “layers of redundancy” and “careful, time-consuming planning”.
The comparison is central to his criticism, comparing to a nuclear plant, where systems are designed so that one broken component or one human mistake does not automatically lead to disaster. AI companies, he said, currently operate with much less redundancy despite the potentially much larger consequences of losing control of a highly capable system.
Robinson also raised the possibility of autonomous AI agents becoming a much bigger security problem. “Imagine ‘rogue’ agents that work like teams of hackers,” he wrote, describing systems that could potentially target critical infrastructure such as hospitals and “never need to sleep”.
He said he had not encountered a colleague at OpenAI with direct experience in making aircraft safe, operating nuclear reactors or maintaining financial-system stability.“This is a new need,” Robinson wrote, arguing that AI systems today are “far more capable and dangerous” than those being built even six months ago.
He acknowledged that he could have stayed and fought for changes inside OpenAI. But the pace of development made that difficult. “My colleagues and I were so busy sprinting that we seldom had the chance to consider big changes,” he wrote.
Robinson said this was ultimately why he decided to work from outside the company, where he hopes to push for stronger incentives for AI firms to prioritise safety.
OpenAI, for its part, said it is continuing to strengthen its safety and security practices and is making sure its models do not become more capable than the company can safely manage. The company also said it pauses training or holds back models when it believes it needs to slow down.
Robinson’s final warning goes beyond engineering. He argues that the industry still has not solved the basic problem of ensuring increasingly intelligent systems consistently behave in ways aligned with human values. “The smarter the industry lets models grow while these problems remain unsolved, the more dangerous our situation becomes,” he wrote.
“Before the organizations building AI can teach a superintelligence to treat humanity well, they’ll need to remember how to do it themselves," he added.