OpenAI safety leader David Robinson quit, saying the current path for AI is unacceptable. Robinson, who joined in 2023, claimed the company’s culture is broken. He wrote, “The future depends on wisdom that Silicon Valley lacks.” He warned that the firm’s trial and error approach guarantees growing failures as systems get smarter.

David Robinson, OpenAI’s former safety lead, has joined a long list of AI researchers across various labs who have quit. Announcing his exit, Robinson said that OpenAI’s “culture is broken” as the industry as a whole faces growing risks around more advanced models.

The OpenAI safety lead was one of the company’s longest-tenured employees. “I led the writing of the safety reports we published with each major launch,” David Robinson wrote in an essay for The Atlantic. “Now I’m joining a parade of former colleagues—at OpenAI and the industry’s other leaders—who have decided that the current path is unacceptable.”

According to Robinson, companies building AI are not being as careful as needed. However, he argued that it was time for us to look deeper. “But I believe that we need to look deeper than specific rules or new laws. We need to talk about culture,” he added.

David Robinson joined OpenAI in 2023, first as head of policy planning, and later became a leader on its Safety Systems team. In earlier public remarks, he had defended gradual deployment, saying in 2023, “We believe in deploying gradually and then learning as we go.” He studied philosophy at Princeton and later completed a degree in philosophy, politics and economics at Oxford.

AI companies lack wisdom to handle risky tech

David Robinson pointed out that the future of AI lies amongst the companies building it. But there was one aspect central to this future. “The future depends on wisdom that Silicon Valley lacks,” he wrote. “Wisdom about how to handle dangerous technology and, more fundamentally, wisdom about what it means to care for people.”

The former OpenAI safety lead claimed that this wisdom would require a certain degree of humility that wasn’t “natural for people who have succeeded through their extreme confidence.” Robinson wrote, “My former colleagues at OpenAI were prescient: They came to understand the scaling laws that meant bigger AI systems would be smarter—and so they went all in on building bigger systems, at great cost.”

David Robinson said OpenAI’s approach had been shaped by “trial and error,” or what the company calls “iterative deployment,” with guardrails improved after problems emerge. “But this approach, by its very nature, guarantees periodic failures — and the scale of those failures is growing as systems get more capable,” he wrote.

Robinson pointed out recent incidents of AI agents going rogue as examples of when things can go rogue. He wrote that during the Hugging Face incident, “OpenAI let a swarm of agents out by mistake.” David Robinson said the company later reported another failure when “a model in training bypassed restrictions on internet access,” with a monitoring system alerting staff but not shutting the model down automatically “as it was supposed to.” Anthropic had also acknowledged accidentally turning off its own safeguards because of a misconfiguration. Robinson said such mistakes were “typical of the industry, given the speed and flexibility with which people operate.”

Time for trial and error is over

David Robinson said this was not a suitable environment in which to build systems that could become more capable than humans. Referring to OpenAI board member Paul Christiano’s recent warning that there is “a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term,” Robinson wrote that “the time for trial and error is over.” He added, “Achieving something much closer to perfection the first time is essential because iteration after a mistake may not be possible.” That is, companies cannot afford to stick to this older method as it may not be possible to fix something if things go wrong in the future.

The former OpenAI safety lead called for two urgent changes – greater reliance on safety expertise from other high-risk fields, and “new science” to ensure more capable models “will make safe choices when we aren’t looking.” To make sure this happens. David Robinson insisted AI companies needed to run “like nuclear-power plants or busy airports, with layers of redundancy and careful, time-consuming planning.” This would ensure that the occasional and inevitable human error does not open a door to disaster, he said.

Robinson said he had not taken the decision to leave lightly and still believed AI “can be useful and valuable”. However, he said stronger pressure from outside companies would now be needed. “Now I plan to work on the outside, in the hope that I can help more people understand the risks I saw, and strengthen the incentives OpenAI and other firms have to be safer,” he wrote.

The resignation comes amid a wider debate over AI safety. Researchers like Jacob Coxon have warned that AI may end humanity within a few years. Even Anthropic CEO Dario Amodei and OpenAI’s Sam Altman have acknowledged such concerns. Though some, like Meta’s Mark Zuckerberg and Nvidia CEO Jensen Huang, have been less concerned.

- Ends