OpenAI safety employee quits, challenging the culture behind its safeguards
David Robinson’s criticism puts attention on the gap between published AI safety procedures and how a company applies them under pressure.
An OpenAI safety employee has resigned and publicly challenged the company’s operating culture, arguing that increasingly powerful artificial intelligence demands a more deliberate approach than learning through mistakes after the fact.
David Robinson set out his criticism in an October 3 essay in The Atlantic. He said he had led drafting of OpenAI’s current Preparedness Framework and overseen safety reports for twelve frontier-model launches. His account is a former employee’s assessment of the organization, rather than an independent investigation establishing that a particular model is unsafe.
Reuters reported that Robinson had worked at OpenAI for three and a half years. In response to the criticism, the company said it could pause training or hold models back when necessary and sought to prevent capabilities from exceeding its ability to manage and secure them. The disagreement therefore concerns how safety commitments operate in practice as much as whether the company has written commitments at all.
OpenAI’s published framework provides context for that dispute. Its April 2025 update describes a process for identifying advanced capabilities that could cause severe harm. It prioritizes risks judged plausible, measurable, severe, new and either instantaneous or difficult to reverse. The framework is a company-designed governance system, not an external certification that its models cannot cause harm.
That update separates two questions: what a model is capable of and whether safeguards adequately control the resulting risk. It describes capability reports alongside dedicated safeguards reports. A Safety Advisory Group reviews those materials, assesses remaining risk and recommends to leadership whether deployment is sufficiently safe.
The distinction is consequential. A strong evaluation result does not by itself establish that defenses will work outside a test, while a written mitigation is different from evidence that it is effective. The framework’s own emphasis on layered defenses and assessment of residual risk recognizes that neither capability testing nor a single protective measure settles the deployment decision.
In May 2026, OpenAI also published a Frontier Governance Framework linking its practices to emerging regulatory requirements. The company said that document covered cyber offense, chemical, biological, radiological and nuclear risks, harmful manipulation and loss of control. It also included incident response, external expertise, security management and updates to the framework.
Those published documents make it possible to ask specific accountability questions: what findings reached decision-makers, which safeguards were tested, what risks remained and who authorized a release. Robinson’s resignation does not answer those questions for an individual model, but it directs attention to the organization responsible for answering them.
OpenAI’s April framework committed to publishing preparedness findings alongside frontier releases. Its May governance document said it would evolve as capabilities, evaluations and regulatory requirements changed. Those are identifiable public commitments against which subsequent releases and disclosures can be assessed.