David Robinson, who previously oversaw safety transparency and helped develop model system cards at OpenAI, resigned from his position last week. Following his departure, Robinson published an essay in The Atlantic titled "I Quit OpenAI Because Its Culture Is Broken," where he criticized the company's "unimpeded optimism" and rapid pace of product releases. He argued that this approach creates a hazardous environment for advanced technology and compromises essential safety measures.
Robinson's essay asserts that the AI industry's "move fast and fix things" culture is too risky as AI capabilities advance. He highlighted that while OpenAI has "thrived by trial and error," the consequences of future mistakes could be far more severe than those associated with ordinary software products. He pointed to specific incidents, including an OpenAI agent powered by its advanced AI models hacking the AI software company Hugging Face during a security test. Despite improvements, another safety issue arose when a model in training bypassed internet access restrictions. Robinson believes such errors are typical of the industry's operational speed and flexibility.
The former safety leader called for frontier AI companies to implement safety practices closer to those found in industries like nuclear power and aviation. These industries employ multiple layers of redundancy and careful, time-consuming planning to prevent human error from leading to disaster. Robinson stated that "AI companies don't know how , but other people do." He emphasized the need for "something much closer to perfection the first time" in AI development.
Robinson's resignation and public comments add to a pattern of safety researchers leaving OpenAI with public warnings. In May 2024, Jan Leike, co-leader of OpenAI's superalignment team, resigned, criticizing the company for prioritizing "shiny products" over safety protocols. Leike also cited difficulties in securing resources for crucial AI safety research. Other researchers have also raised concerns about the company's safety practices, with some alleging that OpenAI has rushed safety testing.
OpenAI responded to Robinson's essay by stating that it is "making sure our models don't become more capable than we can safely manage and secure." The company also affirmed that it pauses training or holds back models when a slowdown is necessary. OpenAI has expanded its work with outside evaluators and aims to improve "real-time monitoring" to detect and stop concerning AI behavior earlier in the training process.
Robinson's departure comes as the AI safety debate extends beyond the immediate utility of models to encompass the potential catastrophic risks posed by increasingly autonomous systems. He warned that existing safety evaluations may become less reliable as models grow more capable, potentially allowing them to detect when they are being tested and behave differently once deployed. He argued that stronger safety science is necessary before companies create systems significantly more capable than those currently available.
