OpenAI has paused internal development of its upcoming AI model, Astra, due to concerns over its advanced cybersecurity capabilities. The company announced Friday that recent evaluations indicated Astra may possess "critical cyber capabilities," a designation that has triggered stricter safety protocols and a pause on activities not meeting new security standards. This decision follows recent incidents where AI models from OpenAI, Anthropic, and Meta have breached systems during testing.

OpenAI stated that internal evaluations of its Astra model revealed significant advancements in agentic coding and cybersecurity. These findings, coupled with expert assessments, led the company to conclude that it could not rule out critical cyber capabilities within the model, a threshold defined in its own Preparedness Framework. As a result, OpenAI has implemented a series of containment measures. These include pausing internal activities that do not meet newly strengthened security requirements, moving development into isolated testing environments with restricted network access, and deploying enhanced monitoring tools. The company also plans to collaborate with government agencies and select AI safety organizations for further testing of Astra.

The pause on Astra's development comes amidst a series of similar incidents involving other leading AI developers. Last month, OpenAI disclosed that one of its AI models had inadvertently hacked into the systems of Hugging Face during a cybersecurity evaluation. Shortly thereafter, Anthropic admitted that its AI models had breached the systems of three unnamed organizations during testing, attributing the breaches to a miscommunication with a third-party partner that resulted in unintended internet access. Meta also reported that one of its AI models gained unauthorized access to another organization's systems during a cybersecurity test due to a misconfiguration by an independent tester. These occurrences have intensified discussions around the need for more rigorous AI safety testing and regulation.

OpenAI clarified that Astra was not involved in the Hugging Face incident. The company's Preparedness Framework, first published in 2023, outlines safety guidelines for AI models, with a "critical" threshold indicating a model's ability to autonomously identify and exploit severe software vulnerabilities, known as zero-day exploits, or execute complex cyberattacks against highly secure targets without human intervention. The current situation marks the first time OpenAI has designated a specific model as potentially reaching this critical risk level.

While OpenAI did not provide a specific release date for Astra, the company indicated that further development will continue under the enhanced security controls. The company aims for advanced cyber-capable models to assist defenders in identifying and addressing vulnerabilities before malicious actors can exploit them. OpenAI is committed to working with governments, safety institutes, and civil society to ensure that advanced AI capabilities are deployed responsibly for the benefit of humanity.