Islamabad (GNP): OpenAI has said it cannot rule out that its upcoming AI model, Astra, possesses “critical” cybersecurity capabilities, prompting the company to pause certain internal development work and activate its safety protocols.
Under OpenAI’s safety guidelines, a model is classified as reaching the critical threshold if it can autonomously identify and exploit severe, real-world software vulnerabilities known as zero-day exploits, or carry out complex cyberattacks against highly secure targets without human intervention.
Preliminary evaluations conducted over the past several days, along with assessments from outside experts, indicated that Astra may be capable of performing increasingly sophisticated cyber tasks autonomously, according to OpenAI.
The company said its ongoing benchmarking and assessment work has shown strong enough performance that it cannot currently rule out the model reaching the critical capability level. In response, OpenAI said it has scaled up its security controls and paused internal activities involving Astra that do not meet its newly strengthened requirements, with the model’s continued development now being moved into isolated testing environments featuring restricted network access and sandboxed execution.
Also Read: China Opens World’s First Dedicated School for Robots in Hangzhou
The development follows an earlier report that OpenAI had discovered additional instances of autonomous AI agents escaping containment, as the company widened its investigation into a hacking incident at technology firm Hugging Face that drew global attention in July.
OpenAI clarified that Astra was not involved in the Hugging Face hack. In recent weeks, OpenAI, Anthropic and Meta Platforms have each disclosed that their AI models broke into other companies’ systems during cybersecurity testing, underscoring how rapidly advancing AI capabilities are testing developers’ ability to keep such systems reliably contained.
Despite the risks identified in preliminary testing, OpenAI chief executive Sam Altman said on social media platform X that the company is working toward making Astra generally available, stating that OpenAI does not consider it good strategy to restrict access to powerful models to a select few. OpenAI said it will partner with government agencies and select AI safety organisations to further test the model’s capabilities as evaluation work continues.





