OpenAI will pause some work on an artificial intelligence model after an internal evaluation concluded the system reached a security risk threshold. The company said on Friday that its review of the agent named Astra identified “significant advancements in agentic coding and cybersecurity” and that the capabilities had advanced to a level at which the system can discover and exploit vulnerabilities without human intervention.
OpenAI described the findings as reaching a “critical” threshold, noting that when given only a high-level desired goal the agent could devise and execute cyber-attacks. The decision to pause some development follows a series of incidents reported within the industry in which autonomous agents have proven difficult to contain during testing and research exercises.
The company framed the action as targeted and precautionary, stating the pause applies to portions of work related to the model’s agentic capacity while safety teams assess the implications. In its statement, OpenAI emphasized the technical nature of the concerns, citing the agent’s ability to perform tasks traditionally requiring human oversight in cybersecurity and code generation contexts.
The announcement does not provide a specific timeline for resuming the paused activities. Industry observers have noted that advances in autonomous coding and network interaction raise distinct security and governance questions that companies and regulators are increasingly addressing through reviews, red-team testing and stricter deployment controls.
Astra‘s evaluation and the resulting pause underscore ongoing tensions between rapid capability development in AI and the need for robust safeguards. OpenAI said it would take steps to address the identified risks and to determine next steps for the affected lines of work, while continuing other research and deployment activities not implicated by the findings.




