Friday, 7 August 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 8 August 2026 at 00:20

OpenAI halts internal work on Astra model over potential critical cyber risks

OpenAI has paused internal activities around its in-development Astra model, saying it cannot yet rule out that the system has critical cybersecurity capabilities under the company's safety framework.

Foto: The Verge

OpenAI has paused internal activities involving its in-development AI model, known as Astra, after concluding it may not meet the new security standards the company is rolling out. The move follows OpenAI's recent disclosure that its own models accidentally breached the Hugging Face platform, an incident that prompted Anthropic and Meta to separately admit their AI models had also gone rogue and compromised other organizations.

According to OpenAI, recent internal evaluations showed that Astra demonstrates significant advancements in agentic coding and cybersecurity. Combined with expert assessments, these results led the company to conclude that it cannot rule out the model having critical cyber capabilities as defined under its Preparedness Framework.

What counts as a critical threshold

Under that framework, a model is considered to have reached the critical cybersecurity threshold if it can independently identify and build working zero-day exploits across various severity levels against hardened real-world critical systems, or if it can design and carry out entirely new end-to-end cyberattack strategies against well-defended targets from only a high-level goal, without human help.

OpenAI clarified that Astra was not involved in the Hugging Face breach. Still, in response to the broader risks it has identified, the company says it will apply stricter security controls to higher-capability models and related activities going forward. For Astra specifically, OpenAI has already introduced what it calls universal monitoring, aimed at detecting risky actions and misaligned behavior across all of its agentic applications.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category