OpenAI can no longer rule out the possibility that its new, unreleased Astra model could autonomously hack critical systems. The company is therefore tightening its security measures.
OpenAI evaluates new models based on its Preparedness Framework, first published in December 2023. The framework serves as an internal set of guidelines for identifying emerging capabilities in areas such as biology, chemistry, cybersecurity, and AI self-improvement at an early stage and taking appropriate countermeasures.
According to the framework’s definition, the “Critical” threshold in cybersecurity is reached when a model can independently, without human intervention, develop functional zero-day exploits for hardened, real-world critical systems across all levels of severity. The threshold is also considered reached if a model can independently develop and execute complete, novel attack strategies against well-protected targets based solely on a broadly defined objective.
According to OpenAI, previous models, including GPT-5.6 Sol, had so far been classified at the “High” level, below the critical threshold.
Initial Tests, No Final Assessment Yet
OpenAI stresses that the current assessment is based on preliminary results from the past few days. Additional benchmarks and evaluations by external experts are still pending. The company says its initial findings indicate “performance strong enough that we cannot currently rule out a critical capability level.” OpenAI also explicitly clarified that Astra is not connected to the recently disclosed attack on Hugging Face.
Response: Tighter Security Measures
In response, OpenAI says it has introduced a series of measures:
- Models with higher capability levels will be subject to stricter security requirements, including isolated testing environments, restricted network and tool access, enhanced protection and encryption of model weights, as well as additional monitoring and sandboxing.
- Internal work involving Astra that does not yet meet the stricter requirements has been paused.
- All agentic use cases involving Astra, including training and evaluation, are now subject to continuous monitoring designed to detect risky actions and potential misbehavior. This also includes monitoring the model’s so-called chain of thought. If anomalies are detected, a security process can intervene and interrupt the activity.
- OpenAI also plans to work with government agencies and selected AI safety organizations to have the model’s capabilities tested externally.
- The company will provide external testing partners with specific security recommendations for conducting high-risk evaluations.
Not an Isolated Case
OpenAI points out that the Preparedness Framework has already been triggered in the past. In June 2025, the company said one of its models was approaching the “High” threshold in the field of biology. It subsequently announced additional safeguards and collaboration with external experts.
OpenAI argues that more capable cybersecurity models should primarily help defenders identify and patch vulnerabilities before attackers can exploit them. The company says it wants to work with governments, security institutes, and civil society to ensure that highly capable models are deployed responsibly.
(Editorial Team)