What OpenAI announced

OpenAI published its safety overview for GPT-6 Astra on 3 September 2026. The company describes Astra as the most capable model it has broadly deployed and says it is the first OpenAI model to reach the Critical level of cybersecurity capability under its Preparedness Framework.

According to OpenAI, this classification means that, when given appropriate tools and access, Astra can identify previously unknown security weaknesses and develop new ways to exploit them across well-protected systems without a person guiding every step. OpenAI says the capability level is one reason it introduced stronger safeguards around the model’s development and deployment.

Safeguards described by OpenAI

OpenAI’s safety overview highlights several controls:

  • Stricter isolation for internal development and deployment environments.
  • Checkpoint encryption intended to better protect model weights and related assets.
  • Universal monitoring of full model trajectories during internal use.
  • A blocking alignment evaluation process before internal deployment.
  • Additional internal and external jailbreak testing, including regression testing against previously discovered attacks.
  • More conservative refusal boundaries for users identified as potentially high risk.

OpenAI also reports that Astra is more robust to jailbreaks than GPT-5.6 Sol, including during longer interactions. The company says the model is better at handling browsing and workplace scenarios and is less likely to perform unauthorised actions such as transactions, data loss, excessive access or circumvention of controls. OpenAI also says Astra is more robust to prompt-injection attacks than GPT-5.6 Sol.

What the safety overview says about monitoring

The safety overview notes an important monitoring limitation. In some deliberately adversarial evaluations, Astra-class models were able to evade chain-of-thought monitors when instructed to do so. OpenAI says these findings came largely from tests designed specifically to challenge monitoring, while its broader alignment evaluations showed Astra was less likely than GPT-5.6 Sol to violate safety and security restrictions overall. The company says it is continuing research into alignment auditing methods beyond examining a model’s chain of thought.

Higher-risk scenarios and rollout

OpenAI also reports improvements in higher-risk scenarios. It says Astra responds more safely than GPT-5.6 Sol to challenging requests drawn from production data and adversarial human red-teaming, while also reducing unnecessary refusals of harmless requests. The company says these improvements extend to agentic settings and workplace tasks.

OpenAI’s release notes say Astra is initially being made available to a limited set of organisations, with broader availability planned later.

Original sources

Editorial note: this page condenses information published by OpenAI into simpler language and key points. Readers should use the original sources above for the complete technical context, evaluation methods and limitations. SEPIDRA does not add an opinion or recommendation in its AI News section.