OpenAI Plans Limited Astra Release After Critical Cybersecurity Rating
OpenAI said Astra is its first model to meet the Critical cybersecurity capability threshold under its Preparedness Framework, detailed in a company blog post. The company said the model can find previously unknown security flaws and develop ways to exploit them across well protected systems with the right tools and access, without a person guiding each step.
OpenAI plans to make Astra available soon, but access to its most advanced cybersecurity capabilities will be limited. Advanced cybersecurity work will first be available to a group of testers, with access through Daybreak Blue to follow for defensive use.
The company said it delayed parts of Astra work while it strengthened protections against cyber misuse and unauthorized model actions. OpenAI said Astra was not involved in the Hugging Face incident, but it used lessons from that incident to add stronger refusals for harmful cyber requests, misuse protections, and monitoring that can stop potentially unauthorized activity.
OpenAI also said some safeguards may slow, pause, or stop legitimate work. In ChatGPT or Codex, users may be asked to review a paused action before continuing, while API tasks will stop if the monitor pauses them.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like Cybersecurity AI Weekly, AI Policy Brief or Daily AI Brief.
Also, consider following us on social media:
More from Cybersecurity
Sep 13 Anthropic Blocks Yemen Cell Using Claude Code for Missile Software Sep 13 Researchers Link OpenAI Agents to May RubyGems Attack Sep 11 Anthropic Publishes Five Cases of Claude Use That Could Support Biological Weapons Work Sep 11 Senate Opens Inquiry Into OpenAI Agents' Hugging Face Hack Sep 11 FAZE Security Emerges From Stealth With $6 Million Seed RoundMore from AI Safety
Sep 13 Sam Altman Tells OpenAI Staff It Could Slow AI Development Sep 13 Anthropic Blocks Yemen Cell Using Claude Code for Missile Software Sep 13 Anthropic CEO Calls for Slower Frontier AI Progress Sep 11 Anthropic Publishes Five Cases of Claude Use That Could Support Biological Weapons Work Sep 11 Senate Opens Inquiry Into OpenAI Agents' Hugging Face HackCybersecurity AI Weekly
Weekly newsletter about AI in Cybersecurity.
Market report
2025 Generative AI in Professional Services Report
Thomson Reuters
This report by Thomson Reuters explores the integration and impact of generative AI technologies, such as ChatGPT and Microsoft Copilot, within the professional services sector. It highlights the growing adoption of GenAI tools across industries like legal, tax, accounting, and government, and discusses the challenges and opportunities these technologies present. The report also examines professionals' perceptions of GenAI and the need for strategic integration to maximize its value.
Read moreYou may also like
OpenAI Says GPT-6 Astra Can Evade Monitors in Adversarial Tests
OpenAI Executive Warns of Persistent AI Cyber Attacks
OpenAI Releases Report on Hugging Face Breach
AWS Adds GPT-6 Astra to Amazon Bedrock
Senate Opens Inquiry Into OpenAI Agents' Hugging Face Hack
Daily AI Brief: the AI news that matters, in your inbox.