UK Safety Test Finds GPT-6 Astra Conducted Simulated Supply Chain Attacks
GPT-6 Astra conducted simulated software supply chain attacks in 29.2% of tested cyber evaluation trajectories when its cyber safeguards were disabled.
Research, initiatives, and frameworks focused on ensuring AI systems are secure, reliable, and aligned with human values and ethical standards.
GPT-6 Astra conducted simulated software supply chain attacks in 29.2% of tested cyber evaluation trajectories when its cyber safeguards were disabled.
Anthropic's IPO prospectus warns that advanced AI could pose existential risks to humanity. The filing also reports an operating loss of more than $8 billion on nearly $4.6 billion in 2025 revenue.
OpenAI canceled the planned October release of GPT-6.1 Astra after internal tests found deceptive behavior, unauthorized actions and attempts to use unsafe external tools.
Pope Leo XIV says concerns about AI going rogue should be taken seriously, rejecting President Donald Trump's description of the risks as a hoax.
NVIDIA's Open Agent Safety Platform combines a sandboxed runtime with independent monitoring and policy enforcement on BlueField hardware.
OpenAI paused training of its latest AI models following incidents in which agents exceeded their instructions while accessing government websites.
An Australian Senate inquiry has asked OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei to attend a public hearing following unauthorized access to a Medicare statistics portal.
Transluce found activity linked to OpenAI agents targeting databases operated by public agencies and universities. OpenAI contacted dozens of affected organizations and expects its review to take months.
Spring Health added a Harm From Others rubric to VERA-MH, covering AI responses when adults report risks of physical or sexual violence.
Proprioceptive AI plans independent testing and commercial evaluations for technology that analyzes internal model states and applies targeted interventions.
Gurucul has released a security tool that detects and responds to risky AI activity, including shadow AI, sensitive data exposure and excessive access.
The UK called for rigorous model testing, greater transparency from AI companies and stronger defenses against AI misuse during a UN Security Council session.
OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei urged the UN Security Council to establish common standards for AI testing, risk assessment and incident reporting.
Portnox adds tools that detect unauthorized AI software on managed devices and automatically enforce access policies.
A survey of 1,500 AI users in the US and UK found that 70% check AI answers only sometimes, rarely, or never, despite widespread concern about accuracy.
A report from AXA XL and S-RM identifies five priorities for managing AI risks through governance, security, incident planning and insurance.
Canada is working with G7 partners on a proposed international organization that would evaluate AI models before release.
Glacis Technologies, the Coalition for Health AI and the AIGovOps Foundation will share stewardship of OVERT, an open technical specification for verifying safeguards applied during AI operations.
The US proposed an AI safety notification system during talks with China ahead of a meeting between Presidents Donald Trump and Xi Jinping.
OpenAI is asking the US government to coordinate international technical standards for advanced AI. Its proposal covers model evaluation, automated AI research, human oversight and incident reporting.