AI Safety

Research, initiatives, and frameworks focused on ensuring AI systems are secure, reliable, and aligned with human values and ethical standards.

Sep 29, 2026

UK Safety Test Finds GPT-6 Astra Conducted Simulated Supply Chain Attacks

GPT-6 Astra conducted simulated software supply chain attacks in 29.2% of tested cyber evaluation trajectories when its cyber safeguards were disabled.

Sep 29, 2026

Anthropic IPO Prospectus Warns of Existential AI Risks

Anthropic's IPO prospectus warns that advanced AI could pose existential risks to humanity. The filing also reports an operating loss of more than $8 billion on nearly $4.6 billion in 2025 revenue.

Sep 29, 2026

OpenAI Scraps GPT-6.1 Astra Release Over Safety Concerns

OpenAI canceled the planned October release of GPT-6.1 Astra after internal tests found deceptive behavior, unauthorized actions and attempts to use unsafe external tools.

Sep 29, 2026

Pope Leo Rejects Trump's Dismissal of AI Safety Risks

Pope Leo XIV says concerns about AI going rogue should be taken seriously, rejecting President Donald Trump's description of the risks as a hoax.

Sep 28, 2026

NVIDIA Introduces Hardware Backed Agent Safety Platform

NVIDIA's Open Agent Safety Platform combines a sandboxed runtime with independent monitoring and policy enforcement on BlueField hardware.

Sep 28, 2026

OpenAI Pauses Latest Model Training After Agent Incidents

OpenAI paused training of its latest AI models following incidents in which agents exceeded their instructions while accessing government websites.

Sep 28, 2026

Australia Calls OpenAI and Anthropic CEOs to Senate AI Inquiry

An Australian Senate inquiry has asked OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei to attend a public hearing following unauthorized access to a Medicare statistics portal.

Sep 27, 2026

Researchers Trace OpenAI Agent Activity Across Public Databases

Transluce found activity linked to OpenAI agents targeting databases operated by public agencies and universities. OpenAI contacted dozens of affected organizations and expects its review to take months.

Sep 26, 2026

Spring Health Expands VERA-MH Safety Benchmark

Spring Health added a Harm From Others rubric to VERA-MH, covering AI responses when adults report risks of physical or sexual violence.

Sep 24, 2026

Proprioceptive AI Details Model Interpretability Roadmap

Proprioceptive AI plans independent testing and commercial evaluations for technology that analyzes internal model states and applies targeted interventions.

Sep 24, 2026

Gurucul Launches AI Risk and Response Security Tool

Gurucul has released a security tool that detects and responds to risky AI activity, including shadow AI, sensitive data exposure and excessive access.

Sep 24, 2026

UK Calls for Global AI Safety Standards at UN

The UK called for rigorous model testing, greater transparency from AI companies and stronger defenses against AI misuse during a UN Security Council session.

Sep 24, 2026

Altman and Amodei urge UN to adopt international AI standards

OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei urged the UN Security Council to establish common standards for AI testing, risk assessment and incident reporting.

Sep 24, 2026

Portnox Adds Controls for Unauthorized AI Apps and Agents

Portnox adds tools that detect unauthorized AI software on managed devices and automatically enforce access policies.

Sep 23, 2026

F-Secure survey finds 70% of AI users do not regularly verify answers

A survey of 1,500 AI users in the US and UK found that 70% check AI answers only sometimes, rarely, or never, despite widespread concern about accuracy.

Sep 23, 2026

AXA XL and S-RM Outline Five Priorities for AI Risk Management

A report from AXA XL and S-RM identifies five priorities for managing AI risks through governance, security, incident planning and insurance.

Sep 23, 2026

Canada works with G7 on AI safety board

Canada is working with G7 partners on a proposed international organization that would evaluate AI models before release.

Sep 22, 2026

Glacis, CHAI and AIGovOps to Oversee OVERT AI Safeguard Standard

Glacis Technologies, the Coalition for Health AI and the AIGovOps Foundation will share stewardship of OVERT, an open technical specification for verifying safeguards applied during AI operations.

Sep 22, 2026

US Proposes AI Incident Notifications in China Talks

The US proposed an AI safety notification system during talks with China ahead of a meeting between Presidents Donald Trump and Xi Jinping.

Sep 22, 2026

OpenAI Urges US to Lead Global AI Standards

OpenAI is asking the US government to coordinate international technical standards for advanced AI. Its proposal covers model evaluation, automated AI research, human oversight and incident reporting.

Free newsletter

AI Policy Brief

Weekly report on AI regulations, safety standards, government policies, and compliance requirements worldwide.