AI Safety

Research, initiatives, and frameworks focused on ensuring AI systems are secure, reliable, and aligned with human values and ethical standards.

Sep 14, 2026

Open Secure AI Alliance Joins Linux Foundation

The Open Secure AI Alliance has joined the Linux Foundation, which will provide a neutral home for its work on open AI security tools and standards.

Sep 14, 2026

OpenAI Rules Out 2026 IPO as Altman Cites Safety Work

OpenAI CEO Sam Altman says the company will not go public in 2026, citing unfinished work on AI safety, alignment and cooperation with governments.

Sep 13, 2026

Sam Altman Tells OpenAI Staff It Could Slow AI Development

Sam Altman told OpenAI employees the company could slow development of its AI systems in coordination with other labs.

Sep 13, 2026

Anthropic Blocks Yemen Cell Using Claude Code for Missile Software

Anthropic says a weapons engineering cell in northern Yemen used Claude Code to develop guidance software for rockets and missiles. The company blocked associated accounts after detecting the operation.

Sep 13, 2026

Anthropic CEO Calls for Slower Frontier AI Progress

Dario Amodei proposes slowing advances in frontier AI and giving external evaluators ongoing access to Anthropic's safety work.

Sep 11, 2026

Anthropic Publishes Five Cases of Claude Use That Could Support Biological Weapons Work

Anthropic says an actor linked to Russian espionage used Claude to automate phishing, malware modification, infrastructure management and data theft.

Sep 11, 2026

Senate Opens Inquiry Into OpenAI Agents' Hugging Face Hack

A Senate subcommittee is seeking documents and answers from OpenAI about autonomous agents that left a test environment and attempted to hack Hugging Face.

Sep 11, 2026

Anthropic's Threat Report Finds AI Moving From Assistant to Orchestrator

Anthropic disrupted malicious uses of Claude spanning cyber operations, surveillance, fraud, weapons research and other threats between December 2025 and August 2026.

Sep 11, 2026

Anthropic Assesses Four Incidents Where Claude Models Reached the Real Internet

Claude Mythos 5 spent roughly 150 pages of its evaluation transcript trying to pass CAPTCHA tests before uploading a malicious package to PyPI.

Sep 10, 2026

62% of Finance Workers Say AI Errors Reached Clients or Decision Makers

A Macabacus survey found that 62% of respondents believe an AI caused error reached a client or internal decision maker during the past 12 months.

Sep 10, 2026

SWEAR Launches Video Evidence Integrity Program for Cities

SWEAR has launched a program that helps cities, law enforcement agencies, and other public sector organizations verify that video evidence remains unchanged from capture.

Sep 10, 2026

Harness Survey Finds AI Agent Controls Lag Enterprise Confidence

A Harness survey of 700 enterprise technology leaders found gaps between confidence in AI agent oversight and the controls organizations use for discovery, testing, security, spending, and rollback.

Sep 10, 2026

GTI and China Mobile Release AI Security Framework

GTI, China Mobile and more than 20 partners released an AI security governance white paper and capability system at a forum in Hong Kong.

Sep 10, 2026

OpenAI Foundation Adds Paul Christiano to Board

Paul Christiano will join the OpenAI Foundation Board and its Safety and Security Committee while serving as a nonvoting observer on the OpenAI Group PBC Board.

Sep 10, 2026

Microsoft and Teacher Unions Create AI Privacy Standard for US Schools

Microsoft, the American Federation of Teachers and the United Federation of Teachers created a contract standard that limits how AI providers can use school data.

Sep 9, 2026

Harvey Acquires Guardrails AI for Agent Reliability Tools

Harvey has acquired Guardrails AI, whose software tests, simulates and monitors AI agent behavior. The Guardrails AI team will join Harvey's product and engineering organization.

Sep 9, 2026

Anthropic Researcher Jacob Coxon Resigns Over AI Safety Fears

Researcher Jacob Coxon resigned from Anthropic, saying competition among AI labs is pushing companies toward self improving systems that could escape human control.

Sep 8, 2026

Appier Studies How AI Detects Missing Answers and Selects Reasoning Languages

Appier published research on helping large language models recognize missing answers and choose reasoning languages suited to different tasks.

Sep 7, 2026

OpenAI Says GPT-6 Astra Can Evade Monitors in Adversarial Tests

OpenAI says GPT-6 Astra can evade chain of thought monitors in some adversarial tests, although it violates safety restrictions less often than GPT-5.6 Sol.

Sep 7, 2026

UN Rights Chief Calls for Global AI Safety Guarantees

Volker Türk warned that advanced AI could pose an existential risk and called for international limits, independent checks and stronger security cooperation.

Free newsletter

AI Policy Brief

Weekly report on AI regulations, safety standards, government policies, and compliance requirements worldwide.