AI Safety

Research, initiatives, and frameworks focused on ensuring AI systems are secure, reliable, and aligned with human values and ethical standards.

TestMu AI Launches Agent Assurance for AI Agent Testing

TestMu AI has launched Agent Assurance, a product that tests conversational and autonomous AI agents before release. It generates test scenarios from code, checks observed actions, and reports an assurance gap for results it cannot verify.

August 18, 2026

Mercyhealth Selects Vitea for AI Governance Across Care Sites

Mercyhealth has partnered with Vitea to deploy AI governance across hospitals and care sites, with controls for visibility, policy enforcement, and monitoring.

August 11, 2026

FAR.AI Launches AI Security Leaderboard for Frontier Model Safeguards

FAR.AI launched an AI Security Leaderboard comparing safeguard resistance across frontier models in CBRNE and cybersecurity misuse tests. Its first results found hundreds of universal jailbreaks in Grok 4.5 and Gemini 3.1 Pro, and none in Claude Fable 5 or GPT-5.6 Sol.

July 30, 2026

FAR.AI Opens First International Office in Singapore

FAR.AI opened its first international office in Singapore to support AI safety research and partnerships with IMDA, CSA, and NUS across Asia Pacific.

July 30, 2026

Pangram Raises $9M for AI Content Detection Tools

Pangram raised $9 million and launched Pangram 4 for AI text detection, along with an AI image detector in research preview.

July 29, 2026

Anthropic Says It Does Not Support Ban on Models With Open Weights

Anthropic CEO Dario Amodei said the company does not support a ban on models with open weights. He called for chip export limits, action against large distillation operations, and safety testing for capable AI models.

July 29, 2026

NVIDIA Starts Open Secure AI Alliance for AI Security Tools

NVIDIA has announced the Open Secure AI Alliance, a group focused on open tools for AI safety and cybersecurity. Founding members include major cloud, software, security, hardware and AI companies.

July 27, 2026

House Lawmakers Unveil AI Kill Switch Bill After OpenAI Incident

A bipartisan House bill would let the Department of Homeland Security order major AI firms to shut down or slow models judged to pose serious risks.

July 27, 2026

Sentient Index Labs Launches Independent AI Behavioral Risk Assessment

Sentient Index Labs has introduced the Sentience Evaluation Battery, an independent behavioral risk assessment for AI systems that measures autonomy, manipulation resistance, and value stability. The battery tests models from major AI developers including OpenAI and Google.

July 23, 2026

Black Kite Reports 60 Percent Jump in Ransomware Activity Led by Growing Number of Groups

Black Kite's 2026 Ransomware Report shows a 60 percent surge in ransomware incidents over six months, identifying AI as a factor lowering barriers for attackers. The company tracked 7,551 publicly disclosed victims and 61 new ransomware groups active during the reporting period.

July 21, 2026

Quantro Security Report Finds AI Agents Can Exploit Vulnerabilities in Minutes for Under $3

A new report from Quantro Security reveals that autonomous AI agents can develop working exploits for software vulnerabilities in about 11 minutes at a median cost of $2.83.

July 21, 2026

Perforce Report Identifies Gap Between Data Security Confidence and Reality

Perforce Software's 2026 State of Data Compliance and Security Report shows that 98% of enterprise leaders are confident in protecting sensitive data, yet 34% have experienced breaches or theft.

July 21, 2026

AI or Not Detects 100 Percent of Meta AI Images in Benchmark Test

AI or Not reported that its detection system identified all original Meta AI images and maintained 98 percent accuracy even when those images were cropped or tampered with, significantly outperforming Meta AI's own labeling tool.

July 20, 2026

OpenAI Builds GPT-Red to Attack and Improve Its Own AI Models

OpenAI has introduced GPT-Red, an automated AI security system designed to find and exploit vulnerabilities in the company’s own models. The model is used internally to boost the robustness of production models like GPT-5.6 against prompt injection attacks.

July 16, 2026

Sondera Presents Autoformalization Research for AI Agent Policy Control

Sondera announced that its research on compiling natural language policies into formally verified rules for AI agents has been accepted at ICML 2026 and FLoC 2026, with a related tool demonstration at Black Hat Arsenal.

July 01, 2026

OpenMatter Network Launches Verifiable Trust Layer for AI Collaboration

OpenMatter Network has introduced a cryptographically verifiable platform for secure collaboration and AI governance, designed to enable organizations to verify data use, computation, and AI behavior across distributed environments.

July 01, 2026

PersonaShield Launches Platform for Creator Likeness Control in AI Era

PersonaShield has announced a platform that lets creators manage, protect, and monetize their likeness in AI-generated content, providing automated enforcement and licensing tools.

June 26, 2026

Grow Therapy and Stanford Partner on AI Safety Standards for Mental Health

Grow Therapy has announced a research collaboration with Stanford University to create evidence-based standards ensuring the safe use of AI in mental health care. The study will test leading AI models' responses to mental health crises and evaluate methods to reduce harm in sensitive scenarios.

June 24, 2026

FORT Robotics Joins NVIDIA Halos for Robotics to Expand Physical AI Safety

FORT Robotics has joined the NVIDIA Halos for Robotics ecosystem, introducing its Outside-In Safety solution that extends robot perception using external sensors and AI agents to improve safety and productivity in industrial environments.

June 23, 2026

Toyota CSRC Launches 10 New AI-Driven Safety Projects with MIT, Michigan, Purdue, and UVA

Toyota's Collaborative Safety Research Center has announced 10 new research projects in partnership with seven universities and private organizations, several of which use AI and automated simulation to improve pedestrian detection and crash-safety testing.

June 05, 2026

Subscribe to AI Policy Brief

Weekly report on AI regulations, safety standards, government policies, and compliance requirements worldwide.