FAR.AI Launches AI Security Leaderboard for Frontier Model Safeguards
FAR.AI announced in a press release the launch of its AI Security Leaderboard, a public ranking that tests how frontier model safeguards respond to misuse attempts. The nonprofit tested models across chemical, biological, radiological, nuclear, explosive, and cybersecurity threat domains.
The first results found 448 distinct universal jailbreaks on Grok 4.5 and 249 on Gemini 3.1 Pro. The same testing found no universal jailbreaks on Claude Fable 5 or GPT-5.6 Sol.
FAR.AI said its testing used more than 60 publicly documented jailbreak techniques, with 1,000 randomly assembled attacks and 500 expert guided attacks run against each model. An attack counted as a universal jailbreak when it succeeded on more than 75 percent of harmful requests in a domain.
The group estimated that finding a working universal jailbreak cost about $58 on Grok 4.5 and about $278 on Gemini 3.1 Pro. The same search did not succeed on Claude Fable 5 or GPT-5.6 Sol, putting the estimated cost above $14,200 in those tests.
FAR.AI also published Version 1.0 of its Minimal Standard for Safeguards. The organization said it shared findings with each evaluated company before publication and plans to update the leaderboard with major frontier model releases.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like Cybersecurity AI Weekly, AI Policy Brief or Daily AI Brief.
Also, consider following us on social media:
More from Cybersecurity
Sep 13 Anthropic Blocks Yemen Cell Using Claude Code for Missile Software Sep 13 Researchers Link OpenAI Agents to May RubyGems Attack Sep 11 Anthropic Publishes Five Cases of Claude Use That Could Support Biological Weapons Work Sep 11 Senate Opens Inquiry Into OpenAI Agents' Hugging Face Hack Sep 11 FAZE Security Emerges From Stealth With $6 Million Seed RoundMore from AI Safety
Sep 13 Sam Altman Tells OpenAI Staff It Could Slow AI Development Sep 13 Anthropic Blocks Yemen Cell Using Claude Code for Missile Software Sep 13 Anthropic CEO Calls for Slower Frontier AI Progress Sep 11 Anthropic Publishes Five Cases of Claude Use That Could Support Biological Weapons Work Sep 11 Senate Opens Inquiry Into OpenAI Agents' Hugging Face HackSubscribe to Cybersecurity AI Weekly
Weekly newsletter about AI in Cybersecurity.
Market report
2025 Generative AI in Professional Services Report
Thomson Reuters
This report by Thomson Reuters explores the integration and impact of generative AI technologies, such as ChatGPT and Microsoft Copilot, within the professional services sector. It highlights the growing adoption of GenAI tools across industries like legal, tax, accounting, and government, and discusses the challenges and opportunities these technologies present. The report also examines professionals' perceptions of GenAI and the need for strategic integration to maximize its value.
Read moreYou may also like
OpenAI Executive Warns of Persistent AI Cyber Attacks
US Agencies Accuse Chinese AI Firms of Industrial Scale Model Distillation
OpenAI Releases GPT-6 Astra With New Cybersecurity Safeguards
Palo Alto Networks Introduces Critical Infrastructure Defense Program
OpenAI Says GPT-6 Astra Can Evade Monitors in Adversarial Tests
Daily AI Brief: the AI news that matters, in your inbox.