AE Studio Research Cited by Anthropic CEO in Open Weight Safety Post
AE Studio said research with Anthropic was cited by Dario Amodei as a possible method for making open weight AI models safer, announced in a press release.
The research focuses on modular training, which separates specific categories of dangerous knowledge into removable modules. AE Studio said the method can keep information such as advanced virology and cyberattack techniques out of a released model while keeping other capabilities intact.
In tests, a single model trained with the method matched the performance of multiple models trained separately from scratch at each evaluated scale. AE Studio said the released model behaved as if it had not learned the removed information, including when simulated attackers tried to train the knowledge back in.
The work was led by AE Studio researchers Ethan Roland, Murat Cubuktepe, and Erick Martinez, in collaboration with Anthropic researchers. The findings are available through Anthropic's alignment research site.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like AI Policy Brief or Daily AI Brief.
Also, consider following us on social media:
More from AI Safety
Sep 19 European AI Firms Reject Calls to Slow Model Development Sep 19 Claude Leads 26% of Anthropic Model Research Work Sep 19 Sam Altman to Brief UN Security Council on AI Risks Sep 19 Newsom Orders Work on AI Kill Switch and Faster Oversight Sep 19 Josh Shapiro Calls for Federal AI GuardrailsAI Policy Brief
Weekly report on AI regulations, safety standards, government policies, and compliance requirements worldwide.
Industry analysis
2025 Global Business Services Agenda: Gen AI Takes Center Stage
This industry analysis by The Hackett Group explores the transformative impact of generative artificial intelligence (Gen AI) on global business services (GBS) in 2025. The study highlights the shift from exploration to acceleration of Gen AI initiatives, with 89% of executives advancing these projects to improve customer satisfaction, innovate products, and reduce costs. The report also discusses the challenges and strategies for successful Gen AI adoption, emphasizing the need for a technology-enabled operating model and the importance of reskilling the workforce.
Read moreYou may also like
Anthropic and OpenAI leave AI evaluator access details open
OpenAI Executive Warns of Persistent AI Cyber Attacks
Anthropic Attributes Its Largest Measured Distillation Campaign to Alibaba
Anthropic Picks Accenture for AI Safety Testing
Anthropic's Threat Report Finds AI Moving From Assistant to Orchestrator
Daily AI Brief: the AI news that matters, in your inbox.