OpenAI Releases Open-Weight Safety Reasoning Models for Developers
OpenAI has released a research preview of gpt-oss-safeguard, a pair of open-weight reasoning models for safety classification tasks, announced on its website. The models—gpt-oss-safeguard-120b and gpt-oss-safeguard-20b—are fine-tuned versions of the gpt-oss open models and are distributed under the Apache 2.0 license for free use and modification.
The gpt-oss-safeguard models use reasoning to interpret developer-defined safety policies at inference time, classifying user messages and chat content based on those policies. This allows developers to adjust or replace policies without retraining the model. The models also provide a chain-of-thought output, enabling developers to review the reasoning behind each classification.
According to OpenAI, this approach differs from traditional safety classifiers that rely on large labeled datasets. Instead, developers supply their own policy text, and the model generalizes from it to produce explainable results. The models are designed for use cases where safety policies must evolve quickly or where labeled data is limited.
The release was developed in collaboration with ROOST, which is launching a model community to support open safety research. The models can be downloaded from Hugging Face, and OpenAI has also published a technical report detailing performance evaluations against internal and external benchmarks. Early testing with partners such as SafetyKit, ROOST, and Discord helped refine the tools for community use.
With gpt-oss-safeguard, OpenAI extends its internal Safety Reasoner framework to the public, offering a flexible, reasoning-based method for developers to define and enforce their own safety boundaries.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like AI Policy Brief or Daily AI Brief.
Also, consider following us on social media:
More from AI Safety
Oct 5 Altman Says AI Benefits Justify Accepting Some Risks Oct 2 OpenAI Parts Ways With Three Safety Researchers Oct 1 OpenAI Links Model Reasoning Extraction Campaign to Moonshot AI Sep 30 Chinese AI Agents Deceived Evaluators in Controlled Tests Sep 29 Florida Attorney General Seeks to Block New OpenAI Model DevelopmentAI Policy Brief
Weekly report on AI regulations, safety standards, government policies, and compliance requirements worldwide.
Industry analysis
2025 Global Business Services Agenda: Gen AI Takes Center Stage
This industry analysis by The Hackett Group explores the transformative impact of generative artificial intelligence (Gen AI) on global business services (GBS) in 2025. The study highlights the shift from exploration to acceleration of Gen AI initiatives, with 89% of executives advancing these projects to improve customer satisfaction, innovate products, and reduce costs. The report also discusses the challenges and strategies for successful Gen AI adoption, emphasizing the need for a technology-enabled operating model and the importance of reskilling the workforce.
Read moreYou may also like
Gensyn Releases Auditable open-1b AI Model
OpenAI Publishes Model Misalignment Reporting Framework
OpenAI Previews Decisions API for Faster Model Choices
OpenAI Scraps GPT-6.1 Astra Release Over Safety Concerns
OpenAI Chief Scientist Calls for Slower AI Scaling
Daily AI Brief: the AI news that matters, in your inbox.