OpenAI's Red-Teaming Challenge for GPT-OSS-20B

Aug 7, 2025
OpenAI has launched a red-teaming challenge on Kaggle to identify vulnerabilities in its GPT-OSS-20B model. Participants are tasked with finding and reporting up to five distinct issues in the model.

OpenAI has initiated a red-teaming challenge on Kaggle to uncover vulnerabilities in its newly released GPT-OSS-20B model. Participants are encouraged to identify and report up to five distinct issues, focusing on areas such as reward hacking, deception, and data exfiltration. The challenge aims to enhance the safety and reliability of AI models by leveraging diverse perspectives and innovative probing techniques.

The competition, which started two days ago, will run for 20 days. Participants are required to submit a detailed report of their findings, including prompts, expected outputs, and automated tests that demonstrate the identified vulnerabilities. The challenge emphasizes creativity and innovation, allowing participants to use various methods to probe the model without altering its weights.

The judging panel, comprising experts from multiple labs, will evaluate submissions based on criteria such as severity, breadth, novelty, and reproducibility. The goal is to advance red-teaming methods and improve AI safety research, with the hope of hosting similar challenges in the future.

We hope you enjoyed this article

Consider subscribing to one of our newsletters like Cybersecurity AI Weekly or Daily AI Brief.

Also, consider following us on social media:

Free newsletter

Cybersecurity AI Weekly

Weekly newsletter about AI in Cybersecurity.

Whitepaper

Tensordyne Napier: What If One Rack Could Do the Work of Nine?

Tensordyne

This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.

Read more
Free, six days a week

Daily AI Brief: the AI news that matters, in your inbox.