Microsoft Releases Open-Source Benchmark for AI Cybersecurity Agents

Oct 20, 2025
Microsoft has launched ExCyTIn-Bench, an open-source benchmark designed to test AI agents on realistic cybersecurity investigations using simulated multi-stage attacks and data from Microsoft Sentinel.

Microsoft has introduced ExCyTIn-Bench, an open-source benchmarking tool for evaluating how AI agents perform in realistic cybersecurity scenarios, announced on its security blog. The benchmark simulates multi-stage cyberattacks within a controlled Microsoft Azure environment to measure how effectively AI systems investigate and reason through complex incidents.

ExCyTIn-Bench includes 57 log tables from Microsoft Sentinel and related services, reflecting the scale and noise of real-world security operations. The tool assesses not only the accuracy of an AI agent’s answers but also the logical steps taken to reach them, offering fine-grained reward signals for each investigative action.

In recent evaluations, OpenAI’s GPT-5 in high reasoning mode achieved the highest average reward score of 56.2%, followed by OpenAI-o3 at 45.6%. Other tested models included xAI’s Grok 4, Alibaba’s Qwen 3-235b-thinking, Meta’s Llama 4-17b-Maverick, and Microsoft’s Phi-4-14B. Google’s Gemini models were excluded due to benchmarking restrictions.

Microsoft is using ExCyTIn-Bench to improve its own security-focused AI products such as Microsoft Security Copilot, Sentinel, and Defender. The benchmark is publicly available, allowing researchers and developers to test and compare AI models for cybersecurity performance and share their findings.

We hope you enjoyed this article

Consider subscribing to one of our newsletters like Cybersecurity AI Weekly or Daily AI Brief.

Also, consider following us on social media:

Free newsletter

Cybersecurity AI Weekly

Weekly newsletter about AI in Cybersecurity.

Industry analysis

2025 Global Business Services Agenda: Gen AI Takes Center Stage

The Hackett Group

This industry analysis by The Hackett Group explores the transformative impact of generative artificial intelligence (Gen AI) on global business services (GBS) in 2025. The study highlights the shift from exploration to acceleration of Gen AI initiatives, with 89% of executives advancing these projects to improve customer satisfaction, innovate products, and reduce costs. The report also discusses the challenges and strategies for successful Gen AI adoption, emphasizing the need for a technology-enabled operating model and the importance of reskilling the workforce.

Read more
Free, six days a week

Daily AI Brief: the AI news that matters, in your inbox.