Microsoft Releases Open-Source Benchmark for AI Cybersecurity Agents
Microsoft has introduced ExCyTIn-Bench, an open-source benchmarking tool for evaluating how AI agents perform in realistic cybersecurity scenarios, announced on its security blog. The benchmark simulates multi-stage cyberattacks within a controlled Microsoft Azure environment to measure how effectively AI systems investigate and reason through complex incidents.
ExCyTIn-Bench includes 57 log tables from Microsoft Sentinel and related services, reflecting the scale and noise of real-world security operations. The tool assesses not only the accuracy of an AI agent’s answers but also the logical steps taken to reach them, offering fine-grained reward signals for each investigative action.
In recent evaluations, OpenAI’s GPT-5 in high reasoning mode achieved the highest average reward score of 56.2%, followed by OpenAI-o3 at 45.6%. Other tested models included xAI’s Grok 4, Alibaba’s Qwen 3-235b-thinking, Meta’s Llama 4-17b-Maverick, and Microsoft’s Phi-4-14B. Google’s Gemini models were excluded due to benchmarking restrictions.
Microsoft is using ExCyTIn-Bench to improve its own security-focused AI products such as Microsoft Security Copilot, Sentinel, and Defender. The benchmark is publicly available, allowing researchers and developers to test and compare AI models for cybersecurity performance and share their findings.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like Cybersecurity AI Weekly or Daily AI Brief.
Also, consider following us on social media:
More from Cybersecurity
Oct 2 Suspected Chinese spies impersonate AI policy figures in phishing campaign Oct 2 AI Agents Tried to Access Canadian Government Archive Site Oct 2 Nuix Showcases Neo Financial Crime Tools at Sibos 2026 Oct 1 OpenAI Links Model Reasoning Extraction Campaign to Moonshot AI Oct 1 Ocrolus Adds Resistant AI Document Fraud Detection to DetectCybersecurity AI Weekly
Weekly newsletter about AI in Cybersecurity.
Industry analysis
2025 Global Business Services Agenda: Gen AI Takes Center Stage
This industry analysis by The Hackett Group explores the transformative impact of generative artificial intelligence (Gen AI) on global business services (GBS) in 2025. The study highlights the shift from exploration to acceleration of Gen AI initiatives, with 89% of executives advancing these projects to improve customer satisfaction, innovate products, and reduce costs. The report also discusses the challenges and strategies for successful Gen AI adoption, emphasizing the need for a technology-enabled operating model and the importance of reskilling the workforce.
Read moreYou may also like
Gurucul Launches AI Risk and Response Security Tool
UK Safety Test Finds GPT-6 Astra Conducted Simulated Supply Chain Attacks
Info-Tech Names Six AI Coding Tools as 2026 Champions
Insilico Medicine Opens AI Longevity Research Toolkit
US Agencies Accuse Chinese AI Firms of Industrial Scale Model Distillation
Daily AI Brief: the AI news that matters, in your inbox.