Atomicwork and New Measure Release ITSMBench for AI Service Desk Tests

August 12, 2026
Atomicwork and New Measure released ITSMBench, an open source benchmark for testing frontier AI models on enterprise IT service management tasks. Initial results show the best performing model completed 50.56% of typical service desk requests.

In a press release, Atomicwork and New Measure announced ITSMBench, an open source benchmark for testing how frontier AI models handle enterprise IT Service Management work. The benchmark evaluates models and agent frameworks from OpenAI, Anthropic, xAI, GLM, and other developers across service desk tasks.

ITSMBench recreates enterprise environments with 42 mocked software systems, nearly 1,800 database tables, and more than 2,000 REST endpoints. It covers 89 L2 and L3 service desk tasks, including identity and access management, device issues, networking, security operations, infrastructure, and engineering escalations.

Initial results show that the best performing model completed 50.56% of requests typically handled by enterprise IT service teams. Grok 4.5 found the right APIs 83.0% of the time but performed worse on task completion, while Opus 5 completed 63.5% of tasks after the correct tools were identified. GPT-5.6 Sol placed between those models on discovery and execution, and GLM-5.2 cost $0.25 per trial but trailed Opus 5 and GPT-5.6 Sol on execution.

The benchmark also found that agent frameworks affected accuracy and cost. Models often stopped after finding a plausible explanation instead of verifying the root cause, and sometimes left related systems, records, or follow up actions incomplete. The benchmark, methodology, results, and environments are available at atomicwork.com/itsm-bench.

We hope you enjoyed this article.

Consider subscribing to one of our newsletters like Enterprise AI Brief or Daily AI Brief.

Also, consider following us on social media:

Subscribe to Enterprise AI Brief

Weekly report on AI business applications, enterprise software releases, automation tools, and industry implementations.

Market report

AI’s Time-to-Market Quagmire: Why Enterprises Struggle to Scale AI Innovation

ModelOp

The 2025 AI Governance Benchmark Report by ModelOp provides insights from 100 senior AI and data leaders across various industries, highlighting the challenges enterprises face in scaling AI initiatives. The report emphasizes the importance of AI governance and automation in overcoming fragmented systems and inconsistent practices, showcasing how early adoption correlates with faster deployment and stronger ROI.

Read more