OpenAI Introduces SWE-Lancer Benchmark for AI in Software Engineering

February 19, 2025
OpenAI has launched SWE-Lancer, a benchmark evaluating AI models on over 1,400 real-world freelance software engineering tasks from Upwork, valued at $1 million in total payouts.
OpenAI Introduces SWE-Lancer Benchmark for AI in Software Engineering
Image: OpenAI

OpenAI has launched SWE-Lancer, a new benchmark designed to evaluate the coding performance of AI models using real-world freelance software engineering tasks from Upwork, valued at a total of $1 million USD in payouts. In a press release, OpenAI detailed that SWE-Lancer includes over 1,400 tasks, ranging from $50 bug fixes to $32,000 feature implementations.

The benchmark not only tests AI models on independent engineering tasks but also on managerial tasks, where models must choose between technical implementation proposals. Independent tasks are graded with end-to-end tests verified by experienced software engineers, while managerial decisions are assessed against the choices of the original hired engineering managers.

Despite the comprehensive nature of SWE-Lancer, current frontier models are still unable to solve the majority of tasks. The best-performing model, Anthropic's Claude 3.5 Sonnet, earned just over $400,000 out of the possible $1 million across all tasks. OpenAI has open-sourced part of the dataset, called SWE-Lancer Diamond, to encourage further research into AI's role in software development.

We hope you enjoyed this article.

Consider subscribing to one of our newsletters like Daily AI Brief.

Also, consider following us on social media:

Subscribe to Daily AI Brief

Daily report covering major AI developments and industry news, with both top stories and complete market updates

Industry analysis

2025 Global Business Services Agenda: Gen AI Takes Center Stage

The Hackett Group

This industry analysis by The Hackett Group explores the transformative impact of generative artificial intelligence (Gen AI) on global business services (GBS) in 2025. The study highlights the shift from exploration to acceleration of Gen AI initiatives, with 89% of executives advancing these projects to improve customer satisfaction, innovate products, and reduce costs. The report also discusses the challenges and strategies for successful Gen AI adoption, emphasizing the need for a technology-enabled operating model and the importance of reskilling the workforce.

Read more