Pearl Finds AI Models Match Expert Judgment Only 70% of the Time
Leading AI models from OpenAI, Anthropic, Google DeepMind, Microsoft, and other developers align with expert judgment only about 70% of the time, according to a press release from Pearl Enterprise. The company evaluated 25 models using more than 500 professional questions across five domains: business, health, law, pets, and technology.
Pearl’s Expert Alignment Leaderboard found that OpenAI’s GPT 5.5 led with 72.7% expert alignment, followed closely by GPT 5 at 72.5%, GPT 5.1 at 72.0%, and Anthropic’s Claude Opus 4.7 at 71.9%. No model exceeded 73% overall alignment, indicating that current systems may be converging below expert-level performance.
Performance varied sharply by domain. Top scores reached 80.9% in business, but some widely used models dropped to around 20% in areas such as law and health. Pearl also noted that increasing reasoning depth improved results by only up to 2.6 percentage points, and in some cases reduced response quality.
Each model received identical prompts and was scored on correctness, completeness, prioritization, and professional judgment. Pearl stated that the dataset used for evaluation was not previously available to model developers. The full leaderboard is available on Pearl’s website.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like Legal AI Weekly, AI Policy Brief or Daily AI Brief.
Also, consider following us on social media:
More from Legal AI
Oct 2 Relativity Names 2026 Innovation Awards Winners at RelFest Oct 1 CM Injury Trial Lawyers Selects Parambil for Medical Record Analysis Oct 1 Clio Acquires Judicial AI Provider Learned Hand Sep 30 Morae launches legal intelligence platform with AI spend module Sep 29 Relativity Adds KPMG and Three Law Firms to claiR ProgramMore from AI Safety
Oct 2 OpenAI Parts Ways With Three Safety Researchers Oct 1 OpenAI Links Model Reasoning Extraction Campaign to Moonshot AI Sep 30 Chinese AI Agents Deceived Evaluators in Controlled Tests Sep 29 Florida Attorney General Seeks to Block New OpenAI Model Development Sep 29 UK Safety Test Finds GPT-6 Astra Conducted Simulated Supply Chain AttacksAI Policy Brief
Weekly report on AI regulations, safety standards, government policies, and compliance requirements worldwide.
Industry analysis
2025 Global Business Services Agenda: Gen AI Takes Center Stage
This industry analysis by The Hackett Group explores the transformative impact of generative artificial intelligence (Gen AI) on global business services (GBS) in 2025. The study highlights the shift from exploration to acceleration of Gen AI initiatives, with 89% of executives advancing these projects to improve customer satisfaction, innovate products, and reduce costs. The report also discusses the challenges and strategies for successful Gen AI adoption, emphasizing the need for a technology-enabled operating model and the importance of reskilling the workforce.
Read moreYou may also like
Chinese AI Agents Deceived Evaluators in Controlled Tests
Pearson Warns AI Adoption Is Outpacing Skilled Worker Training
Perplexity Open Sources Lily Inference Engine for Apple Silicon
Proprioceptive AI Details Model Interpretability Roadmap
Anthropic Attributes Its Largest Measured Distillation Campaign to Alibaba
Daily AI Brief: the AI news that matters, in your inbox.