Pearl Finds AI Models Match Expert Judgment Only 70% of the Time
Pearl Enterprise evaluated 25 AI models from OpenAI, Anthropic, Google DeepMind, Microsoft, and others, finding that even top systems align with expert judgment only about 70% of the time, with performance dropping to as low as 20% in some domains.