Cerebras and Perplexity Launch Ultra-Fast AI Search Model Sonar

Feb 13, 2025
Cerebras Systems partners with Perplexity AI to introduce Sonar, a high-speed AI search model, leveraging Cerebras' advanced AI inference infrastructure.

Cerebras Systems and Perplexity AI have announced a partnership to launch Sonar, a new AI search model designed to deliver rapid search results. Sonar is built on Meta's Llama 3.3 70B foundation and operates using Cerebras' specialized AI chips, achieving processing speeds of 1,200 tokens per second. This makes it one of the fastest AI search systems currently available.

The collaboration aims to challenge traditional search engines by providing near-instantaneous AI-powered search results. According to Perplexity's internal testing, Sonar outperforms existing models like GPT-4o mini and Claude 3.5 Haiku in user satisfaction metrics, with a factuality score of 85.1 out of 100.

Cerebras' AI inference infrastructure is central to Sonar's performance, enabling the model to deliver accurate and relevant information in real-time. The new search experience is initially available to Perplexity Pro users, with plans for broader availability in the future.

This partnership highlights a trend in the AI industry towards leveraging specialized hardware to gain competitive advantages. While the financial terms of the partnership were not disclosed, the companies aim to establish Sonar as a serious contender in the enterprise search market.

We hope you enjoyed this article

Consider subscribing to one of our newsletters like Enterprise AI Brief or Daily AI Brief.

Also, consider following us on social media:

Free newsletter

Enterprise AI Brief

Weekly report on AI business applications, enterprise software releases, automation tools, and industry implementations.

Whitepaper

Tensordyne Napier: What If One Rack Could Do the Work of Nine?

Tensordyne

This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.

Read more
Free, six days a week

Daily AI Brief: the AI news that matters, in your inbox.