Cerebras and Perplexity Launch Ultra-Fast AI Search Model Sonar
Cerebras Systems and Perplexity AI have announced a partnership to launch Sonar, a new AI search model designed to deliver rapid search results. Sonar is built on Meta's Llama 3.3 70B foundation and operates using Cerebras' specialized AI chips, achieving processing speeds of 1,200 tokens per second. This makes it one of the fastest AI search systems currently available.
The collaboration aims to challenge traditional search engines by providing near-instantaneous AI-powered search results. According to Perplexity's internal testing, Sonar outperforms existing models like GPT-4o mini and Claude 3.5 Haiku in user satisfaction metrics, with a factuality score of 85.1 out of 100.
Cerebras' AI inference infrastructure is central to Sonar's performance, enabling the model to deliver accurate and relevant information in real-time. The new search experience is initially available to Perplexity Pro users, with plans for broader availability in the future.
This partnership highlights a trend in the AI industry towards leveraging specialized hardware to gain competitive advantages. While the financial terms of the partnership were not disclosed, the companies aim to establish Sonar as a serious contender in the enterprise search market.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like Enterprise AI Brief or Daily AI Brief.
Also, consider following us on social media:
More from Enterprise
Sep 16 Crunchtime Launches Connected Restaurant Operations Suite Sep 16 Crusoe to Run Perplexity Model Training and Inference Sep 16 Siemens and Salesforce Connect Agentforce With Teamcenter Sep 16 GSMA Warns AI Chip Demand Could Widen Mobile Internet Gap Sep 16 RWS Opens Public Preview of Tridion Agentic PlatformEnterprise AI Brief
Weekly report on AI business applications, enterprise software releases, automation tools, and industry implementations.
Whitepaper
Tensordyne Napier: What If One Rack Could Do the Work of Nine?
Tensordyne
This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.
Read moreYou may also like
Perplexity Open Sources Lily Inference Engine for Apple Silicon
SEOPulse Launches AI Visibility Platform for Enterprise Marketing Teams
HUMAIN and KORA Systems Show Early Linux AI Operating System
Cyberhill Adds Cerebro Context Layer for Claude Enterprise
Enigmata Raises $6.5 Million for Encrypted AI Data Technology
Daily AI Brief: the AI news that matters, in your inbox.