Inception Unveils Mercury: A New Era for AI Models
Inception, a Palo Alto-based company founded by Stanford professor Stefano Ermon, has introduced a new type of AI model called Mercury, which is part of their diffusion-based large language models (dLLMs). Announced on their website, Mercury is designed to be significantly faster and more cost-effective than existing large language models (LLMs).
The Mercury family of models leverages diffusion technology, traditionally used in image and audio generation, to enhance text generation capabilities. Unlike traditional LLMs that generate text sequentially, Mercury's diffusion models generate and refine text in parallel, allowing for faster processing speeds. This approach enables Mercury to achieve speeds of over 1000 tokens per second on standard hardware, a feat previously only possible with specialized chips.
Inception offers Mercury through an API and supports on-premises deployments, making it accessible for various enterprise applications. The company claims that its models can run up to 10 times faster and at a fraction of the cost of traditional models, providing a significant advantage in latency-sensitive applications. Mercury Coder, a model optimized for code generation, is already available for testing and has shown impressive performance on standard coding benchmarks.
With the introduction of Mercury, Inception aims to set a new standard for AI models, offering enhanced speed and efficiency that could transform how language models are utilized across industries.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like Daily AI Brief.
Also, consider following us on social media:
Daily AI Brief
Daily report covering major AI developments and industry news, with both top stories and complete market updates
Whitepaper
Tensordyne Napier: What If One Rack Could Do the Work of Nine?
Tensordyne
This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.
Read moreYou may also like
Google Launches Gemini 3.8 Live Voice Models
Perplexity Open Sources Lily Inference Engine for Apple Silicon
Enigmata Raises $6.5 Million for Encrypted AI Data Technology
ModelBest Releases MiniCPM5-2B Model for Edge Devices
Thomson Reuters Launches Proprietary AI Model Thomson
Daily AI Brief: the AI news that matters, in your inbox.