Inception Unveils Mercury: A New Era for AI Models

Feb 26, 2025
Inception has introduced Mercury, a diffusion-based large language model (dLLM) that promises to be up to 10 times faster and cheaper than current models. The Mercury family aims to revolutionize text generation with its speed and efficiency.

Inception, a Palo Alto-based company founded by Stanford professor Stefano Ermon, has introduced a new type of AI model called Mercury, which is part of their diffusion-based large language models (dLLMs). Announced on their website, Mercury is designed to be significantly faster and more cost-effective than existing large language models (LLMs).

The Mercury family of models leverages diffusion technology, traditionally used in image and audio generation, to enhance text generation capabilities. Unlike traditional LLMs that generate text sequentially, Mercury's diffusion models generate and refine text in parallel, allowing for faster processing speeds. This approach enables Mercury to achieve speeds of over 1000 tokens per second on standard hardware, a feat previously only possible with specialized chips.

Inception offers Mercury through an API and supports on-premises deployments, making it accessible for various enterprise applications. The company claims that its models can run up to 10 times faster and at a fraction of the cost of traditional models, providing a significant advantage in latency-sensitive applications. Mercury Coder, a model optimized for code generation, is already available for testing and has shown impressive performance on standard coding benchmarks.

With the introduction of Mercury, Inception aims to set a new standard for AI models, offering enhanced speed and efficiency that could transform how language models are utilized across industries.

We hope you enjoyed this article

Consider subscribing to one of our newsletters like Daily AI Brief.

Also, consider following us on social media:

Free newsletter

Daily AI Brief

Daily report covering major AI developments and industry news, with both top stories and complete market updates

Whitepaper

Tensordyne Napier: What If One Rack Could Do the Work of Nine?

Tensordyne

This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.

Read more
Free, six days a week

Daily AI Brief: the AI news that matters, in your inbox.