Cerebras Systems Launches Qwen3-32B for Real-Time AI Inference
Cerebras Systems has announced the availability of the Qwen3-32B model on its inference platform, enabling real-time AI reasoning at speeds previously unattainable. According to a company blog post, the Qwen3-32B model, developed by Alibaba, can perform advanced reasoning tasks in just 1.2 seconds, significantly faster than competing models.
The Qwen3-32B model operates at an output speed of 2,400 tokens per second, making it over 40 times faster than traditional GPU-based solutions. This performance is made possible by Cerebras' Wafer Scale Engine, which allows the model to be used in a wide range of applications without the latency issues that typically hinder reasoning models.
Cerebras' platform offers the Qwen3-32B model at a cost-effective rate of $0.80 per million output tokens, making it a competitive alternative to models like GPT-4.1. The company is encouraging developers to experiment with the model by providing 1 million free tokens per day, with no waitlist required for access.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like Silicon Brief or Daily AI Brief.
Also, consider following us on social media:
More from Data Centers
Sep 19 Virginia Proposes Data Center Rules and Creates AI Task Force Sep 19 House Passes Bill to Shield Ratepayers From Data Center Power Costs Sep 19 Laminar Joins L'Oreal Sustainability Accelerator Sep 19 Dnotitia Begins Testing VDPU ASIC Samples Sep 19 Nscale Files for US Initial Public OfferingSilicon Brief
Weekly coverage of AI hardware developments including chips, GPUs, cloud platforms, and data center technology.
Whitepaper
Tensordyne Napier: What If One Rack Could Do the Work of Nine?
Tensordyne
This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.
Read moreYou may also like
d-Matrix Raises $275 Million for AI Inference Chips
Perplexity Open Sources Lily Inference Engine for Apple Silicon
Tensordyne Details Napier AI Inference Chip and Claims One Rack Can Replace Nine
Huawei Introduces OceanStor M900 Storage for AI Inference
Acer Introduces Veriton RI110 AI Mini Workstation
Daily AI Brief: the AI news that matters, in your inbox.