Parasail Combines NVIDIA GPUs with D-Matrix Accelerators for Faster Inference
Parasail is deploying D-Matrix Corp. Corsair accelerators with NVIDIA Hopper and Blackwell GPUs to deliver faster and more efficient inference services, announced in a press release. The integration aims to achieve up to ten times faster token generation and improved cost performance for Parasail customers.
The deployment marks one of the first large-scale examples of heterogeneous disaggregated inference, where NVIDIA infrastructure handles compute-intensive prefill tasks and D-Matrix Corsair accelerators manage latency-sensitive decode operations. Parasail said this approach allows it to extend the lifespan and utilization of its existing GPU fleet across data centers.
D-Matrix Corsair accelerators use Digital In-Memory Compute architecture that combines compute and memory on the same chip. This design reduces data movement between components, improving both speed and energy efficiency. According to the companies, Corsair enables up to ten times faster interactive inference and three times better energy efficiency than traditional GPU setups.
Parasail plans to scale this configuration across its network of more than forty data centers in fifteen countries. D-Matrix's Corsair platform is available to select customers, while Parasail's inference cloud services are already in production.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like Silicon Brief, Enterprise AI Brief or Daily AI Brief.
Also, consider following us on social media:
More from Data Centers
Oct 9 GlobalFoundries Plans FDX Fusion Chip Platform in Dresden Oct 9 CollPlant Announces $3.7 Million Private Placement for LightSolver Oct 9 Acbel to Show 100kW and 72kW AI Data Center Power Shelves Oct 9 Viega Piping Products Gain OCP Acceptance for Data Center Cooling Oct 9 AI Chip Market Forecast to Reach $677.59 Billion by 2035More from Enterprise
Oct 9 T-Mobile Uses AI to Guide Network Investment and Storm Response Oct 9 PDW Adds Booz Allen Autonomy Software to AM Drones Oct 9 Adobe brings AI search tools to Heathrow Airport website Oct 9 CaixaBank expands Google Cloud agreement through 2033 Oct 9 Strada Partners With Aon and Unveils Workday AI Cost ToolEnterprise AI Brief
Weekly report on AI business applications, enterprise software releases, automation tools, and industry implementations.
Whitepaper
Tensordyne Napier: What If One Rack Could Do the Work of Nine?
Tensordyne
This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.
Read moreYou may also like
Saturn Cloud Integrates NVIDIA Run:ai for GPU Inference Services
General Compute Signs Cerebras Inference Deal
Scality Launches AI Inference Factory for On Premises Deployment
Cerebras to Supply 100 Megawatts of AI Systems to Gimlet Labs
Spectrum Adds NVIDIA AI Compute to Its Edge Network
Daily AI Brief: the AI news that matters, in your inbox.