Parasail Combines NVIDIA GPUs with D-Matrix Accelerators for Faster Inference

Jul 9, 2026
Parasail has announced it will deploy D-Matrix Corsair accelerators alongside NVIDIA Hopper and Blackwell GPUs to boost its inference cloud performance, promising up to ten times faster and more cost-efficient token generation for customers.

Parasail is deploying D-Matrix Corp. Corsair accelerators with NVIDIA Hopper and Blackwell GPUs to deliver faster and more efficient inference services, announced in a press release. The integration aims to achieve up to ten times faster token generation and improved cost performance for Parasail customers.

The deployment marks one of the first large-scale examples of heterogeneous disaggregated inference, where NVIDIA infrastructure handles compute-intensive prefill tasks and D-Matrix Corsair accelerators manage latency-sensitive decode operations. Parasail said this approach allows it to extend the lifespan and utilization of its existing GPU fleet across data centers.

D-Matrix Corsair accelerators use Digital In-Memory Compute architecture that combines compute and memory on the same chip. This design reduces data movement between components, improving both speed and energy efficiency. According to the companies, Corsair enables up to ten times faster interactive inference and three times better energy efficiency than traditional GPU setups.

Parasail plans to scale this configuration across its network of more than forty data centers in fifteen countries. D-Matrix's Corsair platform is available to select customers, while Parasail's inference cloud services are already in production.

We hope you enjoyed this article

Consider subscribing to one of our newsletters like Silicon Brief, Enterprise AI Brief or Daily AI Brief.

Also, consider following us on social media:

Free newsletter

Enterprise AI Brief

Weekly report on AI business applications, enterprise software releases, automation tools, and industry implementations.

Whitepaper

Tensordyne Napier: What If One Rack Could Do the Work of Nine?

Tensordyne

This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.

Read more
Free, six days a week

Daily AI Brief: the AI news that matters, in your inbox.