Red Hat Launches llm-d for Scalable AI Inference

May 20, 2025
Red Hat has introduced llm-d, an open-source project aimed at enhancing generative AI inference at scale, with contributions from industry leaders like CoreWeave, Google Cloud, and NVIDIA.

Red Hat has launched llm-d, a new open-source initiative designed to address the growing need for scalable generative AI inference, announced in a press release. This project is supported by founding contributors such as CoreWeave, Google Cloud, IBM Research, and NVIDIA, and aims to make AI inference as ubiquitous as Linux.

The llm-d project leverages a native Kubernetes architecture, vLLM-based distributed inference, and AI-aware network routing to optimize compute resources and deliver AI inference at a massive scale. This approach is intended to meet the demanding service-level objectives of production environments without compromising performance.

Key innovations of llm-d include vLLM, which supports a wide range of accelerators, and Prefill and Decode Disaggregation, which separates AI input context and token generation into discrete operations. Additionally, the project features KV Cache Offloading to reduce memory burdens and AI-Aware Network Routing for efficient request scheduling.

The initiative has garnered support from industry leaders such as AMD, Cisco, Hugging Face, Intel, Lambda, and Mistral AI, as well as academic institutions like the University of California, Berkeley, and the University of Chicago. This collaboration underscores the industry's commitment to advancing large-scale AI inference capabilities.

We hope you enjoyed this article

Consider subscribing to one of our newsletters like Silicon Brief or Daily AI Brief.

Also, consider following us on social media:

Free newsletter

Silicon Brief

Weekly coverage of AI hardware developments including chips, GPUs, cloud platforms, and data center technology.

Industry analysis

2025 Global Business Services Agenda: Gen AI Takes Center Stage

The Hackett Group

This industry analysis by The Hackett Group explores the transformative impact of generative artificial intelligence (Gen AI) on global business services (GBS) in 2025. The study highlights the shift from exploration to acceleration of Gen AI initiatives, with 89% of executives advancing these projects to improve customer satisfaction, innovate products, and reduce costs. The report also discusses the challenges and strategies for successful Gen AI adoption, emphasizing the need for a technology-enabled operating model and the importance of reskilling the workforce.

Read more
Free, six days a week

Daily AI Brief: the AI news that matters, in your inbox.