Scality Launches AI Inference Factory for On Premises Deployment

Oct 8, 2026
Scality AI Inference Factory combines open weight models, inference serving, workload controls and shared storage for running AI on infrastructure owned by customers.

Scality launches AI Inference Factory, an open code software stack for running AI inference on infrastructure owned by enterprises, government agencies and cloud providers. Available now as a software license or managed service, it gives customers control over models and data while offering more predictable costs than usage based cloud services.

The stack combines validated open weight models, separate prefill and decode services, and a control plane for authentication, metering, routing and workload scheduling. It supports models including Mistral, Gemma, Qwen, Kimi, GLM and DeepSeek, along with servers from Dell, HPE, Lenovo and Supermicro.

Scality ADI provides shared object storage for models, enterprise data, inference state and key value cache. Scality says its tests showed cache retrieval was 14 times faster than recomputing context for a 14,000 token prompt, while the shared cache held more than 80 times the memory capacity of one GPU. The design lets different GPU pools process prompts and generate tokens while accessing stored context from the same cache.

We hope you enjoyed this article

Consider subscribing to one of our newsletters like Silicon Brief or Daily AI Brief.

Also, consider following us on social media:

Free newsletter

Silicon Brief

Weekly coverage of AI hardware developments including chips, GPUs, cloud platforms, and data center technology.

Whitepaper

Tensordyne Napier: What If One Rack Could Do the Work of Nine?

Tensordyne

This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.

Read more
Free, six days a week

Daily AI Brief: the AI news that matters, in your inbox.