myrtle.ai Sets STAC Records for Gradient Boosted Tree Inference
myrtle.ai said in a press release that VOLLO built on its April low latency inference record, setting new STAC-ML Markets benchmark records for gradient boosted tree inference.
In STAC audited tests, VOLLO recorded 99th percentile latency below 2 microseconds across all three models. For the smallest model, it processed 50 million inferences per second with 99th percentile latency of 1.77 microseconds. Myrtle.ai says the results reduced latency by more than 30% and increased throughput by at least five times compared with the previous best results.
The tests used an AMD Alveo V80LL Compute Accelerator in a Blackcore ICON 3132-SM+ server. Developers can also test their own models on VOLLO without using FPGA tools.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like Silicon Brief or Daily AI Brief.
Also, consider following us on social media:
More from Data Centers
Oct 6 Info-Tech Blueprint Maps AI Infrastructure to Workloads Oct 6 Cedar Investment Group Launches With More Than 2 GW of US Data Center Projects Oct 6 Hammerhead AI and Columbia Study AI Factory Power Control Oct 6 RMX details QuantrusX distributed data center plan Oct 6 DayOne Files for Proposed Nasdaq IPOSilicon Brief
Weekly coverage of AI hardware developments including chips, GPUs, cloud platforms, and data center technology.
Whitepaper
Tensordyne Napier: What If One Rack Could Do the Work of Nine?
Tensordyne
This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.
Read moreYou may also like
LILT Launches AURORA Multilingual AI Leaderboard
Voltropy Unveils Vast-10M With 10 Million Token Context
Delos Data Raises Over $100 Million for AI Infrastructure
Insilico Medicine Opens AI Longevity Research Toolkit
Corvex Launches Token Factory for Open Weight AI Inference
Daily AI Brief: the AI news that matters, in your inbox.