Baidu Releases ERNIE-4.5-VL-28B-A3B-Thinking Multimodal AI Model
Baidu has introduced a new open-source multimodal AI model called ERNIE-4.5-VL-28B-A3B-Thinking, announced on its AI Studio platform. The model is designed to handle text, images, and video inputs while consuming significantly fewer computing resources than comparable systems from other major AI developers.
The model operates on a Mixture-of-Experts architecture with 28 billion total parameters but activates only 3 billion during inference. This selective activation allows it to perform complex reasoning tasks efficiently on a single 80GB GPU. Baidu states that the model matches the performance of leading systems while maintaining lower computational costs.
ERNIE-4.5-VL-28B-A3B-Thinking introduces several capabilities including visual reasoning, STEM problem solving, visual grounding, and video understanding. A distinctive feature, called “Thinking with Images,” enables the model to zoom in and out of images dynamically, improving its ability to analyze fine-grained visual details. It also supports tool calling for functions like image search and external data access.
The model is available under the Apache 2.0 license, allowing unrestricted commercial use. It supports deployment through multiple frameworks such as Transformers, vLLM, and Baidu’s FastDeploy toolkit. Developers can also fine-tune the model using ERNIEKit, Baidu’s training framework built on PaddlePaddle.
According to the official documentation, the model’s context window extends to 128,000 tokens and it supports both Chinese and English. Its design aims to make advanced multimodal reasoning more accessible to enterprises and researchers seeking efficient AI solutions.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like Daily AI Brief.
Also, consider following us on social media:
Daily AI Brief
Daily report covering major AI developments and industry news, with both top stories and complete market updates
Whitepaper
Tensordyne Napier: What If One Rack Could Do the Work of Nine?
Tensordyne
This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.
Read moreYou may also like
Overmind Open Sources Platform for Specialized AI Models
ModelBest Releases MiniCPM5-2B Model for Edge Devices
Abacus.AI Releases Three Smaug Open Weight Models
AutoTrust AI Releases JEV-27B for Self Hosted Agent Decisions
AskELIE Opens Operational AI Platform to Customers and Partners
Daily AI Brief: the AI news that matters, in your inbox.