Google DeepMind Introduces Gemini 3.6 Flash, 3.5 Flash Lite, and 3.5 Flash Cyber

Jul 21, 2026
Google DeepMind has released three new Gemini models focused on efficiency, speed, and cybersecurity. The update includes 3.6 Flash, 3.5 Flash Lite, and 3.5 Flash Cyber, while the long-expected 3.5 Pro model remains in testing.
Google DeepMind Introduces Gemini 3.6 Flash, 3.5 Flash Lite, and 3.5 Flash Cyber

Google's DeepMind team has announced three new additions to its Gemini model family: Gemini 3.6 Flash, 3.5 Flash Lite, and 3.5 Flash Cyber. The models aim to improve efficiency and reliability for developers building large scale AI agents.

Gemini 3.6 Flash is described as the primary model in this release, offering better coding and multimodal performance while cutting token usage by up to 17 percent compared to the previous 3.5 Flash version. It operates at lower cost and includes enhanced safeguards against misuse in sensitive domains.

Gemini 3.5 Flash Lite focuses on low latency and high throughput workloads, producing up to 350 output tokens per second. It is the most cost-effective model in the 3.5 class and is designed for tasks such as search and document processing.

Gemini 3.5 Flash Cyber is a security-focused model created for identifying and resolving software vulnerabilities. It is being released through a limited access pilot for governments and trusted partners.

The company also confirmed that testing continues for Gemini 3.5 Pro, which has not yet been publicly released. Work has already begun on Gemini 4, described as DeepMind's most ambitious training run to date.

We hope you enjoyed this article

Consider subscribing to one of our newsletters like Cybersecurity AI Weekly or Daily AI Brief.

Also, consider following us on social media:

Free newsletter

Cybersecurity AI Weekly

Weekly newsletter about AI in Cybersecurity.

Whitepaper

Tensordyne Napier: What If One Rack Could Do the Work of Nine?

Tensordyne

This Tensordyne whitepaper presents Napier, an inference-focused AI processor and rack-scale system based on the company’s TDN Math logarithmic number system. It examines infrastructure requirements for large mixture-of-experts and agentic models, compares major inference architecture approaches, and details the TDN AIP processor, TDN72 pod, TDN Link fabric, and Napier Ultra configuration. The paper reports simulation-based performance, cost, and accuracy-validation results, including Tensordyne’s projected comparison of one Napier rack with a nine-rack Nvidia Rubin plus Groq deployment; the chip is reported as taped out and in fabrication.

Read more
Free, six days a week

Daily AI Brief: the AI news that matters, in your inbox.