Ahead of Alphabet's earnings release, Google unveiled three new Gemini models, targeting low-cost inference, high-throughput tasks, and cybersecurity scenarios, respectively. The new versions are already appearing in APIs, development tools, enterprise platforms, and some consumer-facing products, indicating that Google is pushing its Flash series towards larger-scale real-world deployments.
3.6 Flash reduces inference costs
The core of this update is Gemini 3.6 Flash. Google claims that this model performs better than the previous generation in coding, multimodal processing, and knowledge-based tasks, while reducing output token consumption. According to data cited in the article, the output token usage of 3.6 Flash is 17% lower than that of 3.5 Flash, with even greater reductions in some programming tests.
In terms of pricing, the Gemini 3.6 Flash has an input price of $1.50 per million tokens and an output price of $7.50 per million tokens, lower than its predecessor. For developers, this means that the cost of model invocation is expected to decrease further for similar tasks.
3.5 Flash-Lite focuses on high throughput

Another new model, Gemini 3.5 Flash-Lite, places even greater emphasis on speed. According to the data in the article, it is the fastest output version in the 3.5 series, with an output rate of up to 350 tokens per second, making it more suitable for tasks such as large-scale search, document processing, and multi-agent collaboration.
In terms of pricing, the input price of 3.5 Flash-Lite is $0.30 per million tokens, and the output price is $2.50 per million tokens. The article states that it significantly outperforms the previous generation 3.1 Flash-Lite in multiple tests, and even surpasses more expensive models in the same series in some proxy task evaluations.
Cybersecurity model limits the scope of openness
The third model, Gemini 3.5 Flash Cyber, is designed for cybersecurity vulnerability detection and remediation. It is based on Flash 3.5 and deployed within Google's CodeMender code security agent, where multiple agents collaboratively generate vulnerability reports.
This model is not fully open to the public. Because vulnerability discovery capabilities have a dual purpose, Google has limited access to government agencies and trusted partners, and is offering it through limited pilot programs. The article argues that this arrangement aims to allow defenders to discover and patch critical vulnerabilities earlier, while reducing the risk of misuse.
Gemini 4 pre-training has begun.
In addition to this announcement, Google also revealed that the Gemini 3.5 Pro is being tested with partners, with plans to expand its availability once conditions are right. Meanwhile, pre-training for the Gemini 4 has already begun.
On the competitive front, the article mentions that Chinese model manufacturers and Anthropic are accelerating the development of related products. Google's simultaneous launch of multiple sub-models also reflects that the competition in the large-scale model market is shifting from a single flagship competition to a comprehensive contest of cost, speed, scene adaptation, and security control.












