Zendoric
← Back to the day · July 23, 2026

Google prioritizes efficiency with Gemini 3.6 Flash, 3.5 Flash-Lite and the specialized 3.5 Flash Cyber model

🕒 Published on Zendoric: July 23, 2026 · 00:24

Google has unveiled three new models in its Gemini family focused on a single goal: making production AI agents more efficient, faster and cheaper to operate at scale.

Google has unveiled three new models in its Gemini family focused on a single goal: making production AI agents more efficient, faster and cheaper to operate at scale. The three models —Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber— build on the foundation of Gemini 3.5 Flash and are designed specifically for agentic workflows, that is, systems where a model chains together multiple steps, tool calls and decisions with barely any constant human oversight.

The first, Gemini 3.6 Flash, is presented as the workhorse of the new generation, with improvements in coding, knowledge work and multimodal tasks. According to the Artificial Analysis index, it reduces output-token consumption by 17% compared with 3.5 Flash, and on specific benchmarks such as DeepSWE (from Datacurve) the reduction reaches 65%, all at a lower cost per output token. Google also reports quality improvements alongside that greater efficiency: on DeepSWE it goes from 37% to 49% accuracy with fewer unwanted code edits and fewer execution loops; on MLE Bench (machine learning research) it rises from 49.7% to 63.9%; and on OSWorld-Verified, which measures computer-use capabilities, it improves from 78.4% to 83.0%. Indeed, "computer use" becomes a native tool integrated into both the Gemini API and Gemini Enterprise. On professional knowledge tasks, measured with GDPval-AA v2, the model goes from 1,349 to 1,421 points. Customers such as Hebbia and Harvey highlight its multimodal ability to analyze documents, charts and data, and to draft reports. The price of 3.6 Flash is set at $1.50 per million input tokens and $7.50 per million output tokens, a figure lower than that of 3.5 Flash, which reduces the total cost per agentic task. On safety, the model incorporates strengthened "Frontier Safety" safeguards in the chemical, biological, radiological and nuclear (CBRN) and cyber-offense domains, with greater resistance to jailbreaks, while it has been trained to minimize refusals on legitimate uses.

The second model, Gemini 3.5 Flash-Lite, is positioned as the fastest and most economical in the 3.5 series, aimed at low-latency, high-volume tasks such as agentic search or mass document processing. According to Artificial Analysis, it reaches 350 output tokens per second. Its price is $0.30 per million input tokens and $2.50 per million output tokens, and Google claims it offers significantly higher quality than the previous generation, 3.1 Flash-Lite. Developers can adjust the model's level of "thinking": minimal or low levels for high-volume tasks that prioritize latency and cost, or higher levels for multi-step subagent workflows. It also incorporates computer use as an integrated tool. As for performance, it improves markedly over 3.1 Flash-Lite on Terminal-Bench 2.1 (54% versus 31%), on long context according to GDM-MRCR v2 (72.2% versus 60.1%) and on real-world task execution as measured by GDPval-AA v2 (1,140 versus 642). A striking data point is that, on several coding and agentic-task evaluations, 3.5 Flash-Lite even surpasses the 3 Flash model (of a higher-numbered generation), as in SWE-Bench Pro (54.2% versus 49.6%) and OSWorld-Verified (74.0% versus 65.1%), making it a faster and more capable alternative to the existing 2.5 and 3 Flash series.

The third announcement is Gemini 3.5 Flash Cyber, a specialized model built on 3.5 Flash and fine-tuned specifically to find and fix cybersecurity vulnerabilities at a lower cost per token than larger models. It is integrated within CodeMender, Google's code security agent, where multiple instances of 3.5 Flash Cyber work in a coordinated way to produce a single combined report. Google states that, on the CyberGym benchmark, the system achieves performance competitive with frontier models. Given the dual-use nature of this technology (it can be used both to defend and to attack systems), the company has opted for a controlled rollout: the model will be available exclusively to governments and trusted partners through CodeMender, as part of a limited-access pilot program, with the stated goal of giving defensive teams an advantage over potential attackers.

As for availability, 3.6 Flash and 3.5 Flash-Lite are already accessible to developers through the Gemini API in Google AI Studio and Android Studio (3.6 Flash also in Google Antigravity), to enterprises through the Gemini Enterprise Agent Platform and the Gemini Enterprise app, and to the general public in the Gemini app; in addition, 3.5 Flash-Lite is also being rolled out in Google Search.

The article closes with two notes on what comes next: Gemini 3.5 Pro is currently in testing with partners and Google plans to release it generally as soon as it is ready, while the team has already begun what they describe as its most ambitious pre-training run to date, aimed at a future Gemini 4, about which no further details are offered for now.

🔗 Related on Zendoric

Sources & references