Zendoric
← Back to the day · July 23, 2026

Google launches cheaper Gemini 3.6 Flash and 3.5 Flash-Lite, and adds a model dedicated to cybersecurity

🕒 Published on Zendoric: July 23, 2026 · 00:24

Google DeepMind's official Gemini models page presents the "Gemini 3.5" family, which already includes several variants: Gemini 3.6 Flash (described as the most token-efficient for programming, knowledge work and multimodality), Gemini 3.5 Flash-Lite (aimed at high-…

Google DeepMind's official Gemini models page presents the "Gemini 3.5" family, which already includes several variants: Gemini 3.6 Flash (described as the most efficient in token consumption for programming, knowledge work and multimodal tasks), Gemini 3.5 Flash-Lite (aimed at high-volume tasks requiring efficiency and speed), Gemini 3.1 Pro (for complex and creative tasks) and Gemini 3.1 Deep Think (designed for scientific, research and engineering challenges). The page itself announces that a "3.5 Pro" is "coming soon," indicating that the rollout of this generation is not yet complete.

A central point of the content is a per-million-token pricing table that reveals a cost-cutting strategy against the previous generation and against competitors. According to that table, Gemini 3.6 Flash costs $1.50 for input and $7.50 for output per million tokens (without caching), while Gemini 3.5 Flash costs $1.50 for input and $9.00 for output, and Gemini 3.1 Pro rises to $2.00 for input and $12.00 for output. The same table compares these prices with those of other models cited in the source: GPT-5.6 ($1.00 / $6.00), a model called "Luna" ($2.00 / $6.00), Grok 4.5 ($3.00 / $15.00) and Claude Sonnet 5 ($2.00 / $10.00, flagged as a "temporary discount"). Google's implicit message is that Gemini 3.6 Flash offers a more competitive price-performance ratio within its own range and against other models on the market.

As for performance, the benchmark table compares the same models on tests such as SWE-Bench Pro (diverse agentic coding tasks), DeepSWE v1.1 (long-horizon software engineering), Terminal-bench 2.1 (agentic coding in the terminal), MLE-Bench (machine learning engineering), GDPVal-AA v2 (knowledge work, measured in Elo), OSWorld-Verified (agentic computer use), CharXiv Reasoning (synthesizing information from complex charts, with and without tools) and GDM-MRCR v2 (long-context performance, with tests at 128k and 1M tokens). The figures show uneven results depending on the test: for example, on SWE-Bench Pro Gemini 3.6 Flash scores 58.7% versus GPT-5.6's 62.7% or "Luna's" 64.7%, while on MLE-Bench Gemini 3.6 Flash scores 63.9%, a high figure though not the highest in the table (the source records a higher result, 66.9%, on that test), and on the long-context tests (GDM-MRCR v2) Gemini stands out particularly at 128k and 1M tokens.

A relevant element for security is the explicit mention, within the Gemini product ecosystem, of "Gemini 3.5 Flash Cyber," described literally as a model designed to "find and fix vulnerabilities quickly and efficiently." This positions cybersecurity as one of the specific applications Google promotes within the Gemini 3.5 family, alongside other specialized models such as Gemini 3.1 Deep Think (science and research), Gemini Omni (multimodal generation), Gemini Image/Nano Banana (image generation and editing), Gemini Audio, Gemini Robotics and Gemini Embedding 2.

The page also includes testimonials from customers and partners evaluating these models in production. Ashler highlights that Gemini 3.5 Flash-Lite offers the intelligence, speed and cost efficiency needed for information retrieval and tool-use tasks across fragmented infrastructures. Figma notes that Gemini 3.6 Flash balances quality, speed and cost for iterating on prototypes with its Figma Make tool. Harvey, specialized in legal documentation (capital markets and corporate M&A), states that Gemini 3.6 Flash completes tasks 12% faster on average than its predecessor. Hebbia indicates that this model outperformed other reference baselines in evidence retrieval within citation-intensive financial research. JetBrains (through its Junie product) mentions that Gemini 3.6 Flash improves low-reasoning coding performance by 10% to 20% over the previous Flash generation. Palo Alto Networks describes Gemini 3.5 Flash-Lite as a major leap over Gemini 3.1 Flash-Lite in speed and reliability, and Ramp notes that this model sits on the "Pareto frontier" in its receipt-extraction benchmark, balancing accuracy, latency and cost.

The source also reports enterprise adoption cases of Gemini 3.5 Flash: Shopify uses it to run parallel subagents that analyze complex long-term data and improve merchant growth forecasts on a global scale; Macquarie Bank is testing the model to speed up client onboarding by reasoning over documents longer than 100 pages; Salesforce integrates it into Agentforce to automate complex business tasks through multiple subagents with context memory; Ramp employs it to improve optical character recognition (OCR) on complex invoices; Xero uses it to autonomously manage administrative workflows spanning several weeks, such as identifying suppliers for 1099 tax forms; and Databricks applies it in agentic workflows to monitor real-time information and diagnose problems in large datasets.

Finally, the page frames this entire launch within Google's ecosystem of tools for developers and businesses: Google Antigravity (AI development platform), Google AI Studio (to go from prompt to production), the Gemini API and Gemini Enterprise (a platform to build, scale and govern agents). It should be clarified that the analyzed content corresponds to a product/marketing page from Google DeepMind, with abundant navigation and promotional material, and not to a narrative journalistic article; the data summarized here are those that appear explicitly on that page, without adding figures, dates or claims that are not in the original text.

🔗 Related on Zendoric

Sources & references