Aakib Ansari.
Back to articles
News Brief

Gemini 3.6 Flash Ships Cheaper and Faster, But Not Smarter

Md Aakib Ansari
Md Aakib AnsariWeb Developer & AI Tools Reviewer
Updated 3 min readModel: Gemini 3.6 Flash
Gemini 3.6 Flash Ships Cheaper and Faster, But Not Smarter

Google shipped Gemini 3.6 Flash on July 21, alongside a smaller sibling, Gemini 3.5 Flash-Lite, and a specialized security model, Gemini 3.5 Flash Cyber, aimed at governments and trusted partners. The company positioned the release around token efficiency, not a raw intelligence jump.

The headline number is a flat line: Gemini 3.6 Flash scores 50 on the Artificial Analysis Intelligence Index, identical to its predecessor and behind GLM-5.2, Claude Sonnet 5, Grok 4.5, and GPT-5.6 Luna. What Google is actually selling is cost and speed. The model uses roughly 17% fewer output tokens than 3.5 Flash, cuts token usage by up to 65% on some DeepSWE coding tasks, and now costs $1.50 per million input tokens and $7.50 per million output tokens — undercutting the $9 output price it replaces. Coding precision improved too, with DeepSWE climbing from 37% to 49%. The knowledge cutoff moves forward, from January 2025 to March 2026.

This is Google's second Flash-tier update since May, arriving while the promised Gemini 3.5 Pro — teased at I/O for June — remains in partner testing with no new date. Google confirmed it has begun pretraining Gemini 4, suggesting the flagship gap is deliberate sequencing rather than delay.

The release lands in a month already crowded with launches — GPT-5.6, Grok 4.5, and Claude Sonnet 5 all shipped in the past three weeks — reinforcing a wider trend: as top-line scores converge, efficiency and cost are becoming the real competitive axis.

Related Articles

Meta Releases Muse Glimmer: Apache 2.0 Licensed 30B Local Agent Model
News Brief2 min read
Meta Releases Muse Glimmer: Apache 2.0 Licensed 30B Local Agent Model

Meta Superintelligence Labs has released Muse Glimmer, a 30-billion parameter open-weight model distilled from its proprietary Muse Spark flagship. Published under an Apache 2.0 license, Glimmer is purpose-built for offline, on-device agentic workloads like coding, debugging, and file management on consumer hardware.

Ant Group's inclusionAI Team Releases Ling 3.0 Flash FP8 Under MIT License
News Brief2 min read
Ant Group's inclusionAI Team Releases Ling 3.0 Flash FP8 Under MIT License

Ant Group's inclusionAI team has released Ling 3.0 Flash FP8, a highly efficient 124-billion parameter Mixture-of-Experts (MoE) model. Featuring an MIT license and a custom hybrid attention architecture, the model reduces active parameters to 5.1 billion per token, matching the performance of much larger models while dramatically lowering operational costs.

Liquid AI's LFM2.5-2.6B Matches Models Three Times Its Size on Agentic Tasks
News Brief2 min read
Liquid AI's LFM2.5-2.6B Matches Models Three Times Its Size on Agentic Tasks

Liquid AI released LFM2.5-2.6B on August 4, a 2.69B-parameter on-device model purpose-built for agentic workloads. Using a hybrid architecture of short convolution blocks and grouped query attention, it runs under 2.5 GB of memory and reaches approximately 220 tokens/s on Apple M5 Max — while matching or exceeding Qwen3.5-9B on tool use and instruction-following benchmarks.