Aakib Ansari.
Back to articles
News Brief

Gemini 3.5 Flash-Lite: Google's Cheapest Model Undercuts 3.6 Flash by 5x on Input

Md Aakib Ansari
Md Aakib AnsariWeb Developer & AI Tools Reviewer
Updated 3 min readModel: Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite: Google's Cheapest Model Undercuts 3.6 Flash by 5x on Input

Google DeepMind released Gemini 3.5 Flash-Lite on July 21, alongside Gemini 3.6 Flash and a restricted security model, positioning it as the budget entry point of the Flash family.

Flash-Lite costs $0.30 per million input tokens and $2.50 per million output — roughly a fifth of 3.6 Flash's input price and a third of its output price. It keeps the full 1-million-token context window shared across the Flash lineup, accepts text, image, speech, and video input, and scores 36 on the Artificial Analysis Intelligence Index, well above the median for models in its price tier. Output speed lands near the top of Artificial Analysis's rankings at over 460 tokens per second.

The tradeoffs show up outside the headline numbers. Long-context recall on GDM-MRCR v2 comes in at 72.2%, a sizable drop from 3.6 Flash's 91.8%, meaning the million-token window is present but less reliable at the far end. Time-to-first-token also runs long for a "fast" model — Google's own comparisons put it around 6-12 seconds depending on thinking mode, which makes it better suited to background batch jobs than live chat.

Google is pitching Flash-Lite for exactly that kind of work: high-volume, low-reasoning tasks like document processing, translation, search grounding, and subagents that handle one narrow job inside a larger multi-agent pipeline. It's a continuation of Google's pattern with 3.1 Flash-Lite before it — rather than compete on flagship reasoning, Google is pushing hardest on the tier where enterprises run the largest call volumes and price per task matters more than peak intelligence.

Related Articles

Meta Releases Muse Glimmer: Apache 2.0 Licensed 30B Local Agent Model
News Brief2 min read
Meta Releases Muse Glimmer: Apache 2.0 Licensed 30B Local Agent Model

Meta Superintelligence Labs has released Muse Glimmer, a 30-billion parameter open-weight model distilled from its proprietary Muse Spark flagship. Published under an Apache 2.0 license, Glimmer is purpose-built for offline, on-device agentic workloads like coding, debugging, and file management on consumer hardware.

Ant Group's inclusionAI Team Releases Ling 3.0 Flash FP8 Under MIT License
News Brief2 min read
Ant Group's inclusionAI Team Releases Ling 3.0 Flash FP8 Under MIT License

Ant Group's inclusionAI team has released Ling 3.0 Flash FP8, a highly efficient 124-billion parameter Mixture-of-Experts (MoE) model. Featuring an MIT license and a custom hybrid attention architecture, the model reduces active parameters to 5.1 billion per token, matching the performance of much larger models while dramatically lowering operational costs.

Liquid AI's LFM2.5-2.6B Matches Models Three Times Its Size on Agentic Tasks
News Brief2 min read
Liquid AI's LFM2.5-2.6B Matches Models Three Times Its Size on Agentic Tasks

Liquid AI released LFM2.5-2.6B on August 4, a 2.69B-parameter on-device model purpose-built for agentic workloads. Using a hybrid architecture of short convolution blocks and grouped query attention, it runs under 2.5 GB of memory and reaches approximately 220 tokens/s on Apple M5 Max — while matching or exceeding Qwen3.5-9B on tool use and instruction-following benchmarks.